DeepSeek hints at R2 model and introduces new inference scaling method SPCT
DeepSeek
OpenAI
DeepSeek AI has published a paper on inference-time scaling for general reward models, introducing SPCT (Self-Principled Critique Tuning). The company hints at its upcoming R2 model. The research addresses reward sparsity and aims to improve scalability of reward models during inference.
DeepSeek AI published a paper titled 'Inference-Time Scaling for Generalist Reward Modeling' introducing a new method called Self-Principled Critique Tuning (SPCT). SPCT enhances general reward models (GRMs) by dynamically generating principles and critiques through rejection fine-tuning and rule-based online reinforcement learning. The approach aims to improve scalability during inference, addressing reward sparsity—a major hurdle in scaling reinforcement learning for LLMs. DeepSeek's R1 series already validated pure RL training; the company now hints at the next-generation R2 model, expected to build on these advances. The shift in LLM scaling from pre-training to post-training, especially inference phase, follows models like OpenAI's o1, which uses extensive internal chain-of-thought reasoning.
- Сокращения
- LLM = Large Language Model — большая языковая модель
- RL = Reinforcement Learning — обучение с подкреплением
- GRM = General Reward Model — общая модель вознаграждения
- SPCT = Self-Principled Critique Tuning — самопринципиальная критическая настройка
- RM = Reward Model — модель вознаграждения
Source: Synced —
original
