Managing Reasoning Levels in Large Language Models
The article explains how reasoning models like DeepSeek-R1 and GPT-5.6 learn to produce reasoning traces via reinforcement learning with verifiable rewards (RLVR). It details the training process, inference scaling, think tokens, and how reasoning effort modes (low/medium/high) are implemented through supervised fine-tuning and tokenizer switches.
OpenAI
DeepSeek
Moonshot AI
