State of LLMs in 2025: Progress, Challenges, and Predictions
In 2025, LLM development was dominated by reasoning models using reinforcement learning with verifiable rewards (RLVR) and the GRPO algorithm, sparked by DeepSeek R1. Key trends include a focus on inference-time scaling, MoE architectures, and efficiency tweaks like Gated DeltaNets. Academic research highlighted GRPO variants with improvements such as zero gradient filtering and token-level loss.
DeepSeek
OpenAI
Alibaba/Qwen
Moonshot AI
NVIDIA
Google/DeepMind

