QwQ-32B: Harnessing the Power of Reinforcement Learning
Alibaba Qwen unveiled the QwQ-32B model with 32 billion parameters, which achieves performance comparable to DeepSeek-R1 (671B parameters) through scalable reinforcement learning (RL). The model is open-sourced under the Apache 2.0 license and combines reasoning with agent capabilities.
Alibaba/Qwen
DeepSeek
OpenAI


