Qwen2.5-Max: Exploring the Capabilities of a Large-Scale MoE Model
Alibaba/Qwen
Alibaba Qwen released Qwen2.5-Max, a large language model based on the Mixture-of-Experts architecture, trained on over 20 trillion tokens. The model outperforms DeepSeek V3 on several benchmarks and is available through Alibaba Cloud APIs and Qwen Chat.
The Qwen team has released Qwen2.5-Max, a large language model based on the Mixture-of-Experts (MoE) architecture, pre-trained on over 20 trillion tokens. After pre-training, supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) were applied. The model was compared against DeepSeek V3, GPT-4o, and Claude-3.5-Sonnet on the MMLU-Pro, LiveCodeBench, LiveBench, Arena-Hard, and GPQA-Diamond benchmarks. Qwen2.5-Max outperformed DeepSeek V3 in Arena-Hard, LiveBench, LiveCodeBench, and GPQA-Diamond, while showing competitive results in MMLU-Pro. For base models, comparisons were made with DeepSeek V3, Llama-3.1-405B, and Qwen2.5-72B — Qwen2.5-Max demonstrated significant advantages in most tests. The model is available in Qwen Chat and via the Alibaba Cloud API (model name: qwen-max-2025-01-25). The API is compatible with the OpenAI API.
Source: Alibaba Qwen —
original
