ModelsOpen Source 🇨🇳 28.07.2026 17:04

Alibaba Extends Qwen2.5-Turbo Context Length to 1 Million Tokens

Alibaba/QwenAlibaba/Qwen OpenAIOpenAI
Alibaba releases Qwen2.5-Turbo with a context length extended from 128K to 1M tokens. The model achieves 100% accuracy on the 1M Passkey Retrieval task and scores 93.1 on the RULER benchmark, surpassing GPT-4. Inference speed is improved by up to 4.3x using sparse attention, while costs remain at ¥0.3 per 1M tokens.
Alibaba has announced the Qwen2.5-Turbo model, extending its context length from 128K to 1M tokens, equivalent to about 1 million English words or 1.5 million Chinese characters. The model achieves 100% accuracy in the 1M length Passkey Retrieval task and scores 93.1 on the long text evaluation benchmark RULER, outperforming GPT-4's 91.6 and GLM4-9B-1M's 89.9. Using sparse attention mechanisms, the time to first token for a 1M-token context is reduced from 4.9 minutes to 68 seconds, a 4.3x speedup. The price remains ¥0.3 per 1M tokens, allowing Qwen2.5-Turbo to process 3.6 times more tokens than GPT-4o-mini at the same cost. The model maintains competitive short-sequence capabilities, on par with GPT-4o-mini. It is available via Alibaba Cloud Model Studio API, HuggingFace Demo, and ModelScope Demo.
Сокращения
API = Application Programming Interface — интерфейс прикладного программирования
TTFT = Time to First Token — время до первого токена
Source: Alibaba Qwen — original
Our earlier posts on this topic ↓
Fresh news