Alibaba/Qwen

Latest AI news, models and releases from Alibaba/Qwen. ['CodeQwen1.5', 'Cotype Light 3', 'Cotype Pro 3', 'Fun-Realtime-TTS', 'Hanguang', 'Hanguang 800', 'HappyOyster 1.0', 'ICN Switch 1.0', 'not specified', 'Panjiu', 'Qianwen', 'Qoder Security', 'QVQ-72B-Preview', 'QVQ-Max', 'Qwen', 'Qwen 2', 'Qwen2.5', 'Qwen2.5-0.5B', 'Qwen2.5-14B', 'Qwen2.5-14B-Instruct-1M', 'Qwen2.5-1.5B', 'Qwen2.5-32B', 'Qwen2.5-3B', 'Qwen-2.5 72B', 'Qwen2.5-72B', 'Qwen2.5-7B', 'Qwen2.5-7B-Instruct-1M', 'Qwen2.5-Coder', 'Qwen2.5-Coder-0.5B', 'Qwen2.5-Coder-0.5B-Instruct', 'Qwen2.5-Coder-14B', 'Qwen2.5-Coder-1.5B', 'Qwen2.5-Coder-1.5B-Instruct', 'Qwen2.5-Coder-32B', 'Qwen2.5-Coder-32B-Instruct', 'Qwen2.5-Coder-3B', 'Qwen2.5-Coder-3B-Instruct', 'Qwen2.5-Coder-7B', 'Qwen2.5-Coder-7B-Instruct', 'Qwen2.5-Coder-Instruct']

Research 🇺🇸

New DeepSeek-V3 technical report: how hardware-software co-design enables low-cost training of large models

A new technical paper from the DeepSeek team, with CEO Wenfeng Liang as co-author, explores hardware-aware model co-design to reduce LLM training costs. Using DeepSeek-V3 trained on 2048 NVIDIA H800 GPUs as a case study, the paper details innovations in memory efficiency (MLA), sparse computation (DeepSeekMoE), FP8 training, and interconnect-aware routing.

DeepSeekDeepSeek NVIDIANVIDIA Alibaba/QwenAlibaba/Qwen MetaMeta
Synced27.07 · 18:04
Fresh news