ModelsAgents 🇨🇳 04.08.2026 10:02

Between Kimi K3 and DeepSeek V4: The Time Gap of Native Multimodality

Moonshot AIMoonshot AI DeepSeekDeepSeek OpenAIOpenAI Alibaba/QwenAlibaba/Qwen ByteDanceByteDance TencentTencent AppleApple
The article discusses the strategic importance of native multimodal capabilities in AI models, using Kimi K3 as a case study. It highlights how visual feedback is becoming crucial for agent tasks, and contrasts different approaches among Chinese AI labs regarding when and how to integrate vision into large language models.
The article from 36Kr discusses how native multimodal capabilities create a competitive gap between models like Kimi K3 and DeepSeek V4. A multimodal researcher notes that for long-chain tasks, code-level feedback alone leads to error accumulation, while visual feedback is more accurate and closer to user intent. The piece explains that native multimodality means images and text jointly shape the main model from pre-training, continuing through post-training and agent perception. Kimi K3 topped the Arena Frontend Code leaderboard with 1679 points, and in a test by Puter, it identified all five visual deviations without false positives. The article contrasts this with DeepSeek's approach, which focuses on text-only models and external tools, noting that despite DeepSeek V4's strong coding and reasoning, it lacks native vision. Other Chinese labs like Alibaba and ByteDance are also following the native multimodal path. The piece discusses the costs: integrating vision can dilute language, reasoning, and coding abilities, requiring more compute and data. K3 uses a 2.8-trillion-parameter model with 104 billion activated parameters per token, and a custom vision encoder MoonViT-V2. The article notes the trade-off between investing in coding now versus preparing for multimodal futures, and mentions that talent hunters are flocking to Moonshot AI's vicinity after K3's release.
Abbreviations
LLM = Large Language Model — большая языковая модель
OCR = Optical Character Recognition — оптическое распознавание символов
VLM = Vision Language Model — визуально-языковая модель
MCP = Model Context Protocol — протокол контекста модели
Source: 36Kr — original
Our earlier posts on this topic ↓
Fresh news