AI Model Benchmark for 1C Development: Updated Ranking v2
Anthropic
OpenAI
Moonshot AI
xAI
Cursor
DeepSeek
Alibaba/Qwen
MiniMax
An updated benchmark for AI models in 1C development has been released, comparing models like Claude Opus 5, GPT 5.6 Sol, and Kimi K3. The benchmark categorizes tasks and measures token usage, with Opus 5 leading overall but GPT 5.6 Sol excelling with MCP. The author provides personal recommendations for model selection.
The benchmark for AI models in 1C development has been updated to version 2, including models such as Claude Opus 5, Claude Fable 5, GPT 5.6 Sol, Kimi K3, Claude Opus 4.8, Grok 4.5, Composer 2.5, GPT 5.5, DeepSeek v4, GLM 5.2, Qwen 3.8 Max, Claude Sonnet 5, Mimi 2.5, and MiniMax M2.7. Tasks are divided into categories like algorithms, architecture, platform mechanisms, performance, metadata work, and form development. The benchmark tracks tokens spent (input + output) excluding cache. Models were tested in their vendor environments: Claude Code for Claude, Codex for GPT, Kimi Code for Kimi, Cursor for Grok and Composer. A variant with MCP and skills/rules is also included. The benchmark was executed by Hermes with GPT 5.6 Sol, with some corrections. Claude Opus 5 is the overall leader, but GPT 5.6 Sol is best with MCP due to better instruction following. Kimi K3 is promising but slow. DeepSeek v4 and GLM 5.2 are only for budget use. Qwen 3.8, Sonnet 5, Mimi 2.5, and MiniMax M2.7 are not recommended for 1C development. The author's personal stack includes Cursor with Claude Opus for complex tasks, Grok/Composer for everyday work, and GPT 5.6 for review and autonomous agents.
- Сокращения
- MCP = Model Context Protocol — Протокол контекста модели
- ACP = Agent Communication Protocol — Протокол взаимодействия агентов
Source: Habr — хаб ИИ —
original
