New Benchmark Confirms: AI Models Still See Poorly
Moonshot AI
OpenAI
Anthropic
Google/DeepMind
Alibaba/Qwen
Moonshot AI has introduced PerceptionBench, a benchmark that isolates visual perception from reasoning and world knowledge. In tests, all leading models including GPT-5.6 Sol, Kimi K3, and Claude Fable 5 showed significant weaknesses, with none exceeding 60% accuracy. The authors conclude that many supposed reasoning errors are actually caused by faulty visual perception, confirming earlier studies.
Moonshot AI, the team behind Kimi, has released PerceptionBench, a benchmark designed to test visual perception in multimodal language models in isolation. Unlike common benchmarks, PerceptionBench breaks down vision into ten atomic sub-skills, with each question answerable by looking alone, without reasoning or external knowledge. The taxonomy was derived from real model errors, categorized into areas such as Visual Relation, Counting, Attribute, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. Moonshot AI published 3,000 questions from an internal pool of over 17,000 verified ones. Across 16 frontier models, the highest overall accuracy is 59.7% by GPT-5.6 Sol, followed by Kimi K3 at 58.5%, Claude Fable 5 at 57.2%, and Gemini 3.1 Pro at 56.2%. Open-source models lag significantly. The hallucination category is the weakest on average: GPT-5.6 Sol gets only 26.9% there, while the weaker Gemini 3.5 Flash gets 50.6%. The authors argue that many perceived reasoning errors actually occur at the perception level. Moonshot AI had previously released WorldVQA and contributed to BabyVision, which also showed frontier models failing at basic visual tasks, with humans outperforming them. The researchers attribute this to a verbalization bottleneck.
- Abbreviations
- OCR = Optical Character Recognition — распознавание символов
Source: The Decoder (DE) —
original
