RSS Search All 🟢 Status
ResearchMedia Generation 🇷🇺 28.07.2026 08:04

AI Model Developers: Are They Pelicanmaxing?

OpenAIOpenAI AnthropicAnthropic Google/DeepMindGoogle/DeepMind xAIxAI Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek
A researcher tested seven modern AI models using a popular informal benchmark: generating an SVG of a pelican on a bicycle. The results show no evidence that models are optimized for this specific prompt, addressing suspicions of benchmark maximization.
Simon Willison has been testing major LLM releases with the prompt 'Generate an SVG of a pelican on a bicycle', which became a well-known informal benchmark. To investigate whether AI labs are engaging in 'pelicanmaxing' (optimizing models for popular benchmarks), a researcher conducted an experiment. They generated 1008 SVGs across seven models: GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro, using 48 prompts (8 animals × 6 vehicles). Each SVG was rendered and evaluated by an LLM judge. Analysis found no evidence that pelican-on-bicycle images scored higher than expected, suggesting no targeted optimization for this benchmark.
Сокращения
LLM = Large Language Model — Большая языковая модель
SVG = Scalable Vector Graphics — Масштабируемая векторная графика
Source: Habr — хаб ИИ — original
Our earlier posts on this topic ↓
Fresh news