What We Learned by Reproducing 2,200 ICML Papers
Hugging Face, in partnership with alphaXiv, ran the ICML 2026 Open Reproductions challenge, where participants used AI coding agents to attempt reproduction of 2,200 papers. The results: 51% of examined papers had at least one claim verified, 23% had a claim falsified or contested, and 242 papers had conflicting verdicts. Human oversight proved crucial, as agents hit limits and false falsifications occurred; the event is considered the largest claim-level audit of an ML conference to date.
Hugging Face
Anthropic
OpenAI
Cursor
alphaXiv
Hugging Face blog13.08 · 20:03
