6 Lessons for Creating Benchmarks from NeurIPS Tutorial
The authors of SWE-bench and GPQA at NeurIPS discussed key mistakes in creating benchmarks: tyranny of metrics, subjectivity of crowdsourcing, danger of using LLMs as evaluators, and the need for dynamic tests. Main advice — avoid saturated benchmarks, check multimodality, and publish results with confidence intervals.
Google DeepMind

