Supabase Open Sources Evals: Benchmarking AI Coding Agents on Real Tasks
Supabase has open-sourced Supabase Evals, a benchmark and framework that tests how well AI coding agents build with Supabase. It runs agents like Claude Code, Codex, and OpenCode against real tasks, scoring results with a mix of deterministic checks and LLM-as-a-judge. The framework is available under Apache-2.0 and powers a public leaderboard.
Anthropic
OpenAI
Moonshot AI





