How frontier AI models are evaluated. Leaderboards, scores, and methodology for every major benchmark in the industry.
Evaluates AI models on code editing tasks across multiple programming languages using the Aider coding assistant framework. Tests real-world edit accuracy, not just generation.
Crowdsourced human preference rankings where users rate blind side-by-side model responses. The gold standard for measuring real-world conversational quality.
Graduate-Level Google-Proof Q&A — expert-crafted questions in physics, chemistry, and biology that require PhD-level domain knowledge and multi-step reasoning.
Tests commonsense reasoning by asking models to select the most plausible continuation of everyday scenarios. Designed to be trivial for humans but challenging for AI.
Measures functional correctness of AI-generated Python code across 164 hand-written programming problems with unit test verification.
Tests AI models on 5,000 competition-level mathematics problems spanning algebra, geometry, number theory, counting and probability, and precalculus.
Massive Multitask Language Understanding — tests knowledge across 57 academic subjects including STEM, humanities, social sciences, and professional domains.
Massive Multi-discipline Multimodal Understanding — tests vision-language models on college-level problems requiring interpretation of images, diagrams, charts, and figures across 30 subjects.
Evaluates multi-turn conversational ability using GPT-4 as a judge across writing, roleplay, reasoning, math, coding, extraction, STEM, and humanities categories.
Evaluates AI models on real-world software engineering tasks drawn from GitHub issues in popular Python repositories. The "Verified" subset uses human-validated test cases.
6:30 AM PT, Monday through Friday. Free.