BENCHMARKS

Benchmarks

How frontier AI models are evaluated. Leaderboards, scores, and methodology for every major benchmark in the industry.

BENCHMARKS
10
tracked
CATEGORIES
5
evaluation types
ENTRIES
75
model scores
LABS
6
represented
CATEGORY 10 benchmarks
Aider Polyglot Coding

Evaluates AI models on code editing tasks across multiple programming languages using the Aider coding assistant framework. Tests real-world edit accuracy, not just generation.

MODELS 7 TOP 82.7% VOL ~4K/mo
Chatbot Arena Elo General

Crowdsourced human preference rankings where users rate blind side-by-side model responses. The gold standard for measuring real-world conversational quality.

MODELS 8 TOP 1407Elo VOL ~25K/mo
GPQA Diamond Reasoning

Graduate-Level Google-Proof Q&A — expert-crafted questions in physics, chemistry, and biology that require PhD-level domain knowledge and multi-step reasoning.

MODELS 7 TOP 74.9% VOL ~5K/mo
HellaSwag Reasoning

Tests commonsense reasoning by asking models to select the most plausible continuation of everyday scenarios. Designed to be trivial for humans but challenging for AI.

MODELS 8 TOP 97.1% VOL ~3K/mo
HumanEval Coding

Measures functional correctness of AI-generated Python code across 164 hand-written programming problems with unit test verification.

MODELS 8 TOP 94.5% VOL ~8K/mo
MATH Reasoning

Tests AI models on 5,000 competition-level mathematics problems spanning algebra, geometry, number theory, counting and probability, and precalculus.

MODELS 8 TOP 83.2% VOL ~6K/mo
MMLU Language

Massive Multitask Language Understanding — tests knowledge across 57 academic subjects including STEM, humanities, social sciences, and professional domains.

MODELS 8 TOP 90.2% VOL ~15K/mo
MMMU Multimodal

Massive Multi-discipline Multimodal Understanding — tests vision-language models on college-level problems requiring interpretation of images, diagrams, charts, and figures across 30 subjects.

MODELS 7 TOP 72.7% VOL ~3K/mo
MT-Bench General

Evaluates multi-turn conversational ability using GPT-4 as a judge across writing, roleplay, reasoning, math, coding, extraction, STEM, and humanities categories.

MODELS 8 TOP 9.52/10 VOL ~4K/mo
SWE-bench Verified Coding

Evaluates AI models on real-world software engineering tasks drawn from GitHub issues in popular Python repositories. The "Verified" subset uses human-validated test cases.

MODELS 6 TOP 72.5% VOL ~20K/mo