Home · Benchmarks · GPQA Diamond
BENCHMARK

GPQA Diamond

Graduate-Level Google-Proof Q&A — expert-crafted questions in physics, chemistry, and biology that require PhD-level domain knowledge and multi-step reasoning.

Reasoning UNIT % HIGHER IS BETTER MODELS 7
Leaderboard 7 models
# Model Lab Score Date
1 Claude Opus 4 Anthropic
74.9%
MAY 2025
2 Gemini 2.5 Pro Google DeepMind
71.4%
MAR 2025
3 GPT-4.1 OpenAI
66.3%
APR 2025
4 Claude Sonnet 4 Anthropic
65%
MAY 2025
5 Grok-3 xAI
62.7%
FEB 2025
6 DeepSeek V3 DeepSeek
59.4%
JAN 2025
7 GPT-4o OpenAI
53.6%
MAY 2024
About this benchmark

GPQA (Graduate-Level Google-Proof Q&A) Diamond is the benchmark that most convincingly tests whether AI models possess genuine scientific reasoning — not just pattern matching or factual retrieval, but the ability to apply deep domain knowledge across physics, chemistry, and biology to solve problems that PhD-level experts designed to be unsearchable. Created by David Rein and colleagues at New York University in 2023, GPQA Diamond has become the gold standard for evaluating expert-level scientific reasoning in language models, precisely because its questions resist the shortcuts that inflate scores on easier benchmarks.

What GPQA Diamond Measures

GPQA tests cross-disciplinary scientific reasoning at the graduate and professional level. Each question was written by a researcher holding a PhD in the relevant field — physics, chemistry, or biology — and validated by a second PhD-holding expert. The defining design principle is “Google-proofness”: questions are constructed so that someone without domain expertise cannot find the answer through web search, textbook lookup, or surface-level reasoning. Answering correctly requires integrating multiple concepts, applying theoretical frameworks, and performing multi-step reasoning chains.

The Diamond subset is the hardest tier of GPQA, containing 198 questions selected for maximum difficulty and reliability. The full GPQA dataset includes easier tiers (Main and Extended), but Diamond has become the standard reporting target because it provides the most meaningful separation between frontier models. Each question offers four answer options, establishing a 25% random-chance baseline.

The benchmark’s human calibration data reveals its difficulty. Domain experts — PhD holders answering questions in their own field — average approximately 81% accuracy. But when those same experts answer questions outside their specialty, accuracy drops to around 34%, barely above random chance. This means GPQA questions genuinely require deep expertise, not general scientific literacy.

Current Leaderboard Analysis

The GPQA Diamond leaderboard shows meaningful stratification that other saturated benchmarks no longer provide. Claude Opus 4 leads at 74.9%, followed by Gemini 2.5 Pro at 71.4% — a 3.5-point gap that represents a tangible difference in scientific reasoning capability. GPT-4.1 at 66.3% and Claude Sonnet 4 at 65.0% form a clear second tier, roughly 8-10 points behind the leaders.

The spread from top (74.9%) to bottom (53.6% for GPT-4o) is 21.3 percentage points — far wider than MMLU’s 7.1-point spread or HumanEval’s 7.4-point range. This confirms that GPQA Diamond still provides strong discriminative power between models of different capability levels.

A striking observation: Claude Opus 4’s 74.9% approaches in-domain expert accuracy (81%) while dramatically exceeding out-of-domain expert accuracy (34%). The model appears to function as a competent cross-domain scientist — something no individual human expert can achieve across all three disciplines simultaneously. This represents a qualitatively different capability from anything measured by knowledge-recall benchmarks.

The gap between GPT-4o (53.6%, from May 2024) and the latest models demonstrates the speed of progress. In roughly one year, the best-performing model improved by 21.3 percentage points, driven primarily by advances in chain-of-thought reasoning and extended thinking capabilities.

Methodology Deep Dive

The Diamond subset contains 198 multiple-choice questions, each with four answer options. Questions are distributed across three scientific disciplines: physics (including quantum mechanics, statistical mechanics, and particle physics), chemistry (organic, inorganic, and physical chemistry), and biology (molecular biology, genetics, and ecology).

Question creation follows a rigorous protocol. A PhD expert writes a question intended to be answerable only by someone with genuine domain expertise. A second PhD expert in the same field validates the question by attempting to answer it. The question must meet several criteria: it should be answerable by experts in the field (greater than 65% accuracy for in-domain experts), unanswerable by non-experts through search (below 40% accuracy for out-of-domain experts), and have exactly one correct answer that is not debatable among experts.

Scoring is straightforward accuracy — the percentage of questions answered correctly. With only 198 questions in the Diamond set, each question is worth approximately 0.5 percentage points, meaning that statistical noise from a handful of questions can shift scores by 1-2 points. This limited sample size means differences smaller than roughly 3 percentage points should be interpreted cautiously.

Most labs evaluate GPQA with chain-of-thought prompting, allowing the model to reason through the problem step-by-step before selecting an answer. This is important because direct answering (without reasoning traces) typically produces scores 5-10 points lower, and the benchmark is explicitly designed to require multi-step reasoning.

Why This Benchmark Matters

GPQA Diamond matters because it is one of the few benchmarks that measures genuine reasoning capability rather than knowledge retrieval. Many benchmarks that appear to test understanding can actually be solved through sophisticated pattern matching or memorization of training data. GPQA’s Google-proof design specifically neutralizes these shortcuts — a model cannot score well by retrieving a memorized answer from its training data because the questions are novel and require synthesizing multiple concepts.

The benchmark also provides insight into a capability that has significant practical implications: AI-assisted scientific research. A model that scores above 70% on GPQA Diamond is demonstrating the ability to reason about graduate-level science problems across multiple disciplines. This has direct relevance for applications in drug discovery, materials science, climate modeling, and other fields where AI could accelerate research by connecting insights across disciplinary boundaries.

GPQA is also one of the benchmarks most closely watched by AI safety researchers, because expert-level scientific reasoning is a capability threshold that has implications for AI risk assessment. Models that can reason at PhD level across multiple sciences represent a qualitatively different capability profile than models that merely recall facts.

Known Limitations and Criticisms

The most significant limitation is sample size. With only 198 questions, statistical confidence intervals are wide. A model answering 148 questions correctly (74.7%) versus 145 (73.2%) is within the range of normal variance, yet these scores might be reported as meaningfully different. Researchers should present confidence intervals alongside point estimates, though most public reporting does not.

The disciplinary coverage, while broader than many benchmarks, is still limited to physics, chemistry, and biology. Mathematics, computer science, engineering, and social sciences are not represented. A model could score poorly on GPQA while excelling at mathematical reasoning (as measured by MATH) or vice versa.

Question quality, despite the expert-creation protocol, is not perfectly uniform. Some questions have been criticized for being ambiguous, having debatable correct answers, or relying on conventions specific to one subfield. With only 198 questions, even a few problematic items can meaningfully affect scores.

Data contamination is less of a concern for GPQA than for older benchmarks, because the questions are more novel and less likely to appear verbatim in training data. However, the underlying scientific concepts and problem types may be well-represented in training corpora, meaning models may benefit from indirect exposure even if they have not seen the exact questions.

How Scores Have Changed Over Time

YearTop ModelScoreKey Insight
2023GPT-439.7%Barely above random (25%), showed difficulty was genuine
2024Claude 3 Opus50.4%First model to meaningfully exceed random baseline
2024GPT-4o53.6%Incremental gains, reasoning still limited
2024o1-preview73.3%Chain-of-thought reasoning yielded dramatic improvement
2025Claude Opus 474.9%Approaching in-domain expert level (81%)

The most important inflection point was the introduction of extended reasoning models in late 2024. The jump from GPT-4o’s 53.6% to o1-preview’s 73.3% — a 19.7-point gain — demonstrated that chain-of-thought and extended thinking architectures unlock scientific reasoning capabilities that standard generation models cannot match. Claude Opus 4’s 74.9% continues this trend, suggesting that further gains may depend more on reasoning architecture than on scale alone.

GPQA vs Other Benchmarks

BenchmarkDomainDifficultyQuestionsExpert Human Baseline
GPQA DiamondPhD-level scienceVery Hard19881% (in-domain)
MMLUBroad academicsMedium~15,900~89%
ARC-ChallengeGrade-school scienceEasy2,590~95%
MATHCompetition mathHard5,000~90% (varies)
ScienceQAK-12 scienceEasy21,208~90%

GPQA occupies the hardest end of the knowledge benchmark spectrum. While MMLU tests breadth across many subjects at moderate difficulty, GPQA tests depth in three scientific disciplines at expert level. A model can score 90% on MMLU through broad but shallow knowledge, while scoring 55% on GPQA — the two benchmarks measure fundamentally different capabilities.

GPQA and MATH are complementary: MATH tests formal mathematical reasoning and symbolic manipulation, while GPQA tests scientific reasoning that integrates conceptual understanding with quantitative analysis. Strong performance on both indicates a model with robust analytical reasoning across domains.

Practical Implications

For researchers considering AI as a scientific reasoning tool, GPQA scores provide the most relevant signal available. A model scoring above 70% can be expected to provide useful analysis of graduate-level scientific questions, though its reasoning should always be verified by domain experts. At 74.9%, Claude Opus 4’s performance suggests it could serve as a capable first-pass research assistant across physics, chemistry, and biology.

For general users, GPQA scores correlate with a model’s ability to handle complex, multi-step reasoning in any domain. Models that score well on GPQA tend to perform better on other tasks requiring careful logical deduction, even outside science. The reasoning capabilities tested by GPQA — multi-step inference, integration of multiple concepts, resistance to superficially plausible but incorrect answers — transfer to legal analysis, financial modeling, and other expert domains.

For AI safety and governance, GPQA Diamond scores serve as a meaningful capability threshold marker. Models exceeding 75% on GPQA are demonstrating scientific reasoning that rivals trained humans, which has implications for how these models should be governed, deployed, and monitored.

The practical limitation is that GPQA only covers three natural sciences. Users working in mathematics, engineering, or computer science should supplement GPQA scores with MATH, SWE-bench, or domain-specific evaluations.

Frequently Asked Questions

What does “Google-proof” mean in the context of GPQA?

Google-proof means the questions are designed so that someone without domain expertise cannot find the correct answer through web search or textbook lookup. The questions require integrating multiple specialized concepts and performing multi-step reasoning — skills that cannot be replicated by searching for keywords. When non-expert humans attempted to answer GPQA questions using unlimited internet access, they averaged only about 34% accuracy (barely above the 25% random baseline), confirming the design succeeded.

How does GPQA Diamond differ from regular GPQA?

The full GPQA dataset contains three tiers: Extended (546 questions), Main (448 questions), and Diamond (198 questions). Diamond is the hardest subset, containing only questions that met the strictest validation criteria — high in-domain expert accuracy, low out-of-domain expert accuracy, and clear consensus on the correct answer. Diamond scores are consistently lower than Main or Extended scores for the same model.

Why is the human expert baseline only 81%?

Even PhD experts make mistakes on questions in their own field, particularly when questions span subfield boundaries (e.g., a particle physicist answering a condensed matter question). The 81% in-domain accuracy reflects the genuine difficulty of these questions — they are not straightforward recall but require applying multiple concepts and performing calculations or reasoning chains where errors can compound.

Can models game GPQA through process of elimination?

To some extent. With four answer options, a model can sometimes eliminate obviously wrong answers to improve its odds. However, GPQA’s distractors are designed by domain experts to be plausible — each wrong answer represents a common misconception or a result you would get if you made a specific reasoning error. This makes process-of-elimination less effective than on benchmarks with lower-quality distractors.

Will GPQA become saturated like MMLU?

Not soon. The gap between the best model (74.9%) and expert human performance (81%) means there is still meaningful headroom. More importantly, the hardest GPQA questions — those requiring five or more reasoning steps or integration across subfields — remain challenging even for the best models. Saturation would likely require advances in reasoning architecture beyond current capabilities.