Home · Benchmarks · Chatbot Arena Elo
BENCHMARK

Chatbot Arena Elo

Crowdsourced human preference rankings where users rate blind side-by-side model responses. The gold standard for measuring real-world conversational quality.

General UNIT Elo HIGHER IS BETTER MODELS 8
Leaderboard 8 models
# Model Lab Score Date
1 Gemini 2.5 Pro Google DeepMind
1407Elo
APR 2025
2 Claude Opus 4 Anthropic
1399Elo
MAY 2025
3 GPT-4.1 OpenAI
1389Elo
APR 2025
4 Grok-3 xAI
1383Elo
FEB 2025
5 Claude Sonnet 4 Anthropic
1372Elo
MAY 2025
6 GPT-4o OpenAI
1348Elo
MAY 2024
7 Llama 4 Maverick Meta
1340Elo
APR 2025
8 DeepSeek V3 DeepSeek
1335Elo
JAN 2025
About this benchmark

Chatbot Arena Elo is the most widely watched and influential ranking system in AI, measuring which language models real humans actually prefer when given blind, side-by-side comparisons. Operated by the LMSYS team at UC Berkeley — led by Lianmin Zheng, Wei-Lin Chiang, and Ying Sheng — the Arena works by presenting users with anonymous responses from two different models to the same prompt and asking them to vote for the better one. Elo ratings are computed from millions of these pairwise comparisons using the same rating system used in competitive chess, producing a single number that captures real-world human preference across a diverse range of tasks and users. No other AI evaluation comes close to the Arena’s scale, ecological validity, or influence on model development and adoption decisions.

What Chatbot Arena Elo Measures

Chatbot Arena measures human preference — which model’s response a real user considers better for their actual task. This is fundamentally different from automated benchmarks, which measure specific capabilities like accuracy on test questions or code correctness. The Arena captures the holistic quality judgment that determines whether a user finds an AI assistant genuinely useful: Does the response answer the question? Is it clear and well-organized? Is the tone appropriate? Does it demonstrate the right level of detail?

Users submit their own prompts — anything from creative writing requests to coding problems to philosophical questions — and receive responses from two anonymous models. They vote for Model A, Model B, or tie. The models’ identities are revealed only after the vote, preventing brand bias from influencing judgments. This blind evaluation protocol ensures that ratings reflect genuine response quality rather than preconceptions about specific models or labs.

Elo ratings are derived using the Bradley-Terry model, which converts pairwise win/loss records into a continuous scale where a 100-point difference corresponds to approximately a 64% win probability. Ratings are anchored to a starting baseline and updated as new comparisons accumulate. Category-specific leaderboards for coding, math, creative writing, instruction following, and other domains provide more granular analysis than the overall ranking.

The Arena has collected millions of votes from hundreds of thousands of unique users, making it the largest-scale human evaluation of AI models ever conducted. This scale provides statistical robustness that no other evaluation methodology can match.

Current Leaderboard Analysis

The current Elo rankings reveal a tightly competitive frontier with a clear hierarchy. Gemini 2.5 Pro leads at 1407, followed by Claude Opus 4 at 1399 and GPT-4.1 at 1389. The top two are separated by just 8 Elo points — a margin where a user in a blind test would be unable to consistently distinguish which model produced which response. An Elo gap of less than 20 points is generally considered imperceptible in practice.

Grok-3 (1383) and Claude Sonnet 4 (1372) occupy clear fourth and fifth positions, with a more significant gap to GPT-4o (1348), Llama 4 Maverick (1340), and DeepSeek V3 (1335) at the bottom of the tracked models. The 72-point spread from top (1407) to bottom (1335) is substantial in Elo terms — it means Gemini 2.5 Pro would be expected to win approximately 60% of head-to-head comparisons against DeepSeek V3.

The three-way race at the top — Gemini 2.5 Pro, Claude Opus 4, and GPT-4.1 — represents the most competitive period in the Arena’s history. A year ago, GPT-4o at 1348 held a comfortable lead. The fact that three labs are now clustered within 18 Elo points of each other confirms that the frontier is genuinely competitive, with no single lab holding a decisive advantage in overall human preference.

The leaderboard ordering differs from specialized benchmarks in revealing ways. Claude Opus 4 leads on SWE-bench and GPQA Diamond but trails Gemini 2.5 Pro on Arena Elo. GPT-4.1 leads MMLU but sits third on Arena. These discrepancies highlight that human preference integrates many factors — response quality, style, helpfulness, appropriate detail level — that no single specialized benchmark captures.

Methodology Deep Dive

The Arena operates on a simple protocol: a user types a prompt, receives two anonymous responses, and votes. Behind the scenes, the system employs several mechanisms to ensure rating quality and statistical validity.

Model selection is randomized but weighted to prioritize comparisons between models with similar Elo ratings, which produces the most statistically informative comparisons. The system also ensures that new models receive sufficient comparisons to stabilize their ratings quickly.

Elo computation follows the standard Bradley-Terry model with modifications for the AI context. Starting ratings are set at 1000, and ratings update after each comparison. The magnitude of each update depends on the expected outcome — an upset (lower-rated model winning) produces a larger rating change than a predicted outcome. Confidence intervals are published alongside point estimates, and the LMSYS team has found that ratings typically stabilize after approximately 10,000-20,000 comparisons.

The LMSYS team has introduced several refinements over time. Style-controlled ratings attempt to neutralize verbosity bias by adjusting for response length. Category-specific leaderboards (coding, math, creative writing, instruction following, hard prompts) provide domain-specific rankings. A “hard prompts” category specifically filters for challenging user queries that better separate frontier models.

Votes are filtered for quality — extremely rapid votes (suggesting the user did not read both responses), votes from accounts with suspicious patterns, and votes on trivially simple prompts are weighted down or excluded. The team has published analyses showing that these quality controls improve the correlation between Arena rankings and expert human judgment.

The Arena interface is available at lmarena.ai (formerly chat.lmsys.org) and requires no account to use, lowering the barrier to participation and ensuring a diverse user base.

Why This Benchmark Matters

Chatbot Arena matters more than any automated benchmark because it measures what ultimately determines model success: whether real humans prefer a model’s responses for real tasks. Automated benchmarks can measure specific skills — mathematical accuracy, code correctness, factual knowledge — but they cannot capture the holistic quality judgment that determines user satisfaction, retention, and willingness to pay for AI services.

The Arena has become the benchmark that labs openly compete on and celebrate in launch announcements. When a new model debuts, its Arena Elo ranking is among the first metrics cited by both the lab and the media. Enterprise buyers, developers choosing which API to integrate, and individual users selecting a subscription all reference Arena rankings as a primary decision factor.

The scale of the evaluation — millions of votes from diverse users on diverse tasks — provides a breadth of coverage that no designed benchmark can match. Users bring prompts that benchmark designers never imagined: niche domain questions, culturally specific requests, creative challenges, ambiguous instructions, multi-step workflows. The Arena captures model performance across this entire distribution of real-world use, not just the narrow slice that benchmark designers choose to test.

For the AI research community, Arena Elo provides a ground-truth signal against which automated benchmarks can be calibrated. When an automated benchmark produces rankings that diverge significantly from Arena Elo, it suggests the automated benchmark may be measuring something different from what users actually care about.

Known Limitations and Criticisms

Verbosity bias is the most documented concern. Research has consistently shown that users in blind comparisons tend to prefer longer, more detailed responses, even when the additional detail is not more informative or helpful. This creates a systematic advantage for models that produce longer outputs and may not reflect genuine quality differences. The LMSYS team introduced style-controlled ratings to mitigate this, but the raw Elo (which is the most-cited number) does not incorporate this correction.

User demographics skew toward English-speaking, technically oriented individuals — predominantly software developers, researchers, and tech enthusiasts. This means Arena ratings may not reflect the preferences of general consumers, non-English speakers, or users in specific professional domains (medicine, law, education). Models that are optimized for the Arena’s user base may not be the best choice for other populations.

The evaluation protocol tests single-turn responses to user prompts. While some users engage in multi-turn conversations before voting, the rating primarily reflects first-response quality. Models that excel at sustained multi-turn dialogue may be underrated relative to models that produce impressive single responses.

Gaming and manipulation are ongoing concerns. Labs that submit their models to the Arena may tune them to produce the kind of responses that win blind comparisons — longer, more comprehensive, more stylistically polished — at the expense of other qualities like conciseness, accuracy, or appropriate uncertainty. There have been allegations (though none proven) of coordinated voting campaigns to boost specific models.

Sample efficiency varies by model. New models need thousands of comparisons before their ratings stabilize, and during the stabilization period, early ratings may not be representative. The speed of the evaluation depends on user traffic to the Arena, meaning some models accumulate votes faster than others based on factors unrelated to their quality.

How Scores Have Changed Over Time

YearTop ModelElo ScoreKey Insight
2023GPT-4 (launch)~1250Established clear lead over all competitors
2023Claude 2~1150Significant gap to GPT-4, closest competitor
2024GPT-4 Turbo~1290Incremental improvement, still dominant
2024Claude 3.5 Sonnet~1270Gap to GPT-4 narrowed significantly
2024GPT-4o1348Strong lead, but competitors closing in
2025Gemini 2.5 Pro1407First non-OpenAI model to lead Arena

The most significant development in Arena history was the narrowing of the gap between labs. In 2023, OpenAI held a 100+ Elo point lead. By 2025, three labs are within 18 points of each other. Gemini 2.5 Pro’s ascent to the top position is notable as the first time a non-OpenAI model has held the overall Arena lead. This competitive convergence reflects the maturation of the field — the low-hanging fruit in model quality has been picked, and marginal improvements now require increasingly specialized optimization.

Chatbot Arena vs Other Benchmarks

EvaluationMethodScaleSpeedWhat It Captures
Chatbot ArenaHuman preference votesMillions of comparisonsWeeks to stabilizeHolistic user preference
MT-BenchLLM judge (GPT-4)80 questionsMinutesMulti-turn conversational quality
MMLUAutomated accuracy~15,900 questionsHoursFactual knowledge breadth
AlpacaEval 2.0LLM judge805 questionsHoursSingle-turn instruction following
HumanEvalAutomated code tests164 problemsHoursCode generation correctness

Chatbot Arena is unique in using real human judgments at massive scale. Every other major benchmark uses either automated scoring or LLM-as-judge approaches. This makes Arena the most ecologically valid evaluation — it measures what users actually prefer — but also the slowest and most expensive. The correlation between Arena Elo and MT-Bench (r approximately 0.9) suggests that automated evaluations capture most of the signal, but the remaining 10% often represents exactly the kind of qualitative differences (tone, helpfulness, appropriate detail) that matter most to users.

Arena occupies a different niche from capability-specific benchmarks. MMLU, HumanEval, and MATH measure specific skills. Arena measures the integrated quality of a model’s responses across all skills simultaneously, weighted by what real users actually ask about. This makes Arena the best single-number summary of overall model quality but the worst tool for understanding why a model ranks where it does.

Practical Implications

For individual users choosing a model for general-purpose use, Arena Elo is the single most informative ranking. The top-ranked model on Arena is, by definition, the one that the most diverse group of real users preferred for the most diverse set of real tasks. When the Elo gap is small (less than 20 points), users should instead base their choice on secondary factors: pricing, speed, availability, privacy policies, and performance on specific tasks they care about.

For enterprise buyers, Arena Elo serves as a strong default ranking for general-purpose assistant applications (customer support, internal productivity tools, content generation). However, enterprises with specialized needs should supplement Arena rankings with domain-specific evaluations. A company deploying a coding assistant should weight SWE-bench and Aider Polyglot scores more heavily than overall Arena Elo.

For AI researchers and developers, Arena Elo provides the ground truth against which other evaluations should be calibrated. If a new automated benchmark consistently disagrees with Arena rankings, the automated benchmark is more likely to be capturing an artifact than to have identified a genuine quality dimension that millions of human judges missed.

The category-specific Arena leaderboards are often more actionable than the overall ranking. The coding leaderboard, math leaderboard, and creative writing leaderboard may produce different orderings that better match specific use cases. Users should check these domain-specific rankings when their primary use case falls clearly into one category.

The practical interpretation of Elo differences: a gap of 50+ points is noticeable to most users in blind comparisons. A gap of 20-50 points is perceptible to attentive users. A gap of less than 20 points is effectively imperceptible and should not drive model selection decisions.

Frequently Asked Questions

How does the Elo rating system work?

Elo ratings work exactly like chess ratings. Every model starts at a baseline rating. When two models are compared and a user votes for one, both models’ ratings are updated. The winning model gains points and the losing model loses points. The magnitude of the update depends on the expected outcome — beating a higher-rated model produces a larger gain than beating a lower-rated one. Over thousands of comparisons, ratings converge to stable values that reflect each model’s true win rate against the field. A 100-point Elo difference corresponds to approximately a 64% expected win rate for the higher-rated model.

Is Chatbot Arena biased toward any particular model or lab?

The blind evaluation protocol eliminates brand bias — users do not know which model produced which response until after voting. However, several systematic biases exist: verbosity bias (users tend to prefer longer responses), demographic bias (the user base skews toward English-speaking tech workers), and recency bias (newer models may receive more attention and comparisons). The LMSYS team has introduced style-controlled ratings and other adjustments to mitigate these effects, but they affect the raw Elo ratings to some degree.

How many votes does a model need before its rating is reliable?

Ratings typically stabilize after 10,000-20,000 pairwise comparisons. During the initial period, ratings may fluctuate significantly. The LMSYS team publishes confidence intervals alongside point estimates, and models with fewer comparisons show wider confidence bands. Users should be cautious about interpreting ratings for newly added models that have not yet accumulated sufficient comparison data.

What is the difference between overall Elo and category-specific Elo?

The overall Elo rating reflects performance across all user prompts, weighted by what users naturally ask about. Category-specific ratings (coding, math, creative writing, instruction following, hard prompts) filter comparisons to specific task types. Category rankings often differ from the overall ranking — for example, a model might lead on coding but trail on creative writing. Category-specific ratings are more informative for users with specific use cases.

Why do Arena rankings sometimes disagree with automated benchmarks?

Automated benchmarks measure specific, objectively scorable capabilities (factual accuracy, code correctness, mathematical reasoning). Arena measures holistic human preference, which includes subjective factors like response tone, helpfulness, appropriate level of detail, and conversational quality. A model that scores highest on MMLU or HumanEval may not produce the most preferred responses in a blind comparison because technical accuracy is only one component of what users value. The disagreement between Arena and automated benchmarks is informative — it reveals which quality dimensions automated benchmarks fail to capture.