What Is the Chatbot Arena?

In this article

  1. How it works, no tricks
  2. What MMLU and GPQA can't measure
  3. The cracks in the method
  4. The ranking today and what it says, and what it doesn't
  5. Using Arena as a filter before paying
  6. The political question
  7. To go deeper
  8. You might also like

Definitions · References · Elsewhere

Today we're talking about the Chatbot Arena, the only language-model evaluation that resembles reality. The thesis is rough. When the press says "the best AI model is X," the question back should always be "by what measure?", and it almost never gets asked. Why do we believe a ranking they tell us about without showing us the gauge? My take is biased because I've been using Arena as a sanity check for two years before paying for subscriptions, and it has saved me two of them. Build your own. Let's take it slowly: how it works, why it matters, what's fallen off along the way, and when you should distrust Arena too.

I admit that the first time I went onto the LMSYS site I thought it was an academic toy. Two text boxes, two anonymous answers, a button to vote. It didn't have the solemnity of the benchmarks the press reproduced with figures to two decimal places. It took me weeks to realise that the lack of solemnity was precisely the virtue. Benchmarks impress; Arena measures.

How it works, no tricks

The system was built by a group of UC Berkeley researchers inside the LMSYS initiative, now reconverted into an organisation called LMArena. The mechanics are simple. The user goes onto the site, types any question, and two anonymous models —identified as Model A and Model B— answer at once. The user picks which of the two seems better, or declares a tie. Only after voting are the models revealed. Their vote enters an Elo-type scoring system —the same one from chess, adapted by Arpad Elo in 1960 for the US Chess Federation— that recalculates the ranking continuously.

The paper that describes the method in detail is Chiang et al., Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference (arXiv:2403.04132), published in March 2024. In that paper the authors document that the system had accumulated, in its first year, over 240,000 votes on dozens of models, in more than a hundred languages. As of today the figure runs to several million. That's magnitude enough for statistical noise to drop to acceptable levels, even though the voter sample still carries the biases we'll see below.

The important thing about the mechanics isn't the Elo. It's the double blind. The voter doesn't know which models they're comparing. The models don't know who votes for them. That blindness is what kills —as far as possible— brand bias. When a journalist writes that "the new Claude impresses," they're reporting on the effect of the name as much as on the model. In Arena the name doesn't exist until the vote is already cast.

What MMLU and GPQA can't measure

Technical benchmarks measure concrete things, and for that very reason they fall short. MMLUMassive Multitask Language Understanding, proposed by Hendrycks et al. in 2020 (arXiv:2009.03300)— measures the ability to answer exam-style questions across 57 disciplines, from history to medicine. GPQA —Rein et al., arXiv:2311.12022, November 2023— measures PhD-level reasoning in physics, chemistry and biology; the questions were designed so that experts in the field would get them right and laypeople with internet access would get them wrong. HumanEval —Chen et al., arXiv:2107.03374, 2021— measures the ability to write programming functions that pass unit tests.

Each one measures something, yes. But none measures what the user does when they pay for the subscription. The user doesn't examine the model on doctoral chemistry. They ask it to rewrite an email. They ask it to explain a concept. They ask it to summarise a PDF and then answer questions about what was summarised. They ask it to write a short Python script. They ask it to help draft a complaint letter to the bank. For all that, MMLU has nothing to say.

Chatbot Arena does. The distribution of question types in Arena reproduces, with its nuances, what people actually ask of a model. And the final metric —which answer does a human prefer seeing the two blind?— captures the integration of many dimensions at once: correctness, tone, suitable length, clarity, absence of circumlocution and of needless warnings. It's not an analytic metric. It's aggregative, like almost every decision that matters.

Stanford HELM —Holistic Evaluation of Language Models, a Stanford CRFM initiative— does attempt a broader evaluation than MMLU, with several scenarios and multiple metrics per model. But HELM still evaluates against defined tasks, not against real human preference. Which makes it useful, but not a substitute.

The cracks in the method

Here I have to slow down, because I sold Arena as a cure and I haven't yet mentioned the side effects.

The first crack is demographic. Arena's voters are AI enthusiasts who have voluntarily come onto a specialised site. The sample is overrepresented in English speakers, in men, in technical profiles, in the US and European time zones. LMSYS's own estimates acknowledge this bias. If I'm a lawyer in Cádiz who needs a model that writes well in legal Spanish, Arena guides me less than it appears to.

The second crack is style. The LMSYS team published in August 2024 an analysis of their own —Does Style Matter? Disentangling Style and Substance in Chatbot Arena— in which they acknowledged that voters systematically reward longer answers, with markdown headers and lists, over answers in plain prose, even when the prose contains the same information. When the researchers statistically controlled for length and format —a Bradley-Terry regression with style variables— the ranking moved. Models like Claude 3.5 Sonnet and Llama-3.1-405B rose. GPT-4o-mini and Grok-2-mini fell. The data doesn't invalidate Arena. It invalidates the simplistic reading of the top spot.

The third crack is access. In April 2025, a group of researchers from Cohere Labs, AI2, Princeton, Stanford, Waterloo and the University of Washington published The Leaderboard Illusion (Singh et al., arXiv:2504.20879). They analysed around two million battles and 243 models from 42 providers, over a sixteen-month window (January 2024 to April 2025). Their conclusions were uncomfortable. Google and OpenAI had received, respectively, 19.2% and 20.4% of the total votes. The set of 83 open-weight models —from all providers combined— accumulated only 29.7%. On top of that, the big private labs could test variants in private before publishing the one that would enter the definitive ranking. The authors identified 27 private variants Meta had tested in Arena in the months before the Llama-4 launch. With enough access, relative gains in position could reach 112%.

That last one worries me most. It doesn't invalidate the blind-voting mechanics, which remain valid in themselves. But it degrades the competitive reading of the leaderboard, because the models don't reach the ring on equal terms. Some did twenty-seven warm-up runs. Others, none.

I was going to say that even so Arena remains the least bad of the rankings. But I'll correct myself, because "least bad" is a phrase that ages poorly. Let me put it this way: Arena measures something real (blind human preference), with an imperfect instrument (the biased sample and the unequal access). Any other ranking measures something less real with an instrument that looks cleaner because it hides its value judgements inside the benchmark's design.

The ranking today and what it says, and what it doesn't

If you go onto the LMArena site today, the top box of the leaderboard is usually occupied by three or four names that take turns every few weeks: Anthropic with its Claude line, Google with Gemini, OpenAI with GPT, and since 2025 also Chinese models like DeepSeek and the Qwen family. The difference between first and fifth in Elo score is usually, depending on the week, between 20 and 60 points. For perspective, a 100-point Elo difference in chess translates to roughly 64% wins of the higher over the lower. That is: when the press headlines that "model X crushes it," the statistical reality is usually "model X wins six times out of ten against the runner-up."

That proportion is what media coverage doesn't convey. And it's the one that most changes the consumer's rational decision. Paying double for an improvement from 50% (the tie) to a 60% preference probability can be reasonable if the task is critical. It isn't if the task is writing an email.

The ranking also branches by category. There's a general Arena, yes, but also separate categories: coding, long instructions, Spanish, Russian, creative tasks. A model can be first overall and fourth in coding. That granularity matters more than it seems and almost no one consults it.

Using Arena as a filter before paying

This is what I'm here for, and where the article stops being descriptive to turn practical.

When someone —a marketing firm, a popular explainer, a comparison site— says model X is better than Y, the first thing is to go onto LMArena.ai and look at the Elo difference between the two. There are three useful mental thresholds.

If the difference is under 30 Elo points, the difference is marginal and very probably sits inside the system's statistical noise. Either of the two models will be indistinguishable for 90% of tasks. The decision should be made on price, on data privacy, or on integration with the workflow. Not on quality.

If the difference is 30 to 100 points, there's a real but modest aggregate preference. It's worth looking at the specific category for the concrete task. If I write in Spanish and the leading model in Spanish is the runner-up of the general ranking, I should pick the runner-up, not the first.

If the difference exceeds 100 Elo points and holds for several weeks, there's a substantive reason behind it. It's worth reading what the winning model's technical note says, and whether there's a change in architecture or training method.

And if someone sells me a difference above 200 Elo points against the first serious competitor, the default suspicion should be that they're selling smoke, or that the comparison is made against a model two generations old.

The political question

This is personal opinion, but I see it backed by the logic of the sector's incentives. Who measures and how matters. It matters because the rankings —Arena included— govern companies' purchasing decisions and, by extension, the flows of capital toward the labs. The lab that rises in the ranking raises a round. The one that falls trims headcount. That immediate translation turns any evaluation criterion into a political instrument.

That's why, when a benchmark is owned by a company competing in the same market, there's a structural conflict of interest. And that's also why an academic initiative like LMSYS —now LMArena, registered as a non-profit organisation in California in 2024— has a value that's hard to replace, even if its methods are open to criticism. What The Leaderboard Illusion criticises isn't bad faith on the part of the LMSYS team, but the asymmetry of access the big labs exploit within a system designed with good intentions.

The reasonable response isn't to distrust Arena. It's to distrust Arena too, adding it to the inventory of instruments worth reading with your guard up. Along with HELM. Along with the technical benchmarks. Along with directly testing the model on the task you have in hand, which remains the most reliable evaluation there is.

What isn't reasonable is what the Spanish press, and most of the international press, does: copy the top spot of the ranking as if it were dogma, without looking at the Elo difference, without looking at the category, without looking at the sample bias, and without mentioning that the runner-up, in many cases, is a stone's throw away and costs half.

As I write this, the latest available analysis —The Leaderboard Illusion, Singh et al., arXiv:2504.20879, April 2025— documents that with preferential access to the system a provider could achieve relative position gains of up to 112%. That figure, more than any rhetorical conclusion, sums up the size of the crack and the urgency of always looking behind the leaderboard.

Definitions

Elo system: a method of relative scoring proposed by Arpad Elo in 1960 to rank chess players of the US Chess Federation. It calculates each competitor's score from the result of their direct matchups and the opponent's score; a 100-point difference predicts roughly a 64% win rate for the higher player.

Benchmark: a standardised test used to compare a model's performance on a specific task. MMLU, GPQA and HumanEval are benchmarks. Arena isn't a benchmark in the strict sense, but a system of evaluation by aggregated human preference.

Open-weight model: a language model whose trained parameters are published and anyone can download, run and, in many cases, modify. Distinct from open-source model in the strict sense, which also requires publishing training data and code.

Bradley-Terry: a classic statistical model for estimating the relative strength of competitors from pairwise comparisons. LMSYS uses it to control for style variables (length, format) in its ranking-robustness analyses.

References

Wei-Lin Chiang et al., Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference (arXiv:2403.04132, March 2024). The system's founding paper.

Shivalika Singh et al., The Leaderboard Illusion (arXiv:2504.20879, April 2025). An empirical critique of how the ranking works; documents the figures on unequal access and the Meta case before Llama-4.

LMSYS Org, Does Style Matter? Disentangling Style and Substance in Chatbot Arena (blog, August 2024). The team's own internal analysis of style bias and the Bradley-Terry correction.

Dan Hendrycks et al., Measuring Massive Multitask Language Understanding (arXiv:2009.03300, 2020). The original MMLU paper.

David Rein et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark (arXiv:2311.12022, 2023). The original GPQA paper.

Mark Chen et al., Evaluating Large Language Models Trained on Code (arXiv:2107.03374, 2021). The original HumanEval paper.

Percy Liang et al., Holistic Evaluation of Language Models (HELM, Stanford CRFM, 2022 onward). A multidimensional evaluation initiative as an alternative to Arena.

To go deeper

Arpad E. Elo, The Rating of Chessplayers, Past and Present (Arco Publishing, 1978). The mathematical foundation of the scoring system Arena adopts.

Brian Christian, The Alignment Problem (Norton, 2020). Context on why evaluating AI models is an open problem not solved by technical benchmarks.

Stuart Russell & Peter Norvig, Artificial Intelligence: A Modern Approach (Pearson, 4th ed., 2020). The chapter on evaluating and validating learning systems, useful for understanding why human preference is hard to operationalise.

Simon Willison, Understanding the recent criticism of the Chatbot Arena (simonwillison.net, 30 April 2025). A nimble, honest reading of the Cohere Labs paper and its implications.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment