Tip: once it’s playing, tap any line to jump the audio there.
Read an AI benchmark like a test score.
A benchmark score says how a model did on one fixed test, under one set of choices. It doesn't say how it will do on your client's job.
A client forwards a headline: this Artificial Intelligence (AI) model just hit number one on the leaderboard, so can we switch to it? The honest reply starts with a question of our own. Number one on which test, and does that test look anything like the work we do for them? A benchmark is a school exam for a model: one fixed set of questions, one scoring rule, one number at the end. The number is real. What it measures is narrower than the headline makes it sound.
What does a benchmark measure?
A benchmark is a fixed set of questions plus a rule for scoring the answers. Run a model over the questions, tally the result, and you get a number you can line up against other models. The top number on a given test gets called state of the art, and that phrase does a lot of quiet work. A strong score on a popular benchmark is widely understood as indicative of progress1 toward general ability, and Raji and colleagues spend a whole paper on the construct validity issues1 in reading it that way (preprint). Stanford's HELM project makes the same case by design: it reports a grid of metrics beyond accuracy2 across many scenarios rather than one flattering headline, in the name of reproducible and transparent evaluation2.
Why do leaderboards mislead?
The score is honest about one thing: how the model did on that test, that day, under those settings. Every step from there to a real decision is a place the number can drift away from the truth you care about. Four quick checks catch most of the drift.
Same test?
does the benchmark resemble the job you'd hire the model for
Studied?
the test questions can leak into training data and inflate the score
Real gap?
scores this close at the top are often inside the margin of error
Your job?
the only score that settles it is the one from your own task
A top-of-the-board number clears none of these on its own; you do.
What's Real?
- Good benchmarks earn their keep, run right. Stanford's HELM reports metrics beyond accuracy2 across many scenarios, so bias, cost, and reliability sit next to the headline number instead of hiding behind it. A transparent test, read with its caveats, tells you something worth knowing.
- The test can leak into the training data. When benchmark questions end up in what a model learned from, it can score high by memory rather than skill. A survey of the problem defines it as a model inadvertently incorporating evaluation benchmark information from its training data, leading to inaccurate or unreliable performance3 (preprint). It is the open-book exam nobody admits to sitting.
- The gaps at the top are often smaller than the noise. Evaluations are experiments, and experiments need error bars. Anthropic's work on the statistics of evals treats each test as questions drawn from an unseen super-population4 and shows how to report results in a way that minimizes statistical noise4 (preprint). Two models a point apart on a single run may not be reliably different at all.
- Hard tests get easy fast. Stanford's AI Index notes that benchmarks built in 2023 to stretch the best systems (MMMU, GPQA, SWE-bench) saw scores climb by 18.8, 48.9, and 67.3 percentage points5 within a year. A benchmark everyone aces has stopped telling models apart, which is why the goalposts keep moving.
- A popularity board measures preference, not correctness. Chatbot Arena ranks models based on human preferences6 through a pairwise comparison approach6 (preprint), which captures what readers enjoy, not whether the answer is right. And even that can be gamed: a 2025 analysis found systematic issues7 and warned of overfitting to Arena-specific dynamics rather than general model quality7 (preprint).
- One number is not a measure of general ability. Raji and colleagues argue that treating a favorite benchmark as a functionally “general” broad measure of progress1 runs into construct validity issues1 (preprint): being good at the test is a different thing from being good at the world.
The “so what” — for anyone serving tech clients
- When a client says “use the best model,” answer with a question: best at what, measured how? Then settle it on their ground. Before you quote a model, run a small eval8 on the client's real task, with their inputs and your own definition of a good answer. That result decides it, not the board.
- Don't rebuild the stack every time the ranking reshuffles. Benchmarks saturate and the order at the top churns; chasing it carries a real cost in rework and risk. Switch models when your own test says the new one is better on the client's job, not when a headline says it topped a chart. Steady beats twitchy when a client's name and your reputation are on the output.
Sources
- Raji et al. — AI and the Everything in the Whole Wide World Benchmark (preprint)
- Stanford CRFM — HELM (Holistic Evaluation of Language Models)
- Xu et al. — Benchmark Data Contamination of Large Language Models: A Survey (preprint)
- Miller (Anthropic) — Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (preprint)
- Stanford HAI — 2025 AI Index Report (Technical Performance)
- Chiang et al. — Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference (preprint)
- Singh et al. — The Leaderboard Illusion (preprint)
- Cortexa Ground Truth Nº 8 — “Evals, in plain English”
- Cortexa Ground Truth Nº 7 — “A great demo isn’t a shipped product”
Next Reads
- Explainer
What “agentic AI” actually means
Agentic AI isn't smarter answers — it's software that takes actions toward a goal. Here's what's real, and why it still needs a human in the loop.
- Explainer
AI doesn’t lie — it guesses
Why Artificial Intelligence (AI) “hallucinates,” in plain English: it isn’t lying — it predicts likely words, and a confident guess can still be wrong.





