Glossary

Cortexa AI Glossary · How it learns

What is an AI benchmark?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Every model launch comes with a chart of test scores. What those tests are, and what they leave out.


The launch chart

When a new phone launches, there's usually a chart showing it's faster than last year's. New artificial intelligence (AI) models launch with the same kind of chart, full of names you've probably never heard of, each with a score beside it. Those names are benchmarks. Once you know what's inside one, the chart gets much easier to read.

A test with known answers

A benchmark is a standard test with known answers, used to compare models on one kind of task. It has a fixed set of questions, the correct answer to each, and a way of scoring. Some are multiple choice. Others ask a model to solve a math problem, fix a bug in some code, or answer a question that takes real expertise. The score is usually the share it got right.13

Why the same test helps

The point is a fair comparison. If every model takes the same test under the same rules, you can line up their scores, the way a standardized exam lets a college compare students from different schools. Researchers also use benchmarks to track progress over time. Run the same test each year, and you can see how far the best models have come.1

One task at a time

Each benchmark measures one kind of skill. A math test says little about how well a model writes a cover letter, and a coding test tells you nothing about whether it gets medical facts right. So a model can top one chart and sit in the middle of another. When you see a score, the first thing to find out is what the test asked.

Tests get used up

Benchmarks wear out. When the best models all score near the top, a test can no longer tell them apart, and researchers say it's saturated. So they write harder ones. A test called Humanity's Last Exam, released in January 2025, was built from expert questions meant to be hard for AI. Stanford's AI Index 2026 reports that top models gained about thirty percentage points on it in a single year. Tests meant to stay hard for years, it says, are now used up in months.24

When the answers leak

There's a quieter problem too. Models learn from enormous amounts of text from the web, and benchmark questions get posted and discussed online. If the questions and their answers end up in a model's training data, the model may remember answers instead of working them out. Researchers call that contamination, and it can make a score look better than the skill behind it.1

Who picks the chart

Remember who made the launch chart. A company shows the tests its new model did well on, which is natural, and it doesn't have to show the rest. It may also have run the tests under its own settings. So a launch chart is a starting point. A comparison run by someone with nothing to sell can tell you more.

The question to keep

The most useful question is a plain one. Is this test anything like my work? If you mostly ask a chatbot to summarize reports, a top score in competition math may not matter much to you. The test that fits best is your own: a handful of real tasks you already know the answers to. Topic 84 shows how to put one together.

Works cited

  1. IBM, "What are LLM benchmarks?" (checked )
  2. Stanford HAI, "The 2026 AI Index Report: Technical performance." (checked )
  3. Google for Developers, "Machine Learning Glossary: Metrics." (checked )
  4. Phan et al., "Humanity's Last Exam" (2025) (checked )