Cortexa AI Glossary · Asking well
What are evals?
From Cortexa Learn, by Cortexa Consulting. Last checked .
Your work app updated overnight. Better or worse? Evals are how teams find out.
The update nobody announced
The app you use at work updated overnight, and this morning its summaries feel a little different. Shorter, maybe. Better? Worse? You can't really say, because all you have is a feeling from two or three examples. Teams that build with artificial intelligence (AI) run into this all the time. The tool they use to turn that hunch into an answer is called an eval.
Your test, from your work
Eval is short for evaluation. It's a set of real tasks, each paired with a clear idea of what a good answer looks like, that you run through an AI tool to see how it does. You may have heard of benchmarks, the standard public tests companies use to compare models. A benchmark is someone else's exam. An eval is yours, written from the work you need done.2
Built from real questions
The best evals start with the questions people really ask. Say your team uses AI to summarize customer emails. Then the eval holds real customer emails, with names and private details taken out: the long rambling ones, the angry ones, the ones that mention two problems at once. A tool can look great on tidy demo examples and stumble on messy ones. So the messy ones belong in the test.4
Decide what right looks like first
Before you read a single answer, decide what a good one is. Maybe a good summary names the customer's problem, says what they're asking for, and stays under a hundred words. Write that down for each task. If you look at the answers first, it's easy to talk yourself into liking whatever came back. A standard set in advance keeps the test fair.4
Run it again
An eval pays off the second time you run it. You run it again whenever something changes: a new model, a reworded instruction, a different setting. Same tasks, same standard, so the results line up side by side. If the new version passes forty of your fifty tasks where the old one passed forty-five, you know something slipped. And you know exactly which five tasks to look at, instead of guessing from the one odd answer you noticed over coffee.23
When the tool changes under you
Sometimes nothing changes on your side. The company behind a tool can update its model, and the answers shift. The questions people send can drift over time too. Running the same eval on a schedule can catch that early. The National Institute of Standards and Technology (NIST), part of the United States government, treats this kind of checking as part of managing the risks of generative AI: testing a system before release, and keeping watch on how it behaves once people are using it.1
A first handful
You don't need special software to start. A handful of real tasks, a note of what good looks like for each, and the answers your current tool gives today, kept in one document, is already the start of an eval. Then, when the next update arrives, you have something to compare it against besides a feeling. Which task from your own week would you put in first?