Cortexa AI Glossary · How it learns
What is training?
From Cortexa Learn, by Cortexa Consulting. Last checked .
"Trained on billions of examples" sounds grand. Underneath is a very patient loop.
The phrase in the ad
You've probably seen the line in an ad or a headline: "trained on billions of examples." It sounds grand, maybe a little mysterious. Underneath it is something patient and plain. Training is how an artificial intelligence (AI) model learns, and it works a lot like practice. Nobody types the answers in.
Where it starts
A model is, at heart, a very large set of numbers. Engineers call them weights or parameters, and a big model has billions of them. Before training begins, those numbers are set more or less at random. Ask an untrained model anything and you get nonsense back. Pure noise. Everything it later gets right comes from adjusting those numbers.17
Guess, check, nudge
Training repeats one small loop.
- Guess: the model is shown an example and makes a prediction, such as the next word in a sentence.
- Check: the guess is compared with the right answer, and the gap between them is measured. Engineers call that gap the loss.
- Nudge: each number is adjusted a tiny amount in the direction that would have made the gap smaller.
It's like a cook tasting a soup, adding a pinch of salt, and tasting again. The cook never writes down a rule for salt. The soup just gets better, one pinch at a time.12
Small steps, many times
Each nudge is tiny on purpose. Push the numbers too far and the model overshoots, the way too much salt ruins the pot. So the steps stay small, and the loop runs over and over, across more examples than a person could read in many lifetimes. No single nudge teaches the model anything you'd notice. Piled up across all those examples, they add up to a skill.23
What comes out
When training stops, what's left is the finished set of numbers. That's the model. It doesn't keep a library of its examples to look things up in, though it can memorize bits that showed up many times. What it learned is spread across the numbers as patterns. So it can write a sentence it never saw, and it usually can't point to the page where it learned something.
Why it costs so much
All that guessing and nudging takes a great deal of computing. The largest models train on thousands of specialized chips running for weeks or months. Stanford's AI Index estimated in 2024 that the computing alone for training OpenAI's biggest model of 2023 cost about seventy-eight million dollars. Prices like that help explain why only a handful of companies train the largest models from scratch, and why they don't retrain them every day.46
Set before you meet it
By the time you use a model, its training is finished, and a lot was settled there. Its skills come from the examples it practiced on, and so do its gaps. If a kind of question rarely came up, the model is likely weaker at it, and it can carry the slants of whoever wrote its examples. So when an ad says "trained on billions of examples," you can ask the next question. Billions of examples of what?
Works cited
- NVIDIA, "What is AI training?" (checked )
- IBM, "What is gradient descent?" (checked )
- IBM, "What is learning rate in machine learning?" (checked )
- Stanford HAI, "The 2024 AI Index Report." (checked )
- IBM, "What is machine learning?" (checked )
- Meta (Llama Team), "The Llama 3 herd of models" (2024) (checked )
- Google for Developers, "Neural networks: Training using backpropagation" (Machine Learning Crash Course) (checked )