Cortexa AI Glossary · How it learns
What is pretraining?
From Cortexa Learn, by Cortexa Consulting. Last checked .
The first and biggest stage of training, and where a chatbot's knowledge and its cutoff both come from.
Finish the line
Type "Twinkle, twinkle, little" into a chatbot and it can hand you "star" before you've finished asking. Give it the first half of a well-known proverb and it can usually finish that too, without searching the web. Nobody taught it nursery rhymes one at a time. It picked them up in the first and largest stage of its training. That stage is called pretraining.
One task, over and over
Pretraining gives an artificial intelligence (AI) model one job: read some text and predict the piece that comes next. The model guesses, compares its guess with the real text, and adjusts itself a tiny amount. Then it does it again. And again, across a huge collection of writing, such as books, articles, websites, and code. Nobody labels any of it, because the text is its own answer key.126
What comes out of it
Out of that one task comes a surprising amount. To predict the next word well, a model has to pick up grammar and spelling, the shape of a recipe or a cover letter, and a great many facts, because facts turn up in writing again and again. It holds none of that as a list of entries. It's all stored as patterns in the model's internal numbers. That's why it can finish your proverb. It's also why it can repeat a mistake that was common in what it read.1
The base model
The model that comes out of pretraining is called a base model, and it's an odd thing to talk to. It continues text. Ask it for the capital of France and it may carry on with a few more quiz questions, since on a web page a question is often followed by more questions. Anthropic's documentation for its Claude models says that pretrained models aren't naturally good at answering questions or following instructions.1
Why it's called "pre"
The "pre" means it comes first. After pretraining, the model goes through shorter stages that teach it to follow requests and to give the kind of answers people rate as helpful. Those stages shape how it behaves, but most of what it knows was already there. Pretraining builds the general base. The rest turns it into an assistant you can talk to.13
Where the cutoff comes from
Pretraining also explains a chatbot's knowledge cutoff. The collection of writing was gathered up to some date, and then the reading stopped. Anything that happened after that isn't in the model unless it's handed new information, for example through a web search. Model makers often publish the date. Anthropic's model pages list two, because the last months before the cutoff are thinly covered, so a model may know less about them than it seems to.4
Done rarely
Pretraining is the biggest and costliest stage. International Business Machines (IBM) describes training a model like this as taking thousands of specialized chips and weeks of computing, typically at a cost of millions of dollars. So companies don't pretrain a new model every time something happens in the world. They do it now and then, and build on the result for months. The next time a chatbot seems behind on the news, ask it directly: what's your knowledge cutoff?5
Works cited
- Claude docs, "Glossary" (pretraining, fine-tuning) (checked )
- IBM, "What is GPT (generative pre-trained transformer)?" (checked )
- IBM, "What are foundation models?" (checked )
- Claude docs, "Models overview" (knowledge cutoffs) (checked )
- IBM, "What is generative AI?" (checked )
- IBM, "What is LLM training?" (checked )