Glossary

Cortexa AI Glossary · How it learns

How is a model like ChatGPT trained?

From Cortexa Learn, by Cortexa Consulting. Last checked .

The tidy paragraphs, the friendly opener. Every habit a chatbot has was trained in its own stage.


The polite reply

Ask a chatbot almost anything and the reply tends to arrive in tidy paragraphs, often with a friendly opener and an offer to help further. None of that manner came free. A chatbot such as ChatGPT is built in stages, and each stage leaves a habit behind. Knowing the stages explains a lot of what you see on the screen.

Stage one, pretraining

The first stage is called pretraining. The model reads an enormous library of text, much of it from the public web, plus books, code, and licensed material. It practices one task the whole time: predicting the next small piece of text. One task, repeated an enormous number of times. To get good at that, it has to soak up grammar, facts, styles of writing, and some patterns of reasoning, because all of them help predict what comes next.3

A text continuer

What comes out of stage one is powerful but odd. It continues text, and that's all it does. Type a question and it might write three more questions, as if finishing a list of quiz items. It has no particular wish to help, because nothing in its training rewarded helping. Clever, in a strange way. Not yet a helper. OpenAI's 2022 paper on its InstructGPT models was written about exactly this gap, between predicting text and following what a person asked.12

Stage two, learning from examples

The second stage is often called instruction tuning. People write examples of requests along with good answers: explain this, summarize that, draft a polite email. The model is trained on thousands of these pairs, and it learns the shape of a helpful reply. Picture a new employee handed a folder of model letters before their first day. For the InstructGPT work, OpenAI hired a team of about forty contractors to write that material and to judge the model's answers.2

Stage three, tuning toward preferences

The third stage uses people's judgment. The model writes several answers to the same request, and people rank them from best to worst. Those rankings train a second model that predicts which answers people would prefer, and the chatbot is then tuned to produce answers that score well. The method is called reinforcement learning from human feedback (RLHF), and it has a topic of its own, next in the order.12

A newer step

Newer reasoning models get extra practice. They're given problems whose answers can be checked, such as math and computer code, and rewarded when they reach the right one. Nobody has to judge a math answer by taste. It's right or it isn't. The Chinese lab DeepSeek described training this way in a paper in 2025. Labs mix, reorder, and change these stages, so treat all of this as the common outline. It's no single product's recipe.4

The marks each stage leaves

You can see each stage in the way a chatbot behaves.

  • What it knows comes from pretraining, which is also why its knowledge stops at a date.
  • How it talks, the tidy format and the helpful tone, comes from the examples in stage two.
  • Its habit of agreeing with you can come from stage three, because people tend to rate agreeable answers higher.

So when a chatbot sounds sure about something recent, ask it when its training data ends.5

Works cited

  1. OpenAI, "Aligning language models to follow instructions" (2022) (checked )
  2. Ouyang et al., "Training language models to follow instructions with human feedback" (2022) (checked )
  3. OpenAI, "How ChatGPT and our foundation models are developed." (checked )
  4. DeepSeek-AI, "DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning" (2025) (checked )
  5. Sharma et al. (Anthropic), "Towards understanding sycophancy in language models" (2023) (checked )
  6. IBM, "What is reinforcement learning from human feedback (RLHF)?" (checked )