Glossary

Cortexa AI Glossary · How it learns

What is distillation?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Why a chatbot's "mini" version can be quick and still pretty capable.


The mini version

Open the model menu in a chatbot and you'll often find two sizes: the main model, and a smaller one labeled something like "mini" or "fast." The small one answers sooner and usually costs less to run. And it may have learned a good part of what it knows from the big one, the way an apprentice picks up a trade by watching a master at work. In artificial intelligence (AI), that method is called distillation.

Teacher and student

Distillation trains a small model, called the student, to copy a large one, called the teacher. The teacher is already trained and good at its job. The student works through the same kinds of questions and is graded on how closely its answers match the teacher's. The aim is a model that keeps much of the teacher's skill at a fraction of its size.1

More than the right answer

The clever part is what the student copies. A teacher model gives more than an answer. Behind every answer sits a spread of guesses, each with a likelihood. Show a picture-sorting model a photo of a car, and it might be almost certain it's a car, faintly tempted by "truck," and almost never by "carrot." Those faint second guesses tell the student which things look alike. A plain right-or-wrong label carries none of that, so the student learns more from each example.13

Why bother

Big models are slow and expensive to run, and most are far too large for a phone or a laptop. A distilled student can answer faster, cost less per question, and need less computing power. That's why distillation is a common way to make the quicker, cheaper versions you see offered next to a company's main model. Some students are small enough to run on your own device.1

A real example

In January 2025, the Chinese company DeepSeek released a reasoning model, one built to work through a problem step by step. Alongside it came six much smaller models. To make them, DeepSeek had its big model write about 800,000 worked solutions, then trained existing open models on those solutions, from the Qwen family made by Alibaba and the Llama family made by Meta. The students picked up much of the teacher's step-by-step style.24

Where it gets argued about

Distillation also turns up in the news, usually in arguments about copying. Anyone who can send a model questions can collect its answers and train a student on them. Some companies' terms of use forbid using their models' output to build competing models. So distilling a rival's model without permission can break an agreement, and when that's alleged, it makes headlines. The facts in those stories often take a while to settle.5

What's lost

A student is rarely as good as its teacher at everything. Squeezing skill into a smaller model usually costs something, and often what goes is breadth. The student may keep up with the teacher on the kinds of tasks it practiced and fall behind on the rest. So next time a chatbot offers a mini version, try both on a question you care about, and see whether the quick one is good enough for you.1

Works cited

  1. IBM, "What is knowledge distillation?" (checked )
  2. IBM, "DeepSeek's reasoning AI shows power of small models, efficiently trained." (checked )
  3. Hinton, Vinyals and Dean, "Distilling the Knowledge in a Neural Network" (2015) (checked )
  4. DeepSeek, "DeepSeek-R1" (model release and distilled models, GitHub) (checked )
  5. OpenAI, "Terms of Use" (what you cannot do: using output to develop competing models) (checked )