Glossary

Cortexa AI Glossary · How it learns

What is overfitting?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Great on the examples it learned from, weak on yours. Why that happens.


The streets around home

Picture someone who learned to drive only on the streets around their house. They know every bump, every stop sign, the exact moment to brake before the corner. Then a detour sends them across town, and suddenly every turn feels new. They learned one neighborhood very well. They never learned driving in general. A machine learning model can make the same mistake, and it has a name: overfitting.

What training is for

A model learns from examples, called its training data. But the examples are a means to an end. Nobody needs a model to sort emails it has already seen, because those answers are known. The point is to handle new ones: next week's emails, a photo it has never met. Google's machine learning glossary calls that ability generalization, making correct predictions on new data the model hasn't seen before.1

Learning the quirks

Overfitting is when a model matches its training examples so closely that it fails on new ones. Google's glossary puts it bluntly: the model memorizes the training set. Instead of the broad pattern, it picks up the quirks of its own examples, like odd lighting that ran through one batch of photos, or a phrase that turned up in one month's spam. Those quirks don't carry over, so on fresh data it stumbles.1

The telltale gap

Overfitting leaves a clear sign, and it's a gap. The model scores very well on the examples it trained on and noticeably worse on examples it hasn't seen. A small gap is normal, because no model is perfect on things it has never seen. A big one means the model has learned its examples instead of the pattern behind them.2

Holding some back

The usual way to catch it is to hold some examples back. Google's crash course suggests splitting the data three ways before training starts.

  • A training set, which the model learns from.
  • A validation set, which checks its progress while it trains.
  • A test set, kept back for one final check at the end.

If the score on the validation set starts getting worse while the training score keeps improving, the model has started memorizing.23

How it's prevented

The fixes are mostly common sense. More examples, and more varied ones, leave fewer quirks to latch onto; our driver would do better after practicing all over town. Teams also stop training at the point where the validation score stops improving, which Google calls early stopping. And there are ways to keep a model from getting too complicated for the data it has. Those go by the name regularization.2

Where you'd notice it

You may never see the word on a product page, but you can see what it does. A tool can look brilliant in a polished demo, on examples much like the ones it was built and tuned with, then struggle with your own messy files. So when someone shows you an impressive result, ask whether it was tested on data the model had never seen. Then try it on a few examples of your own.

Works cited

  1. Google for Developers, "Machine Learning Glossary: ML Fundamentals." (checked )
  2. Google for Developers, Machine Learning Crash Course, "Overfitting." (checked )
  3. Google for Developers, Machine Learning Crash Course, "Datasets: Dividing the original dataset." (checked )