Cortexa AI Glossary · How it learns
What are the three ways machines learn?
From Cortexa Learn, by Cortexa Consulting. Last checked .
Learning from answers, from patterns, or from a score. How machines learn, and why it matters which.
The kitchen
Think about how you learned to cook. Some of it came from a recipe with a photo of how the dish should turn out. Some came from sorting a crowded spice drawer until the things that go together sat together. And some came from burning the toast, again and again, until you stopped. Machines learn in much the same ways, and each way has a name.
Learning from answers
The first is the recipe photo. The machine gets examples with the right answer attached, and learns to match them. A spam filter is the classic case: thousands of emails, each one already marked spam or not spam by people, and the machine learns what separates the two. It's called supervised learning, because the answers act like a teacher checking the work. A lot of the artificial intelligence (AI) that sorts or labels things for you was trained this way.12
Finding patterns
The second is the spice drawer. Nobody gives the machine any answers. It gets a pile of data and looks for structure on its own: things that cluster together, and things that don't fit anywhere. A store might use it to find groups of customers who shop alike. A bank might use it to spot a payment that looks unlike all the others. That's unsupervised learning. The machine can find the groups, but it can't tell you what they mean, so people look at them and decide.12
Learning from a score
The third is the toast. The machine takes an action, gets a score that says how well it went, and tries again. Nobody shows it the right move. It only learns which choices led to better scores, over many, many tries. That's reinforcement learning. It suits problems with no single right answer to copy, only better and worse results, like steering a robot arm or planning a series of moves.12
Where each one fits
Each way suits a different kind of job.
- Supervised learning fits labeling and predicting, when you have examples with answers.
- Unsupervised learning fits grouping things and spotting the unusual, when you don't.
- Reinforcement learning fits step-by-step decisions, where the result only shows up after a series of moves.
Often the data decides. Labeled answers are slow and costly to collect, so people reach for the other two when there aren't enough.1
Real systems mix them
Big systems often use more than one. A chatbot starts by reading huge amounts of text and predicting the next word. The text supplies its own answers, since the real next word is right there, so nobody has to label anything. Researchers call that self-supervised learning. Many chatbots then get a second stage, where people rank the model's answers and it's trained toward the ones they preferred. That's reinforcement learning from human feedback (RLHF). In OpenAI's 2022 study of the method, people preferred the tuned model's answers to the original's about 85 percent of the time.34
Trained against what?
Knowing how a system learned tells you what it was told counted as right. A supervised system learned the labels people gave it, so sloppy labels teach sloppy habits. An unsupervised one found patterns, but people decided what they meant. A system trained on a score learned to chase that score, and it's only as good as the score. So the next time you hear that an AI was trained, a useful follow-up is: trained against what?