Cortexa AI Glossary · How it learns
What is inference, and how is it different from training?
From Cortexa Learn, by Cortexa Consulting. Last checked .
You tap Translate, and a second later it's in English. That second has a name.
The tap
A post shows up in a language you don't read. You tap "Translate," and a second later it's in English. In that second, an artificial intelligence (AI) model did its job. The name for that moment is inference.
Two stages
Every AI model goes through two very different stages. First comes training, the long and expensive stretch where it learns from examples by adjusting its internal numbers. That happens once for each version, usually long before you meet it. Then comes inference: the trained model putting what it learned to work on something new, over and over, for everyone who uses it. Think of the years a translator spends studying, and then the moment they glance at a sign and tell you what it says. Both are skill. Only the first one changes them.12
Inside that second
What happens in that second is mostly arithmetic. Your words are turned into numbers. Those numbers flow through the model's learned numbers, called weights, and out comes a prediction: the likeliest translation. A chatbot repeats that step for each small piece of its answer, which is why you can watch the reply appear a few words at a time. Nothing in the model changes along the way.1
Fixed while you use it
That last part surprises people. During inference, the model isn't learning. If you correct a chatbot, it can use your correction for the rest of that conversation, because your words are part of what it's reading. But the model itself is the same afterward. The next person to open the app meets the same version you did.
Every answer costs something
Training gets the headlines, but inference runs all day, for millions of people. Picture everyone who taps Translate in a single minute. Each answer takes computing power, and so a little electricity. Researchers at International Business Machines (IBM) have estimated that up to ninety percent of a model's working life is spent on inference. That's why companies work hard to make each answer cheaper and faster, and it's one reason some tools cap how much you can ask in a day.3
Where it runs
Most of the time, inference happens in a company's data center. Your request travels there, gets answered, and comes back. But smaller models can run right on your phone. Some captioning and translation features work this way, so they can respond with no signal and keep your words on the device. The biggest chatbots mostly still run in data centers, because their models are too large for a phone.14
When it behaves differently
Knowing this can save you some frustration. Correcting a model again and again doesn't change the model itself, so a better answer comes from what you put in the request: a clearer instruction, an example, the document it needs. And when a tool suddenly answers in a new way, the likely reason is that the company switched to a new version. So the next time an answer misses, ask yourself one thing. What could I add?