Cortexa AI Glossary · How it learns
Is a bigger model always better?
From Cortexa Learn, by Cortexa Consulting. Last checked .
"Biggest ever" makes a good headline. Whether you need it is a different question.
A truck for the milk
Nobody drives a moving truck to pick up a carton of milk, even though the truck can carry far more. Choosing an artificial intelligence (AI) model works much the same way. Size sells. Launch announcements love the words "biggest ever." They tend to leave out the other half of the story: what the extra size costs, and whether your job needs it.
What bigger means
When people call a model bigger, they usually mean it has more parameters, the adjustable numbers inside it that training sets. Bigger models are usually trained on more data too, using more computing power. More parameters give a model more room to store patterns. Think of it as shelf space. That room is a large part of why bigger models can often handle harder and more varied tasks.1
The pattern researchers found
In 2020, a team at OpenAI measured what happens as language models grow. They trained many models of different sizes and found a smooth pattern. As size, data, and computing went up, the models got steadily better at predicting text, across a range of more than a million times in computing. That pattern, often called a scaling law, is one big reason companies kept building larger models. Bigger kept paying off.1
The catch
The same curve has a catch built in. Each further step of improvement takes a much bigger jump in size and computing than the step before. Going from good to slightly better can mean many times the cost. The easy gains come early. And the bill doesn't stop once training is finished and the model is out in the world. A bigger model is slower and pricier every time you ask it something, because more numbers are used in working out each answer.13
Size alone isn't the recipe
In 2022, researchers at DeepMind showed how much the training matters. Their model, Chinchilla, had 70 billion parameters, a quarter the size of their earlier model, Gopher. But with the same computing budget, it learned from about four times as much data. Chinchilla beat Gopher, and several even larger models from other labs, on a wide range of tests. The smaller model won. Their advice was to grow the data along with the model.2
Matching the job
It's the truck and the milk again. A lot of everyday work doesn't use what a giant model offers. Sorting messages into folders, pulling a date out of a form, or summarizing a short report are focused jobs. A smaller model can often do them well, faster, and for far less money. A hard math proof, or untangling a long and messy contract, may be worth the bigger model. Many teams use both: a small model for the routine work, and a large one saved for the hard part.3
A question to ask
So when the next "biggest ever" headline lands, the size does tell you something: more room, and usually a higher bill. Next time you pick a tool, try asking it the other way around. What's the smallest one that does this well?