Cortexa AI Glossary · How it answers
What is a mixture of experts?
From Cortexa Learn, by Cortexa Consulting. Last checked .
How a model can be enormous and still answer quickly.
The front desk
At a big hospital, the front desk doesn't send you to every doctor in the building. It sends you to one or two who fit. Some large artificial intelligence (AI) models are built in a similar way. It helps explain a puzzle you may have noticed in their launch announcements. The design is called a mixture of experts.
Huge and quick at once
Model announcements often boast two things that seem to pull against each other: an enormous number of parameters, and quick replies. Parameters are the numbers a model learns during training. In a typical model, all of them do some work for every word it writes, so more parameters usually means more computing for each reply. How can a model be huge and still fast?
Many smaller parts
A mixture-of-experts model splits part of itself into many smaller sub-networks, called experts. They sit side by side inside the model's layers. Any one of them could process the text. But the model doesn't run them all at once, and that choice is where the savings come from.1
The router
Beside the experts sits a small piece called a router. For each token, which is a word or a piece of a word, the router picks the few experts that should handle it. The rest stay switched off for that token. So the model holds a great many parameters, and each word passes through only some of them. Researchers call this sparse activation.1
One model that published its numbers
One well-documented example is Mixtral, an openly released model from the French company Mistral AI. Each of its layers has eight experts, and the router picks two of them for every token. In all, the model holds about 47 billion parameters. Each token uses only about 13 billion. That's how it can hold a big model's worth of parameters while doing the work of a much smaller one for each word.2
Expert in what?
The name invites a picture of one expert for medicine, one for law, and one for cooking. Mixtral's authors looked for that and didn't find it. Texts from math papers, biology abstracts, and philosophy were routed to the experts in very similar ways. They called that surprising. The router did show patterns, but they followed the structure of the text, such as words in computer code, more than its subject. An expert here is a part that got good at some slice of the work. That slice may not be anything a person would name.2
Older than chatbots
The idea is older than today's chatbots. In 2017, a Google team led by Noam Shazeer showed it could work at a very large scale for language, with a layer of up to thousands of experts and a router picking a few for each input. Their paper was called "Outrageously Large Neural Networks." Some makers now say their models use this design, and others don't say. Next time you read a parameter count, look for a second number: how many are switched on for each token.3