Glossary

Cortexa AI Glossary · The basics

What is a foundation model?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Many different apps can share the same model underneath. Here's what that means for you.


Same factory, different boxes

Two store-brand cereals from different supermarkets can come out of the same factory, with different boxes on the front. Many artificial intelligence (AI) apps work in a similar way. A writing tool, a customer-service chatbot, and a homework helper can all be built on the same model underneath. That shared model has a name. It's called a foundation model.

One model, one job

For most of AI's history, each model was built for a single task. One was trained to translate. Another found faces in photos. Another sorted spam. If you wanted a new skill, you usually gathered new examples and trained a new model from scratch, which took time, money, and a great deal of carefully labeled data.1

Train broadly, then adapt

A foundation model turns that around. It's trained first on a huge, broad mix of data, such as text from books and the web, so it picks up general patterns of language and knowledge. That first stage, called pretraining, isn't aimed at any one job. Afterward, the same model can be adapted to many jobs. That's where the name comes from: it's the foundation other things get built on.12

Three ways to adapt it

Builders usually adapt a foundation model in one of three ways.

  • Fine-tuning: more training on a smaller set of examples for one job, like a company's past support chats.
  • Instructions: directions the model gets before every conversation, often called a system prompt.
  • Added knowledge: handing the model the right documents at the moment you ask, a method called retrieval-augmented generation.

None of these rebuilds the model from the ground up. Each one shapes the same base for a particular use.1

A young name

The term is younger than you might guess. Researchers at Stanford University coined it in August 2021, in a long report called "On the Opportunities and Risks of Foundation Models." Two of its authors later explained why: the older names didn't fit. "Large language model" was too narrow, since these models aren't only about language. And "pretrained model" made it sound as if the important part all happened after pretraining.234

Shared strengths, shared flaws

Because a fairly small number of foundation models sit under a great many products, what's true of the base tends to show up above it. A model that's good at summarizing lends that skill to every app built on it. A flaw travels the same way. The Stanford report warned that the defects of a foundation model are inherited by all the models adapted from it. So if one app gets a certain kind of question wrong, others built on the same base may get it wrong too.3

What sits under your apps

You often can't see which foundation model sits under an app, and the answer can change if the company switches. Some makers say so in a help page or on their website. When an app gives you a strange answer, it's fair to wonder whether the quirk belongs to the app or to the model underneath. Which apps on your phone might be sharing a foundation?

Works cited

  1. IBM, "What are foundation models?" (checked )
  2. Stanford HAI, "Brief definitions of key terms in AI." (checked )
  3. Bommasani et al. (Stanford), "On the Opportunities and Risks of Foundation Models" (2021) (checked )
  4. Stanford CRFM (Bommasani and Liang), "Reflections on Foundation Models" (2021) (checked )