Glossary

Cortexa AI Glossary · How it answers

What is a small model, and can AI run on my phone?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Turn on airplane mode and type a message. Something still suggests the next word.


Airplane mode

Put your phone in airplane mode and start typing a message. It still suggests the next word, and it still fixes your typos. Nothing left the phone to make that happen. The suggestion came from a small artificial intelligence (AI) model that lives on the phone itself, doing its work right there in your hand.4

Small, by comparison

"Small" is relative. The largest language models run in data centers and can have hundreds of billions of parameters, the adjustable numbers that training sets inside a model. A small model has anywhere from a few million to a few billion. Picture a pocket dictionary next to a whole library. Apple, for instance, said in 2025 that the model built into its newer devices has about 3 billion. That's still a lot of numbers. But it's few enough to fit in a phone's memory and run on its chips.13

What it gives up

Shrinking a model has a price. A small model holds less general knowledge, so it's more likely to miss a fact or stumble on a hard, open-ended question. Ask one about an obscure piece of history and it may come up short. Ask it to tidy a sentence and it's quick. For focused jobs like that, or predicting your next word, or sorting notifications, it can be very good, especially when it has been trained or tuned for that one kind of task.1

No trip required

When a model runs on your device, your request doesn't have to travel to a data center and back. So it can work offline, on a plane or in a subway tunnel. No signal needed. And it can answer faster, because there's no wait for the network, however slow the connection is where you're standing. Many newer phones and laptops carry a part of their chip designed for exactly this kind of math, sometimes called a neural processing unit.15

What stays on the phone

Privacy is the other big reason. What a model handles on your device doesn't need to be sent to anyone's server. For your messages, your photos, and what you type, that's a real comfort. Your words stay with you. It doesn't cover everything an app does, though, because the same app can use both kinds of model.1

Less power per answer

Small models also cost less to run. Fewer numbers means less computing for each answer, which means less energy and, for the companies running them, smaller bills. Across millions of requests a day, that adds up. That's one reason many products use a small model for quick everyday jobs and save a larger one for harder requests.12

Which one am I using?

So how do you know which kind you're using? Often the screen won't tell you, because many apps mix both, and there's no badge that says where your request went. Apple, for example, describes a model on its devices alongside a larger one that runs on its own servers for harder requests. An app's settings or privacy page usually says what stays on your device and what gets sent away. Which app on your phone would you check first?3

Works cited

  1. IBM, "What are small language models?" (checked )
  2. IBM, "Are bigger language models always better?" (checked )
  3. Apple Machine Learning Research, "Apple Intelligence Foundation Language Models: Tech Report 2025." (checked )
  4. Apple Newsroom, "iOS 17 makes iPhone more personal and intuitive" (2023) (checked )
  5. IBM, "What is a neural processing unit (NPU)?" (checked )