Glossary

Cortexa AI Glossary · How it answers

What is prompt caching?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Why the same long instructions can cost less the second time you send them.


The refill

When you refill a prescription, you don't bring the doctor's note again. The pharmacy already has it on file. Some artificial intelligence (AI) services do something similar with the long instructions at the start of many requests. It's called prompt caching. It's mostly a tool for people who build with AI, and it helps explain why the same model can cost less, or answer faster, in one app than in another.

The same long opening

Many apps send a model the same opening with every question. It might be pages of instructions on how to behave, called a system prompt. It might be a long document you want to ask about, like a product manual. Only the last few lines, your actual question, change. And in a long conversation, the whole earlier chat is sent again each time you reply.

Reading it all again

Without a cache, the model processes that whole opening from the start, each time. The work and the bill are counted in tokens, the small chunks of text a model reads, and a long opening can be most of them. So an app that asks a hundred questions about one long manual pays to have the manual read a hundred times. There's a wait, too: a long opening takes time to process before the first word of the answer appears.13

Keeping the work on file

Prompt caching keeps the work the model already did on that opening, for a short while. When the next request starts with the same text, the service picks up where that work left off and processes only what's new. Anthropic's documentation describes it as resuming from a prefix, meaning the opening part of a prompt. The reused part is usually billed at a lower rate, and the answer can start sooner. How long a cache lasts, and how much it saves, vary by company and change often.134

Exact match only

There's one strict rule. The saved part has to match exactly, word for word, from the very start. Change one character near the top, and everything after it has to be processed fresh. Even the current date and time, or details about the user, can break the match on every request if they sit near the top. So Anthropic's and OpenAI's guides both tell builders to put the parts that never change first, and the parts that do, like your question, at the end.13

Short-lived and separate

A cache is short-term storage. It's kept for minutes, sometimes up to a day, and then it's cleared. Anthropic and OpenAI both say caches aren't shared between organizations, so another company's app can't reuse yours. And reusing a cache teaches the model nothing. It only skips repeated work. Whether a company may use what you type to train future models is a separate question, answered by its own policies.123

Cached tokens

If you ever see "cached tokens" on a bill or a usage page, now you know what it means: text the service didn't have to process from scratch. If you build with AI, which part of your prompt stays the same every time?

Works cited

  1. Claude docs (Anthropic), "Prompt caching." (checked )
  2. OpenAI, "Prompt caching in the API" (2024) (checked )
  3. OpenAI, "Prompt caching" (API documentation, checked 2026-10-05) (checked )
  4. OpenAI, "Introducing GPT-5.1 for developers" (2025), extended prompt caching (checked )