Glossary

Cortexa AI Glossary · How it answers

What is a context window?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Why a long chat forgets what you said at the start, and what to do about it.


The instruction that slipped

You start a long chat by asking for short answers. For a while, that's what you get. Forty messages later, the answers have grown long again. The chatbot didn't decide to ignore you. Your instruction slipped out of what it could see, or got buried under everything that came after. The reason has a name: the context window.

What fits in the window

A context window is how much text an artificial intelligence (AI) model can work with at one time. Everything shares that space: your messages, its earlier replies, any document you paste or file you upload, instructions the app adds behind the scenes, and the reply it's writing now. Anthropic's documentation calls it the model's working memory. It's separate from what the model learned in training. Think of it as everything in front of the model at this moment.12

Counted in tokens

The window is measured in tokens, the small chunks of text a model reads and writes. In English, a token works out to about three quarters of a word. So a window's size is a count of chunks, and a long report you paste in can use up thousands of them before you've typed a word of your own. Pictures and documents you upload count against the same window.5

Sizes vary, and they grow

Windows differ a lot from model to model, and they've grown fast. As one example, as of October 2026, Anthropic lists a window of one million tokens for its newest Claude models. It estimates that at roughly 555,000 English words. Its smaller Haiku model holds 200,000. Other companies publish their own figures, and those change often. Treat any number you hear as a snapshot.3

When the window fills

So what happens in a very long chat? Something has to give. Depending on the app, the oldest messages may be dropped, or squeezed into a short summary, so the conversation can carry on. You may not see it happen; the chat just keeps going. Anthropic notes that chat apps can manage the window on a rolling "first in, first out" basis. Either way, the earliest details are usually the first to go, and that's what happened to your request for short answers.12

Lost in the middle

Even inside the window, not every part gets equal attention. In 2023, Nelson Liu and colleagues at Stanford tested models on long inputs. The models did best when the key fact sat near the start or the end, and noticeably worse when it was buried in the middle. Anthropic's documentation makes a related point: as the token count grows, accuracy and recall can slip. A bigger window holds more. That doesn't mean every line in it gets the same care.42

Three habits for long chats

Three habits help.

  • Put your most important point first, and if it really matters, repeat it near the end.
  • In a long chat, restate your key instruction now and then.
  • When a conversation drifts, start a fresh chat with a short summary of what matters.

Each one takes a few seconds, in any chatbot you use. Which of your long chats could use a fresh start today?

Works cited

  1. IBM, "What is a context window?" (checked )
  2. Anthropic, "Context windows" (Claude documentation) (checked )
  3. Anthropic, "Models overview" (Claude documentation) (checked )
  4. Liu et al., "Lost in the Middle: How Language Models Use Long Contexts" (Transactions of the Association for Computational Linguistics, 2024) (checked )
  5. OpenAI, "What are tokens and how to count them?" (checked )