Insights

Cortexa Ground Truth

Explainer5 min readBy Joe Coffman
0:00 / 5:12
1×

Tip: once it’s playing, tap any line to jump the audio there.

Prompt engineering was step one. Now engineer the context.

Prompt engineering was about wording the question. Context engineering is about everything the model can see when it answers.


“We wrote the perfect prompt, so why does the Artificial Intelligence (AI) still get it wrong?” Every team building with a Large Language Model (LLM) runs into this. The prompt is a sliver of what the model reads before it answers. The rest is context, and choosing it well has a name now: context engineering.

So what is context engineering?

Prompt engineering1 was the first version of this skill: write a clear request, get better work back. Context engineering is the wider job. It covers everything you put in front of the model before it answers: the instructions, the examples, the documents you retrieve, the tools you expose, the running history of the chat. Anthropic calls it the natural progression of prompt engineering2, and defines it as the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference2.

All of it has to fit inside the context window: the model’s short-term memory for a single task, everything it can hold in view at once, measured in tokens3. Andrej Karpathy described the goal as the delicate art and science of filling the context window with just the right information for the next step4. The window is finite, so the real question is which information earns a place in it.

The same desk both times, and the same size. What changes is which few things earned a place on it.

Is more context always better?

The tempting move is to pour everything in: the whole knowledge base, every past message, all the tools at once. Two things make that backfire. A model has a limited attention budget, so Anthropic warns that context must be treated as a finite resource with diminishing marginal returns2. Past a point, more text buys worse answers. Models also read the middle less carefully than the ends. Studying long-context models, researchers found that performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models5. Bury the one fact that matters in the middle of a huge prompt and the model can walk right past it.

Four levers of a well-fed model

What's Real?

The “so what” — for anyone serving tech clients

  • Debug the context before the prompt. When an AI feature misfires, look at what the model could see: stale documents, missing history, ten tools where two would do. That is usually a faster fix than re-wording the request.
  • Ask vendors how they manage context. “We use AI” says nothing. A partner who can explain how they retrieve6 the right facts, keep the history tight, and hand the model a clean set of tools has thought about reliability. One who only talks up the model has not.
  • Budget for the plumbing. Assembling good context is real engineering, and it is where a promising demo becomes a dependable feature. It is the same gap the demo-to-production brief7 named: a sharp sample is a long way from a shipped product. Building the retrieval, the memory, and the guardrails is the work that ships8.

Next Reads