Glossary

Cortexa AI Glossary · Asking well

What is RAG?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Those small numbered links under a chatbot's answer are a clue.


The little numbers

You open the help chat on a company's website and ask how to reset your router. The answer comes back with small numbered links underneath, each pointing to one of that company's own help pages. Those little links are a clue. The tool didn't answer from memory alone. It looked something up first, then wrote its answer from what it found. The name for that is retrieval-augmented generation (RAG).

What a model doesn't know

An artificial intelligence (AI) model learns from a huge collection of text, and then its training stops. After that, it hasn't seen last week's news. It never read your company's manuals, your policies, or your files. Ask it about them anyway and it may guess, sometimes confidently and wrongly. That's the gap RAG was built to fill. It helps by handing the model the right pages at the moment you ask.123

Retrieve

The first word in the name is retrieval. When your question comes in, the system searches a set of documents, like that company's help pages, and pulls out the few passages that best match what you asked. Often it searches by meaning. So a question about resetting can find a page that only says "restore factory settings." Think of a librarian who hands you the few pages that answer your question, instead of the whole shelf.1

Augment

Augmented means added to. The passages it found are placed in front of the model along with your question, inside what's called the context window: the text a model can see while it writes. So the model isn't working from memory alone, and nothing about the model itself has to change. The pages are right there, the way your notes are on the desk during an open-book test.12

Generate

Then comes generation. The model still does the writing, so the answer reads naturally. But the facts come from the passages it was handed, and many tools add links back to the exact pages it used. Those are the small numbered links you saw under the reply, back in the help chat.1

Why it helps

That simple loop brings three benefits.

  • Current: when the pages are updated, the answers can change the same day, with no retraining.
  • Specific: it can answer from one company's own documents, which no general model has read.
  • Checkable: the links show where each claim came from.

It's how many help bots, workplace assistants, and search tools answer from material the model never trained on. Chatbots that search the web before they answer work in a similar way.12

What can still slip

It can still go wrong. The search can pull the wrong page, or an out-of-date one. And the model can misread the right page, or add something the page never said. Retrieval makes a good answer more likely. It doesn't make one certain. So when an answer matters, open the link and read the sentence it's based on. Which help bot have you used lately that showed its sources?

Works cited

  1. IBM, "Retrieval augmented generation (RAG) architecture pattern." (checked )
  2. IBM Research, "What is retrieval-augmented generation (RAG)?" (checked )
  3. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks" (NeurIPS 2020) (checked )