Retrieval-Augmented Generation (RAG)

Definition
The technique of retrieving relevant documents first, then having a model write and cite its answer from that material.

Why it matters

Without retrieval, a model fills gaps with plausible guesses, which is what a hallucination is. With it, the answer is grounded in a real document that can be named and checked. A support assistant can quote the current returns policy. A sales tool can quote the actual price list instead of an outdated number remembered by a rep. The same idea explains how AI search engines choose what to cite: they retrieve pages, then summarise them, so a page that cannot be found or quoted cannot be cited.

How to apply it

  • Keep the source documents current. The answer is only as reliable as what is retrieved.
  • Write clear, self-contained statements that a chunk can carry on its own, with the heading and the facts together.
  • Check which sources get cited. A wrong answer is usually fixed in the document, not by arguing with the model.

What it is

A language model answers from what it learned in training, which may be out of date and knows nothing about your private files. RAG fixes that by adding a lookup step. When a question arrives, the system searches a body of documents, picks the passages most likely to help, pastes them into the prompt and asks the model to answer using only that material. The documents can be your handbook, a price list, a support archive or the open web.

Common mistakes

  • Assuming RAG removes hallucination. It reduces it, but poor retrieval or contradictory documents still produce wrong answers.
  • Feeding in everything, which crowds the context window and hides the useful passage.
  • Forgetting access rules, so the assistant quotes a document the asker should not see.

How it works

  • The documents are split into short chunks and stored in a searchable index, usually by meaning as well as by keyword.
  • A question is matched against the index, and the best few chunks are returned.
  • Those chunks go into the model's context window with the question.
  • The model writes an answer from them, and the system can show which chunks were used.
Worked example

Suppose a ten-person firm wants an assistant that answers staff questions about its returns policy. Without retrieval, the model answers from its training and may invent a policy that does not exist. The team splits its policy handbook into short sections and indexes them. For each question, the three most relevant sections are pasted into the prompt, and the model is asked to answer only from those sections and to name the one it used. The team builds the assistant on the Anthropic API, calling Claude from its own code. When the policy changes, only the handbook is updated. A person checks the first fifty answers, and any answer without a citation counts as a failure.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    Hallucination

    The failure retrieval is built to reduce.

  2. Article

    Context window

    The space retrieved passages must fit into.

  3. Article

    Context engineering

    The wider craft of deciding what a model sees.