…
Skip to content
Topics
On this page

What is RAG (Retrieval-Augmented Generation)

Retrieval-augmented generation (RAG) is a technique that connects a large language model to an external document collection at inference time. Before generating a response, the system retrieves the most relevant passages and inserts them into the prompt, so the answer is grounded in that source text.

  • RAG meaning: RAG stands for Retrieval-Augmented Generation, a term introduced in 2020 research on knowledge-intensive language tasks.
  • Retriever: A component that searches an indexed document collection and returns the passages most relevant to a query.
  • Augmented prompt: The user query combined with the retrieved passages and an instruction to answer only from them.
  • Grounding: Constraining a generated answer to supplied source text instead of the model's parametric knowledge alone.
  • Primary use: Answering questions over private or frequently updated information without retraining the model.
How retrieval-augmented generation answers a questionA developer asks how to authenticate requests to the Orders API. The system searches the team's API documentation (authentication, rate limits, pagination, webhooks, error codes), picks the Authentication passage as the top match, builds an augmented prompt from the question plus that passage, sends it to the LLM, and returns a grounded answer: send the API key in the Authorization header as a Bearer token.Developer questionAuth for Orders API?Search API docsAuth, limits, webhooksTop passageAuthenticationAugmented promptQuestion + passageLLMAnswers from passageGrounded answerBearer <API key>
How retrieval-augmented generation answers a question

For example, a developer asks how to authenticate requests to the Orders API, and the documentation assistant answers from the authentication section of the team's API docs.

Key Characteristics of RAG

  • External knowledge: Facts are stored in documents outside the model, so editing a document changes subsequent answers.
  • No retraining: The model weights remain unchanged; only the index is rebuilt when the underlying content changes.
  • Source attribution: Each answer can cite the specific passage it used, which simplifies verification.
  • Model independence: The same index can serve several language models, including models from different providers.
  • Retrieval dependence: Answer quality is bounded by the relevance and completeness of the retrieved passages.
  • Reduced hallucination: Grounding in authoritative text lowers the rate of LLM hallucination, although it does not eliminate it.

How Retrieval-Augmented Generation Works

  1. Indexing: Documents are divided into short segments through chunking. Each chunk is converted into embeddings, numerical vectors that represent meaning, and stored in a vector database.
  2. Retrieval: The query is converted into an embedding and compared with the stored vectors. This semantic search returns the chunks closest in meaning, independent of exact spelling.
  3. Augmentation: The highest-ranked chunks, the query and an instruction such as "answer only from these passages" are assembled into a single prompt.
  4. Generation: The language model processes the augmented prompt and generates an answer from the supplied context.
  5. Citation: The system attaches the source passages to the response so that users can verify each statement.

Example: RAG Retrieval in Python

The code below is a minimal RAG example: it scores five API documentation passages against a developer question and assembles the augmented prompt, using only the Python standard library.

Python
import math
import re
from collections import Counter

# Short passages from a team's API documentation.
docs = {
    "authentication": "Authentication: to call the Orders API, send your API key in the Authorization header as a Bearer token.",
    "rate-limits": "Rate limits: each key may make 100 calls per minute. Extra calls return status 429.",
    "pagination": "Pagination: list endpoints such as GET /orders return 50 items per page. Use the cursor parameter for the next page.",
    "webhooks": "Webhooks: a POST request is sent to your URL when an order changes status.",
    "error-codes": "Error codes: 401 means the key is missing or invalid, and 404 means the order was not found.",
}
STOP = {"a", "an", "the", "is", "to", "i", "do", "how", "in", "as", "your", "per", "for", "or", "and", "when"}

def vector(text):
    return Counter(w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP)

def cosine(a, b):
    dot = sum(a[w] * b[w] for w in a.keys() & b.keys())
    norm = math.sqrt(sum(v * v for v in a.values())) * math.sqrt(sum(v * v for v in b.values()))
    return dot / norm if norm else 0.0

question = "How do I authenticate requests to the Orders API?"
q = vector(question)
scores = sorted(((round(cosine(q, vector(t)), 2), k) for k, t in docs.items()), key=lambda s: (-s[0], s[1]))
for score, key in scores:
    print(f"{key:<15} {score:.2f}")

best = scores[0][1]
print("\nAugmented prompt:")
print(f"Answer only from this passage.\nPassage: {docs[best]}\nQuestion: {question}")

Output:

Example
authentication  0.42
pagination      0.12
error-codes     0.00
rate-limits     0.00
webhooks        0.00

Augmented prompt:
Answer only from this passage.
Passage: Authentication: to call the Orders API, send your API key in the Authorization header as a Bearer token.
Question: How do I authenticate requests to the Orders API?
  • Ranking: The authentication passage ranks first because it shares the terms "orders" and "api" with the question.
  • Lexical mismatch: Word counting treats "authenticate" and "authentication" as unrelated tokens, and likewise "requests" and "request", so these related terms contribute nothing to the score.
  • Production approach: Production systems use embeddings and a vector database, which compare semantic similarity instead of surface spelling.

RAG vs Using an LLM Alone

AspectLLM AloneRAG
Source of factsTraining data up to a fixed cut-off dateDocuments retrieved for each individual query
Private information, such as internal API docsNot available to the modelAvailable once the documents are indexed
Answer to "How do I authenticate to the Orders API?"A generic guess about API keys or OAuthThe exact method specified in the team's documentation
Source citationsNot possibleLinks to the documentation section used
Setup effortMinimal configurationRequires an index, a retriever and a prompt template

Applications of RAG

  • Developer documentation: Answering API and SDK questions from a team's reference documentation.
  • Customer support: Resolving billing, account and policy questions from internal support articles.
  • Internal knowledge assistants: Answering employee questions from HR policies and IT operations runbooks.
  • Financial services: Locating specific clauses in loan agreements, card terms and insurance policies.
  • Legal and compliance: Searching contracts and regulations for applicable sections.
  • Agentic systems: Agentic AI systems that decide what to retrieve and when to search again.

Advantages

  • Current answers: A documentation update is reflected in the next generated answer.
  • Lower cost: Indexing documents is considerably cheaper than fine-tuning a model.
  • Traceability: Cited passages allow users to audit each answer against its source.
  • Access control: Retrieval can be restricted to documents a particular user is authorised to read.
  • Fewer fabricated facts: Answers are anchored to authoritative source text.

Limitations

  • Retrieval quality: If the relevant chunk is not retrieved, the model generates its answer from irrelevant context.
  • Chunking trade-offs: Small chunks lose surrounding context, whereas large chunks introduce unrelated text.
  • Latency and cost: Retrieval adds latency, and longer prompts consume more of the context window.
  • Stale index: Answers remain outdated until modified documents are re-indexed.
  • Evaluation complexity: Retrieval and generation must be evaluated separately, which increases testing effort.

The indexing and query pipelines are examined stage by stage in how RAG works. RAG suits knowledge that is private or changes frequently. For adapting a model's style, output format or specialised behaviour, see RAG vs fine-tuning.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In retrieval-augmented generation, what happens just before the model writes its answer?

Frequently Asked Questions

What is RAG in AI?

RAG stands for Retrieval-Augmented Generation. The system retrieves relevant passages from a document collection and adds them to the prompt. The language model then writes its answer from those passages.

Does RAG stop an LLM from hallucinating?

RAG reduces hallucination but does not remove it. If the retriever returns the wrong passage, or the model ignores the passage, the answer can still be wrong. Citations and regular evaluation help catch these errors.

Is a vector database required for RAG?

It is not required for small collections, which can be searched in memory. For thousands of documents or search by meaning, a vector database is the usual choice.

How is RAG different from fine-tuning?

RAG supplies facts at query time from an external index, so an update only needs a document change. Fine-tuning changes the model weights and suits a fixed style, format or narrow skill.

What are the main components of a RAG system?

A RAG system has a document index, a retriever, a prompt template and a language model. Many systems also add a reranker and a citation step.