What is RAG (Retrieval-Augmented Generation)
Retrieval-augmented generation (RAG) is a technique that connects a large language model to an external document collection at inference time. Before generating a response, the system retrieves the most relevant passages and inserts them into the prompt, so the answer is grounded in that source text.
- RAG meaning: RAG stands for Retrieval-Augmented Generation, a term introduced in 2020 research on knowledge-intensive language tasks.
- Retriever: A component that searches an indexed document collection and returns the passages most relevant to a query.
- Augmented prompt: The user query combined with the retrieved passages and an instruction to answer only from them.
- Grounding: Constraining a generated answer to supplied source text instead of the model's parametric knowledge alone.
- Primary use: Answering questions over private or frequently updated information without retraining the model.
For example, a developer asks how to authenticate requests to the Orders API, and the documentation assistant answers from the authentication section of the team's API docs.
Key Characteristics of RAG
- External knowledge: Facts are stored in documents outside the model, so editing a document changes subsequent answers.
- No retraining: The model weights remain unchanged; only the index is rebuilt when the underlying content changes.
- Source attribution: Each answer can cite the specific passage it used, which simplifies verification.
- Model independence: The same index can serve several language models, including models from different providers.
- Retrieval dependence: Answer quality is bounded by the relevance and completeness of the retrieved passages.
- Reduced hallucination: Grounding in authoritative text lowers the rate of LLM hallucination, although it does not eliminate it.
How Retrieval-Augmented Generation Works
- Indexing: Documents are divided into short segments through chunking. Each chunk is converted into embeddings, numerical vectors that represent meaning, and stored in a vector database.
- Retrieval: The query is converted into an embedding and compared with the stored vectors. This semantic search returns the chunks closest in meaning, independent of exact spelling.
- Augmentation: The highest-ranked chunks, the query and an instruction such as "answer only from these passages" are assembled into a single prompt.
- Generation: The language model processes the augmented prompt and generates an answer from the supplied context.
- Citation: The system attaches the source passages to the response so that users can verify each statement.
Example: RAG Retrieval in Python
The code below is a minimal RAG example: it scores five API documentation passages against a developer question and assembles the augmented prompt, using only the Python standard library.
import math
import re
from collections import Counter
# Short passages from a team's API documentation.
docs = {
"authentication": "Authentication: to call the Orders API, send your API key in the Authorization header as a Bearer token.",
"rate-limits": "Rate limits: each key may make 100 calls per minute. Extra calls return status 429.",
"pagination": "Pagination: list endpoints such as GET /orders return 50 items per page. Use the cursor parameter for the next page.",
"webhooks": "Webhooks: a POST request is sent to your URL when an order changes status.",
"error-codes": "Error codes: 401 means the key is missing or invalid, and 404 means the order was not found.",
}
STOP = {"a", "an", "the", "is", "to", "i", "do", "how", "in", "as", "your", "per", "for", "or", "and", "when"}
def vector(text):
return Counter(w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP)
def cosine(a, b):
dot = sum(a[w] * b[w] for w in a.keys() & b.keys())
norm = math.sqrt(sum(v * v for v in a.values())) * math.sqrt(sum(v * v for v in b.values()))
return dot / norm if norm else 0.0
question = "How do I authenticate requests to the Orders API?"
q = vector(question)
scores = sorted(((round(cosine(q, vector(t)), 2), k) for k, t in docs.items()), key=lambda s: (-s[0], s[1]))
for score, key in scores:
print(f"{key:<15} {score:.2f}")
best = scores[0][1]
print("\nAugmented prompt:")
print(f"Answer only from this passage.\nPassage: {docs[best]}\nQuestion: {question}")Output:
authentication 0.42
pagination 0.12
error-codes 0.00
rate-limits 0.00
webhooks 0.00
Augmented prompt:
Answer only from this passage.
Passage: Authentication: to call the Orders API, send your API key in the Authorization header as a Bearer token.
Question: How do I authenticate requests to the Orders API?- Ranking: The authentication passage ranks first because it shares the terms "orders" and "api" with the question.
- Lexical mismatch: Word counting treats "authenticate" and "authentication" as unrelated tokens, and likewise "requests" and "request", so these related terms contribute nothing to the score.
- Production approach: Production systems use embeddings and a vector database, which compare semantic similarity instead of surface spelling.
RAG vs Using an LLM Alone
| Aspect | LLM Alone | RAG |
|---|---|---|
| Source of facts | Training data up to a fixed cut-off date | Documents retrieved for each individual query |
| Private information, such as internal API docs | Not available to the model | Available once the documents are indexed |
| Answer to "How do I authenticate to the Orders API?" | A generic guess about API keys or OAuth | The exact method specified in the team's documentation |
| Source citations | Not possible | Links to the documentation section used |
| Setup effort | Minimal configuration | Requires an index, a retriever and a prompt template |
Applications of RAG
- Developer documentation: Answering API and SDK questions from a team's reference documentation.
- Customer support: Resolving billing, account and policy questions from internal support articles.
- Internal knowledge assistants: Answering employee questions from HR policies and IT operations runbooks.
- Financial services: Locating specific clauses in loan agreements, card terms and insurance policies.
- Legal and compliance: Searching contracts and regulations for applicable sections.
- Agentic systems: Agentic AI systems that decide what to retrieve and when to search again.
Advantages
- Current answers: A documentation update is reflected in the next generated answer.
- Lower cost: Indexing documents is considerably cheaper than fine-tuning a model.
- Traceability: Cited passages allow users to audit each answer against its source.
- Access control: Retrieval can be restricted to documents a particular user is authorised to read.
- Fewer fabricated facts: Answers are anchored to authoritative source text.
Limitations
- Retrieval quality: If the relevant chunk is not retrieved, the model generates its answer from irrelevant context.
- Chunking trade-offs: Small chunks lose surrounding context, whereas large chunks introduce unrelated text.
- Latency and cost: Retrieval adds latency, and longer prompts consume more of the context window.
- Stale index: Answers remain outdated until modified documents are re-indexed.
- Evaluation complexity: Retrieval and generation must be evaluated separately, which increases testing effort.
The indexing and query pipelines are examined stage by stage in how RAG works. RAG suits knowledge that is private or changes frequently. For adapting a model's style, output format or specialised behaviour, see RAG vs fine-tuning.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In retrieval-augmented generation, what happens just before the model writes its answer?
Frequently Asked Questions
What is RAG in AI?
RAG stands for Retrieval-Augmented Generation. The system retrieves relevant passages from a document collection and adds them to the prompt. The language model then writes its answer from those passages.
Does RAG stop an LLM from hallucinating?
RAG reduces hallucination but does not remove it. If the retriever returns the wrong passage, or the model ignores the passage, the answer can still be wrong. Citations and regular evaluation help catch these errors.
Is a vector database required for RAG?
It is not required for small collections, which can be searched in memory. For thousands of documents or search by meaning, a vector database is the usual choice.
How is RAG different from fine-tuning?
RAG supplies facts at query time from an external index, so an update only needs a document change. Fine-tuning changes the model weights and suits a fixed style, format or narrow skill.
What are the main components of a RAG system?
A RAG system has a document index, a retriever, a prompt template and a language model. Many systems also add a reranker and a citation step.
Related Articles
- How RAG WorksHow RAG works in detail: the offline indexing pipeline, the online query pipeline, a traced Python example over API docs, and where each stage can fail.
- RAG vs Fine-TuningRAG vs fine-tuning compared: what each changes, cost, freshness and accuracy trade-offs, when to use each, and one API docs example handled both ways.
- Chunking Strategies in RAGChunking strategies in RAG explained: fixed-size, heading-based and semantic chunking, how to choose chunk size and overlap, with a Python comparison.
- Vector Database in RAGVector database in RAG explained: how embeddings are stored, indexed and searched with cosine similarity, ANN indexes, metadata filters and a Python demo.
- LLM Hallucination: Causes and FixesLearn what LLM hallucination is, why models invent facts and code APIs, and how to reduce it with RAG and validation, with a runnable Python checker.
- Build a RAG Chatbot in PythonBuild a RAG chatbot in Python: chunk API docs, retrieve the closest passage, build a grounded prompt, then call Claude with the Anthropic SDK. Full code.