…
Skip to content
Topics
On this page

How RAG Works

Understanding how RAG works means seeing it as two connected pipelines: an offline indexing pipeline that prepares documents for search, and an online query pipeline that retrieves relevant chunks and generates a grounded answer. The indexing pipeline runs whenever documents change, while the query pipeline runs on every user question.

  • RAG indexing pipeline: The offline pipeline that loads, parses, chunks, embeds and stores documents before any question arrives.
  • RAG query pipeline: The online pipeline that embeds the question, retrieves candidate chunks, assembles a prompt and calls the language model.
  • Consistency: Both pipelines must use the same embedding model, otherwise query vectors and document vectors cannot be compared.
  • Index: The persistent store of chunk text, vectors and metadata that connects the indexing pipeline to the query pipeline.
  • Parameters: Retrieval settings such as top-k and the similarity threshold, which determine how much context reaches the language model.
The indexing pipeline and the query pipeline in RAGTop row, offline indexing: the Orders API Markdown file is chunked, each chunk is embedded, and the vectors are written to a vector store. Bottom row, online query: a developer question is embedded with the same model, a top-k search runs against the vector store, the best chunks are assembled into a prompt, and the LLM generates a cited answer. The layout is illustrative.Indexing pipeline (offline, when docs change)Query pipeline (online, every question)orders-api.mdChunkEmbedVector storeQuestionEmbedTop-k searchPromptLLM
The indexing pipeline and the query pipeline in RAG

For example, a documentation assistant indexes the Orders API reference once, then answers "Which header do I use to authenticate, and what does a 401 mean?" by combining the Authentication and Error codes sections.

Key Characteristics of How RAG Works

  • Phases: Indexing is a batch process with relaxed latency requirements, whereas querying is an interactive process that must respond within seconds.
  • Precomputation: Document embeddings are computed once and reused across requests, so only the incoming query is embedded at request time.
  • Metadata: Each chunk carries its source URL, section heading and API version, which later supports metadata filtering and accurate citation.
  • Narrowing: Retrieval returns a small ranked subset of the index, which keeps the assembled prompt within the model's context window.
  • Modularity: Each stage of the RAG architecture can be replaced independently, for example by switching the vector store, changing the chunker or introducing a reranker.
  • Propagation: An error in an early stage, such as poorly chosen chunk boundaries, degrades the output of every subsequent stage.

How RAG Works in Seven Stages

  1. Parsing: Source files, such as Markdown API reference pages, are loaded and converted into normalised text with their heading structure preserved.
  2. Chunking: The text is divided into retrievable segments, and the selected chunking strategies in RAG determine boundaries, size and overlap.
  3. Embedding: Each chunk is converted into a vector by an embedding model and written to a vector database together with its metadata.
  4. Encoding: At request time, the user question is converted into a query vector using the same embedding model that indexed the documents.
  5. Retrieval: The vector store returns the top-k most similar chunks, and matches below a similarity threshold are discarded. Many production pipelines add reranking in RAG at this point.
  6. Augmentation: The remaining chunks are numbered, labelled with their sources and combined with the question and an instruction to answer exclusively from them.
  7. Generation: The language model produces an answer, and the application maps each citation marker back to its original documentation section.

Example: Tracing a RAG Pipeline in Python

The code below traces both pipelines over a short Orders API document, using a toy concept-based embedding in place of a trained embedding model.

Python
import math
import re

DOC = """## Authentication
Send your API key in the Authorization header as a Bearer token.
## Rate limits
Each key may make 100 calls per minute. Extra calls return status 429.
## Pagination
GET /orders returns 50 items per page. Use the cursor parameter for more.
## Webhooks
A POST request is sent to your URL when an order changes status.
## Error codes
401 means the key is missing or invalid. 404 means the order was not found."""

# A toy embedding: one dimension per concept, so related words share a dimension.
CONCEPTS = {
    "auth": {"authenticate", "authentication", "authorization", "key", "bearer", "token", "401"},
    "limits": {"rate", "limit", "limits", "calls", "minute", "429"},
    "paging": {"pagination", "page", "cursor", "items"},
    "events": {"webhook", "webhooks", "post", "status", "url"},
    "errors": {"error", "errors", "401", "404", "429", "invalid", "mean", "means"},
}

def embed(text):
    words = re.findall(r"[a-z0-9]+", text.lower())
    return [sum(w in vocab for w in words) for vocab in CONCEPTS.values()]

def cosine(a, b):
    norm = math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(x * x for x in b))
    return sum(x * y for x, y in zip(a, b)) / norm if norm else 0.0

# Indexing pipeline (offline): chunk by heading, embed, store.
sections = re.findall(r"## (.+)\n(.+)", DOC)
index = [(title, text, embed(title + " " + text)) for title, text in sections]
print(f"[index] {len(index)} chunks stored, dimension {len(CONCEPTS)}")

# Query pipeline (online): embed, search, filter, build the prompt.
question = "Which header do I use to authenticate, and what does a 401 mean?"
q = embed(question)
print(f"[query] vector {q}")
ranked = sorted(((cosine(q, v), t, x) for t, x, v in index), reverse=True)
for score, title, _ in ranked:
    print(f"[search] {title:<15} {score:.2f}")
context = [(t, x) for s, t, x in ranked[:3] if s >= 0.5]  # top-k=3, threshold 0.5
print(f"[filter] kept {len(context)} of top 3 (threshold 0.5)")

prompt = "Answer only from the sources below and cite them as [n].\n"
prompt += "".join(f"[{n}] {t}: {x}\n" for n, (t, x) in enumerate(context, 1))
prompt += f"Question: {question}"
print("[prompt]\n" + prompt)

Output:

Example
[index] 5 chunks stored, dimension 5
[query] vector [2, 0, 0, 0, 2]
[search] Error codes     0.89
[search] Authentication  0.71
[search] Rate limits     0.23
[search] Webhooks        0.00
[search] Pagination      0.00
[filter] kept 2 of top 3 (threshold 0.5)
[prompt]
Answer only from the sources below and cite them as [n].
[1] Error codes: 401 means the key is missing or invalid. 404 means the order was not found.
[2] Authentication: Send your API key in the Authorization header as a Bearer token.
Question: Which header do I use to authenticate, and what does a 401 mean?
  • Shared concepts: "authenticate" and "Authorization" share the auth dimension, so the lexical mismatch seen in what is RAG no longer blocks the match.
  • Two sources: The question spans two distinct topics, so the prompt combines two numbered sources that the model can cite individually.
  • Threshold filter: The Rate limits chunk entered the top 3 but was excluded by the threshold, which keeps weakly related text out of the prompt.

Applications of the RAG Pipeline

  • API documentation: Answering endpoint, parameter and error-code questions from API reference documentation.
  • Runbook search: Retrieving remediation procedures for production alerts from operational runbooks.
  • Code search: Locating relevant functions and usage examples across a large code repository.
  • Support deflection: Answering product questions from a help centre before a customer opens a support ticket.
  • Answer engines: Combining hybrid search with generation to produce cited summaries of search results.

Advantages

  • Incremental updates: Only modified documents need to pass through the indexing pipeline again.
  • Predictable latency: Heavy processing happens offline, so each request performs one vector search and one model call.
  • Easy debugging: Every stage produces an inspectable intermediate result, such as similarity scores or the final prompt.
  • Flexibility: Chunkers, embedding models and vector stores can be exchanged without redesigning the entire system.

Limitations

  • Cascading errors: Defects from parsing or chunking cannot be corrected by later stages.
  • Parameter tuning: Top-k, chunk size and the similarity threshold interact, and each must be tuned against real questions.
  • Model upgrades: Upgrading the embedding model requires re-embedding the entire document corpus.
  • Multi-hop questions: Questions that require facts from several unrelated documents are often retrieved only in part.
  • Evaluation: Each stage requires its own measurements, as described in RAG evaluation metrics.

The same pipeline, extended with conversation history, forms the basis of the RAG chatbot in Python.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. Which part of a RAG system runs before any user asks a question?

Frequently Asked Questions

What are the main stages of a RAG workflow?

A RAG pipeline has an indexing stage and a query stage. Indexing loads, chunks, embeds and stores documents ahead of time. The query stage embeds the question, retrieves matching chunks, builds a prompt and generates the answer.

Does RAG re-embed all documents for every question?

No. Documents are embedded once during indexing and stored with their vectors. Only the incoming question is embedded at query time, which keeps retrieval fast.

How many chunks should a RAG system retrieve?

Many systems retrieve between 3 and 10 chunks, but the right value depends on chunk size, the context window and the question type. Teams usually tune this number against an evaluation set rather than fixing it in advance.

What happens when a RAG system retrieves nothing relevant?

A well-designed pipeline applies a similarity threshold and drops weak matches. If nothing passes the threshold, the prompt instructs the model to state that the documentation does not cover the question instead of guessing.

Where does reranking fit in a RAG pipeline?

Reranking runs after the first retrieval and before prompt assembly. A more precise model rescores a larger candidate set and passes only the best few chunks to the language model.