Types of RAG (Naive, Advanced, Graph RAG, Agentic RAG)
Types of RAG are the main RAG architectures, the designs that connect retrieval to generation, ranging from a single retrieve-then-generate pass to autonomous agents that plan several searches. Each type changes how context is selected, which affects answer accuracy, latency, cost and overall implementation complexity.
- Naive RAG: One retrieval over one index, followed by one generation step.
- Advanced RAG: Naive RAG plus optimisation steps before and after retrieval, such as query rewriting and reranking.
- Modular RAG: Interchangeable components, such as routers and multiple indexes, assembled into a configurable pipeline.
- Graph RAG: Retrieval over a knowledge graph of entities and relationships extracted from the documents.
- Agentic RAG: An AI agent treats retrieval as a tool and decides when, where and how often to search.
For example, a documentation assistant for the Orders API can answer "What is the rate limit?" with naive RAG, but a question connecting webhooks, pagination and error codes may need graph or agentic RAG.
Overview of Types of RAG
| Type | How it works | Typical use |
|---|---|---|
| Naive RAG | Embed the question, retrieve the top chunks, generate an answer | Prototypes and small, well-structured document sets |
| Advanced RAG | Adds query rewriting, hybrid search, reranking and context compression | Production assistants that need higher precision |
| Modular RAG | Routes each question to specialised indexes and components | Several sources, such as docs, changelogs and tickets |
| Graph RAG | Retrieves connected entities and summaries from a knowledge graph | Questions about relationships across many documents |
| Agentic RAG | An agent plans searches, evaluates results and searches again | Multi-step, ambiguous or open-ended questions |
Naive RAG
- Mechanism: Documents are split through chunking, embedded and stored; each question retrieves the nearest chunks once.
- Strength: It is simple to implement, inexpensive to operate and easy to debug.
- Weakness: Retrieval errors pass directly to the model, because nothing checks or reorders the chunks.
Example: "How do I authenticate to the Orders API?" retrieves the Authentication chunk and answers from it.
Advanced RAG
- Pre-retrieval optimisation: The system rewrites or expands the question, for example adding "Retry-After" to a question about throttling.
- Improved retrieval: Hybrid search combines keyword and vector scores, which helps with exact identifiers such as error codes.
- Post-retrieval processing: Reranking in RAG reorders candidates, and compression removes irrelevant sentences before generation.
Example: "Why am I getting 429 errors?" is expanded to include "rate limit", then reranking places the Rate limits section first.
Modular RAG
- Routing: A classifier or rule sends each question to the appropriate index, such as the API reference, the changelog or the support tickets.
- Composability: Components such as retrievers, rerankers and generators can be replaced independently.
- Fusion: Results from several indexes are merged, often with reciprocal rank fusion, before generation.
Example: "When did pagination change to cursors?" is routed to the changelog index instead of the API reference.
Graph RAG
- Knowledge graph construction: An LLM extracts entities, such as endpoints, events and error codes, and the relationships between them.
- Relational retrieval: The system retrieves a connected subgraph or community summary instead of isolated chunks.
- Trade-off: Graph construction is expensive and must be repeated when documents change.
Example: "Which endpoints emit the order.cancelled webhook, and which errors can each return?" follows links from the event to endpoints and error codes.
Agentic RAG
- Autonomous retrieval: An AI agent decides whether retrieval is necessary and which source to query.
- Iterative refinement: The agent evaluates retrieved passages and issues a new query when the evidence is incomplete.
- Operational cost: Multiple model calls increase latency and cost, and each run can follow a different path.
Example: for "How do I resume a failed export of all cancelled orders?", the agent searches pagination, then webhooks, then error codes before answering. Agentic RAG covers this design in detail.
How to Choose Between Types of RAG
- Start with naive RAG: Build a baseline first, as in build a RAG chatbot in Python, and collect real questions.
- Measure before upgrading: Use RAG evaluation metrics to locate the failure: retrieval recall, ranking precision or answer faithfulness.
- Choose advanced RAG for ranking problems: Add reranking or hybrid search when the relevant chunk is retrieved but ranked too low.
- Choose modular RAG for multiple sources: Add routing when questions depend on distinct collections with different structures.
- Choose graph RAG for relationship questions: Use it when answers depend on links between entities spread across many documents.
- Choose agentic RAG for multi-step questions: Accept the added latency only when single-pass retrieval repeatedly fails.
These designs are not exclusive: an agentic system often calls an advanced retriever, and a modular pipeline can include a graph index. The fundamentals shared by every design are explained in what is RAG.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. A documentation assistant often retrieves the correct Orders API section but ranks it fourth. Which type of RAG addresses this most directly?
Frequently Asked Questions
Naive RAG vs advanced RAG: what is the difference?
Naive RAG embeds the question, retrieves the nearest chunks and passes them straight to the model. Advanced RAG adds steps before and after retrieval, such as query rewriting, hybrid search and reranking, to improve precision.
What is Graph RAG used for?
Graph RAG suits questions that depend on relationships between entities spread across many documents. It builds a knowledge graph of entities and links, then retrieves connected facts instead of isolated chunks.
Is agentic RAG better than standard RAG?
Agentic RAG handles multi-step and ambiguous questions more reliably, because an agent can search again or switch sources. It also adds latency, cost and more ways to fail, so simple lookups rarely need it.
Which type of RAG should a team build first?
Most teams start with naive RAG to get a working baseline and a test set. They then add advanced RAG techniques one at a time and keep only the changes that improve measured retrieval and answer quality.
Related Articles
- What is RAG (Retrieval-Augmented Generation)Retrieval-augmented generation (RAG) explained: how it retrieves document passages to ground LLM answers, its key steps, a Python example and its limits.
- How RAG WorksHow RAG works in detail: the offline indexing pipeline, the online query pipeline, a traced Python example over API docs, and where each stage can fail.
- Reranking in RAGReranking in RAG explained: how a second, stricter scorer reorders retrieved passages, cross-encoders vs bi-encoders, a Python example and the trade-offs.
- Agentic RAGAgentic RAG explained: how an agent decides when to retrieve, grades results and rewrites queries, with a Python incident runbook example and its limits.
- RAG Evaluation MetricsRAG evaluation metrics explained: hit rate, recall@k, MRR, faithfulness and answer relevance, with a Python example on a labelled test set and its limits.
- Hybrid Search (BM25 + Vector)Hybrid search explained: how BM25 keyword results and vector results are merged with reciprocal rank fusion, with a runnable Python runbook search example.