Agentic RAG
Agentic RAG is a retrieval-augmented generation design in which an AI agent controls retrieval instead of a fixed pipeline. The agent decides whether to retrieve, which source to search, whether the results are relevant and when to rewrite the query and search again.
- Retrieval decision: The agent first determines whether a question requires stored documents, a live operational tool or neither.
- Source routing: It selects the appropriate knowledge source, such as operational runbooks, previous incident reports or a monitoring API.
- Relevance grading: It evaluates the retrieved passages and rejects irrelevant matches before any answer is generated.
- Query reformulation: It rewrites an unsuccessful query using the terminology of the documents and then repeats retrieval.
- Bounded retries: A fixed attempt limit terminates the loop with either a cited answer or an escalation to a person.
For example, an on-call assistant asked why checkout returns 502 after a release searches the runbooks, finds weak matches, rewrites the query with "5xx" and "deploy", and then retrieves the matching incident report.
Key Characteristics of Agentic RAG
- Agent control: Retrieval is a tool that the retrieval agent calls through tool calling, not a step that runs on every request.
- Multiple sources: A single agent can query a vector index, a keyword index and live operational APIs within one investigation.
- Iteration: Retrieval can repeat several times per question, following the reason, act and observe cycle of a ReAct agent.
- Self-correction: Weak or irrelevant retrieval results trigger a revised query instead of an unreliable answer, the core idea of self-corrective RAG.
- Grounded answers: The final response cites only the passages that passed the relevance evaluation.
- Escalation: When no reliable context exists, the agent reports the absence of evidence or transfers the question to an engineer.
How Agentic RAG Works
- Question analysis: The agent classifies the question as requiring documentation, live operational data or a direct response.
- Source selection: It chooses a source and writes a search query, often with a metadata filter such as service name.
- Retrieval: The retriever returns the top-k passages together with their similarity scores.
- Grading: The agent, or a separate grader model, evaluates whether the retrieved passages actually answer the question.
- Reformulation: If the relevance grade is below a threshold, the agent rewrites the query and returns to step 3.
- Generation: With relevant passages available, the language model produces an answer that cites its sources.
- Termination: After the configured attempt limit, the agent escalates instead of generating a speculative answer.
Example: An Agentic RAG Loop for Incident Runbooks in Python
The program below routes three on-call questions, grades keyword-overlap retrieval against a threshold and reformulates weak queries once.
import re
# Knowledge base: runbooks and past incident reports.
DOCS = {
"runbook-checkout-5xx": "checkout 5xx errors: compare error rate before and after the last deploy, request rollback approval",
"runbook-db-pool": "database connection pool exhausted: timeouts acquiring connections, raise pool size",
"INC-0412": "incident report: checkout 5xx after deploy, payment client timeout, resolved by rollback",
"INC-0388": "incident report: CDN cache purge slowed static pages, no 5xx errors",
}
METRICS = {"checkout": "5xx rate 9.4% (normal 0.2%)"}
SYNONYMS = {"502": "5xx", "503": "5xx", "release": "deploy", "fix": "resolved"}
STOP = {"the", "a", "is", "on", "what", "why", "did", "we", "do", "last", "time",
"after", "today", "s", "returns", "how", "it"}
def terms(text):
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP}
def retrieve(query, k=2):
q = terms(query)
scored = [(len(q & terms(doc)) / len(q), doc_id) for doc_id, doc in DOCS.items()]
return sorted(scored, reverse=True)[:k]
def reformulate(query):
# Stand-in for the LLM: rewrite the query in the knowledge base's vocabulary.
return " ".join(sorted({SYNONYMS.get(w, w) for w in terms(query)}))
def agent(question, service="checkout", min_score=0.5, max_tries=2):
print(f"Q: {question}")
if "current" in question: # decide: live state comes from a tool, not documents
print(f" route: skip retrieval, get_metrics -> {METRICS[service]}")
return
query = question
for attempt in range(1, max_tries + 1):
hits = retrieve(query)
print(f" retrieve #{attempt} '{query}': " + ", ".join(f"{d} {s:.2f}" for s, d in hits))
good = [d for s, d in hits if s >= min_score]
if good:
print(f" grade: relevant, answer with citations {good}")
return
if attempt < max_tries:
query = reformulate(query)
print(" grade: weak match, reformulate")
print(" grade: no reliable context, escalate to the on-call engineer")
agent("What is the current 5xx rate on checkout?")
agent("Checkout returns 502 after today's release, how did we fix it?")
agent("Why is the search index stale?")Output:
Q: What is the current 5xx rate on checkout?
route: skip retrieval, get_metrics -> 5xx rate 9.4% (normal 0.2%)
Q: Checkout returns 502 after today's release, how did we fix it?
retrieve #1 'Checkout returns 502 after today's release, how did we fix it?': runbook-checkout-5xx 0.25, INC-0412 0.25
grade: weak match, reformulate
retrieve #2 '5xx checkout deploy resolved': INC-0412 1.00, runbook-checkout-5xx 0.75
grade: relevant, answer with citations ['INC-0412', 'runbook-checkout-5xx']
Q: Why is the search index stale?
retrieve #1 'Why is the search index stale?': runbook-db-pool 0.00, runbook-checkout-5xx 0.00
grade: weak match, reformulate
retrieve #2 'index search stale': runbook-db-pool 0.00, runbook-checkout-5xx 0.00
grade: no reliable context, escalate to the on-call engineer- Routing: The live error rate comes from a monitoring tool, because no runbook can report the current state of a production service.
- Reformulation: Mapping "502" to "5xx" and "release" to "deploy" raises the top score from 0.25 to 1.00.
- Escalation: No document describes the search index, so the agent escalates to the on-call engineer instead of fabricating an answer.
In production, an LLM performs the routing, grading and rewriting, and retrieval uses a vector database or hybrid search.
Applications of Agentic RAG
- Incident response: Combining operational runbooks, previous incident reports and live metrics during an outage.
- Engineering support: Answering questions that span API documentation, source code and issue history.
- Research assistants: Decomposing a broad question into sub-queries across several document collections.
- Compliance review: Verifying a proposed configuration change against several security policy documents in sequence.
- Customer support: Searching product documentation first and consulting account data only when necessary.
- Multi-agent pipelines: Serving as the research role inside multi-agent systems.
Advantages
- Higher recall: Reformulated queries retrieve relevant documents that the original wording failed to match.
- Lower overhead: Questions that require no documentation skip retrieval entirely, which reduces latency and cost.
- Fewer unsupported answers: Relevance grading rejects irrelevant context before the generation stage begins.
- Flexibility: New sources are added as tools without redesigning the pipeline, a step beyond the designs in types of RAG.
Limitations
- Latency and cost: Every grading and retry step adds another model call, which increases response time and inference cost.
- Loop risk: Without an attempt limit, the agent can repeat unproductive searches indefinitely.
- Harder testing: Different runs can take different retrieval paths, which requires AI agent evaluation as well as RAG evaluation.
- Untrusted content: Retrieved text can carry hidden instructions, a risk covered in prompt injection.
- Grader errors: An inaccurate grader can reject useful passages or accept irrelevant ones.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In the Python example, why does the agent skip retrieval for the first question?
Frequently Asked Questions
What is agentic RAG, and how does it differ from standard RAG?
Agentic RAG is a RAG design that places an agent in control of retrieval, so it can skip it, choose a source, grade the results and search again with a better query. Standard RAG retrieves once for every question and then generates an answer.
Agentic RAG vs RAG: when should each be used?
Agentic RAG suits questions that need several sources, several lookups or a mix of documents and live tool data. For simple lookups over one well-indexed collection, standard RAG is cheaper and faster.
What is query reformulation in agentic RAG?
Query reformulation is the step where the agent rewrites a search query after a weak result. It replaces vague words with the terms used in the documents, splits a broad question or adds context such as the service name.
Does agentic RAG reduce hallucination?
It can reduce hallucination because the agent grades retrieved passages and escalates when none are relevant, instead of answering from weak context. It does not remove the risk, so answers still need citations and evaluation.
Related Articles
- What is RAG (Retrieval-Augmented Generation)Retrieval-augmented generation (RAG) explained: how it retrieves document passages to ground LLM answers, its key steps, a Python example and its limits.
- Types of RAG (Naive, Advanced, Graph RAG, Agentic RAG)Types of RAG compared: naive, advanced, modular, graph RAG and agentic RAG, how each retrieves context, typical uses, and how to choose the right design.
- Tool Calling (Function Calling) in LLMLearn how tool calling works in LLMs: JSON Schema tool definitions, model tool calls, argument validation and tool results, with a Python DevOps example.
- ReAct Agent ExplainedLearn how a ReAct agent alternates thought, action and observation, with a Python example that traces a failing unit test to a commit and opens a ticket.
- RAG Evaluation MetricsRAG evaluation metrics explained: hit rate, recall@k, MRR, faithfulness and answer relevance, with a Python example on a labelled test set and its limits.
- AI Agent EvaluationAI agent evaluation explained: task success, tool-call accuracy, trajectory checks, cost, latency and regression suites, with a Python trajectory scorer.