Build a RAG Chatbot in Python
A RAG chatbot in Python is a program that answers questions by retrieving the most relevant passage from a document collection and passing it to a language model with the question. It combines chunking, retrieval, prompt construction and generation in a loop, so every answer is grounded in the supplied documentation.
- Document loader: Code that reads the source documentation into memory as text.
- Chunker: A function that splits the documentation into sections that can be retrieved independently.
- Retriever: A scoring function that compares the question with every chunk and returns the closest match.
- Prompt construction: A template that combines an instruction, the retrieved passage and the question.
- Generator: The language model call that writes the answer; Part 1 uses a clearly labelled stand-in.
For example, a developer asks the documentation chatbot how to fetch the next page of orders, and it answers from the Pagination section of the Orders API documentation.
The implementation is divided into two parts. Part 1 (Steps 1 and 2) shows how to build RAG from scratch with the standard library only, and a template quotes the passage in place of the model. Part 2 (Step 3) replaces that stand-in with a real call to Claude through the Anthropic Python SDK.
Prerequisites
- Python 3.11: Any recent Python 3 release is compatible; the examples were tested with version 3.11.
- RAG fundamentals: Familiarity with retrieval-augmented generation and the stages described in how RAG works.
- Terminal familiarity: Basic use of a command-line shell to create directories, execute scripts and define environment variables.
- API key (Part 2 only): An Anthropic API key, stored in an environment variable and never written into the code.
Setup
Create a project directory. Part 1 requires no additional installation; for Part 2, create a virtual environment and install the pinned SDK version.
python3.11 --version
mkdir rag-chatbot && cd rag-chatbot
# Part 2 only
python3.11 -m venv .venv
source .venv/bin/activate
pip install anthropic==1.0.0
export ANTHROPIC_API_KEY="paste-your-key-here"Step 1: Load and Chunk the Orders API Docs
The documentation is stored as one string with a heading per section, and the chunker divides it at every heading. In a production project, the string would be loaded from Markdown files with open() or retrieved from a documentation repository.
# Step 1: the Orders API docs, split into one chunk per "## " heading.
DOCS = """
## Authentication
Send your API key in the Authorization header as a Bearer token. A request without a valid key returns error 401.
## Rate limits
Each API key may send 100 requests per minute. A request over the limit returns error 429 with a Retry-After header.
## Pagination
GET /orders returns 50 orders per page. Pass next_cursor from the response as the cursor parameter to get the next page.
## Webhooks
When an order changes status, the API sends a POST request to your webhook URL. A failed delivery is retried three times.
## Error codes
400 means an invalid request body, 401 means a missing or invalid key, 404 means the order was not found and 429 means too many requests.
"""
def chunk(text):
chunks = []
for block in text.split("## ")[1:]:
title, _, body = block.partition("\n")
chunks.append({"title": title.strip(), "text": body.strip()})
return chunks
chunks = chunk(DOCS)
for c in chunks:
print(f"{c['title']:<15} {len(c['text'].split())} words")Output:
Authentication 21 words
Rate limits 21 words
Pagination 21 words
Webhooks 22 words
Error codes 26 words- Section boundaries: Each heading becomes one chunk, which suits reference documentation with short, self-contained sections, as explained in chunking in RAG.
- Metadata: The section title is stored alongside the text, so every answer can identify and cite its source section.
Step 2: Build the RAG Chatbot in Python
Save the complete program as rag_chatbot.py. It repeats Step 1 so that it runs independently, then adds scoring, retrieval, prompt construction and the stand-in answer.
vector(): Converts text into a word-frequency vector after removing stop words and applying a crude seven-letter stem.cosine(): Calculates the cosine similarity between two vectors, a value between 0 and 1 that measures lexical overlap.retrieve(): Scores every chunk, including its section title, and returns the highest-scoring chunk with its similarity score.build_prompt(): Combines a grounding instruction, the retrieved passage and the question into one augmented prompt.answer(): Represents the language model; in Part 1 it only quotes the passage, so the output is deterministic.
# rag_chatbot.py: a retrieval chatbot over the Orders API docs.
import math
import re
from collections import Counter
DOCS = """
## Authentication
Send your API key in the Authorization header as a Bearer token. A request without a valid key returns error 401.
## Rate limits
Each API key may send 100 requests per minute. A request over the limit returns error 429 with a Retry-After header.
## Pagination
GET /orders returns 50 orders per page. Pass next_cursor from the response as the cursor parameter to get the next page.
## Webhooks
When an order changes status, the API sends a POST request to your webhook URL. A failed delivery is retried three times.
## Error codes
400 means an invalid request body, 401 means a missing or invalid key, 404 means the order was not found and 429 means too many requests.
"""
STOP = {"a", "an", "the", "is", "to", "i", "do", "how", "my", "if", "of", "what", "does", "as", "from", "with"}
MIN_SCORE = 0.15 # below this, the docs do not cover the question
def chunk(text):
chunks = []
for block in text.split("## ")[1:]:
title, _, body = block.partition("\n")
chunks.append({"title": title.strip(), "text": body.strip()})
return chunks
def vector(text):
# Crude stemming: "authenticate" and "authentication" share "authent".
return Counter(w.rstrip("s")[:7] for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP)
def cosine(a, b):
dot = sum(a[w] * b[w] for w in a.keys() & b.keys())
norm = math.sqrt(sum(v * v for v in a.values())) * math.sqrt(sum(v * v for v in b.values()))
return dot / norm if norm else 0.0
def retrieve(question, chunks):
q = vector(question)
scored = [(cosine(q, vector(c["title"] + " " + c["text"])), c) for c in chunks]
return max(scored, key=lambda s: s[0])
def build_prompt(question, passage):
return (f"Answer only from the passage. If it does not answer the question, say so.\n"
f"Passage ({passage['title']}): {passage['text']}\nQuestion: {question}")
def answer(prompt, passage):
# Stand-in for the model: quotes the passage instead of generating text.
return f"[stand-in, no model] From the {passage['title']} docs: \"{passage['text']}\""
def main():
chunks = chunk(DOCS)
questions = [ # in a terminal, use: question = input("You: ")
"How do I authenticate my requests?",
"What happens if I go over the rate limit?",
"How do I get the next page of orders?",
"Does the API support GraphQL?",
]
for turn, question in enumerate(questions, 1):
score, passage = retrieve(question, chunks)
print(f"You: {question}")
if score < MIN_SCORE:
print(f"Bot: The Orders API docs do not cover this. (top score {score:.2f})\n")
continue
prompt = build_prompt(question, passage)
if turn == 1:
print(f"--- prompt sent to the model ---\n{prompt}\n---")
print(f"Bot: {answer(prompt, passage)} (score {score:.2f})\n")
if __name__ == "__main__":
main()Output
You: How do I authenticate my requests?
--- prompt sent to the model ---
Answer only from the passage. If it does not answer the question, say so.
Passage (Authentication): Send your API key in the Authorization header as a Bearer token. A request without a valid key returns error 401.
Question: How do I authenticate my requests?
---
Bot: [stand-in, no model] From the Authentication docs: "Send your API key in the Authorization header as a Bearer token. A request without a valid key returns error 401." (score 0.32)
You: What happens if I go over the rate limit?
Bot: [stand-in, no model] From the Rate limits docs: "Each API key may send 100 requests per minute. A request over the limit returns error 429 with a Retry-After header." (score 0.37)
You: How do I get the next page of orders?
Bot: [stand-in, no model] From the Pagination docs: "GET /orders returns 50 orders per page. Pass next_cursor from the response as the cursor parameter to get the next page." (score 0.77)
You: Does the API support GraphQL?
Bot: The Orders API docs do not cover this. (top score 0.13)- Retrieval scores: The Pagination chunk scores 0.77 because the question shares "get", "next", "page" and "orders" with it; the other questions score lower but still select the correct section.
- Refusal threshold: The GraphQL question scores below
MIN_SCORE, so the chatbot states that the documentation does not cover it instead of guessing. - Placeholder generation: The
[stand-in, no model]label marks text produced by a template, not by a language model. - Similarity measure: Word-count cosine similarity is a simplified substitute for embeddings stored in a vector database.
Step 3: Replace the Stand-in with a Real LLM Call
Part 2 completes RAG with Claude: it imports the existing retrieval functions and replaces only the answer() function with an authenticated request to the Messages API. Save this as rag_chatbot_llm.py in the same folder, with the virtual environment active and ANTHROPIC_API_KEY exported.
# rag_chatbot_llm.py: Part 2, the same chatbot with a real model call.
import os
import anthropic
from rag_chatbot import DOCS, MIN_SCORE, build_prompt, chunk, retrieve
MODEL = "claude-sonnet-5" # check the current model name before running
client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
def answer(prompt):
response = client.messages.create(
model=MODEL,
max_tokens=1024,
system="You answer questions about the Orders API. Use only the passage in the prompt.",
messages=[{"role": "user", "content": prompt}],
)
return "".join(block.text for block in response.content if block.type == "text")
def main():
chunks = chunk(DOCS)
while True:
question = input("You: ").strip()
if question.lower() in {"", "quit", "exit"}:
break
score, passage = retrieve(question, chunks)
if score < MIN_SCORE:
print("Bot: The Orders API docs do not cover this.")
continue
print(f"Bot: {answer(build_prompt(question, passage))}")
if __name__ == "__main__":
main()- Credential management: The key is read from the
ANTHROPIC_API_KEYenvironment variable, so it never appears in the source code or in version control. - Grounding instruction: The system prompt and the prompt template both restrict the model to information in the retrieved passage.
- Interactive conversation:
input()replaces the fixed question list, and an empty line orquitends the session.
Illustrative output (the model's wording varies between runs):
You: What happens if I go over the rate limit?
Bot: Each API key may send 100 requests per minute. A request over that limit returns error 429 with a Retry-After header.
You: quitCommon Errors
ModuleNotFoundError: No module named 'anthropic': The virtual environment is inactive or the package is not installed, so activate it withsource .venv/bin/activateand runpip install anthropic==1.0.0.KeyError: 'ANTHROPIC_API_KEY': The environment variable is not defined in the current terminal session, so export it again in that session before executing the script.anthropic.AuthenticationError: The API returned status 401 because the key is invalid or revoked, so generate a replacement key in the Anthropic Console and export it.anthropic.NotFoundError: The model identifier is incorrect or unavailable to the account, so compare it with the current model list in the Anthropic documentation.anthropic.RateLimitError: The account exceeded its request limit and received status 429; the SDK retries automatically, and persistent failures require a delay between questions.- Every question returns "do not cover this": The similarity threshold is too strict for the documentation, so print the retrieval scores, then lower
MIN_SCOREor confirm that theSTOPset does not remove important terms.
Next Steps
- Quality measurement: Build a labelled test set and compute hit rate, MRR and faithfulness with RAG evaluation metrics.
- Ranking improvement: Retrieve the top five chunks and add reranking in RAG before building the prompt.
- Retrieval improvement: Combine keyword and vector scores with hybrid search so identifiers such as
429still match. - Autonomous retrieval: Let an agent decide when to search again, as described in agentic RAG.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In the RAG chatbot in Python, what happens when the top retrieval score is below MIN_SCORE?
Frequently Asked Questions
Can you build RAG without LangChain in Python?
A RAG chatbot needs only a chunker, a retriever, a prompt template and a model call, which plain Python and one model SDK can provide. Frameworks add connectors and utilities but are not required.
Which vector database should a Python RAG chatbot use?
A small documentation set can be searched in memory, as in the example code. For thousands of chunks or search by meaning, a vector database or a search engine with vector support is the usual choice.
How does a RAG chatbot avoid answering questions outside its documents?
It sets a minimum retrieval score and refuses when no chunk reaches it. The prompt also instructs the model to answer only from the passage and to say when the passage does not contain the answer.
How should the API key be stored for a RAG chatbot?
The key belongs in an environment variable or a secrets manager, never in the source code. The script reads it at start-up, so the same code runs in development and production with different keys.
Related Articles
- What is RAG (Retrieval-Augmented Generation)Retrieval-augmented generation (RAG) explained: how it retrieves document passages to ground LLM answers, its key steps, a Python example and its limits.
- How RAG WorksHow RAG works in detail: the offline indexing pipeline, the online query pipeline, a traced Python example over API docs, and where each stage can fail.
- Chunking Strategies in RAGChunking strategies in RAG explained: fixed-size, heading-based and semantic chunking, how to choose chunk size and overlap, with a Python comparison.
- Reranking in RAGReranking in RAG explained: how a second, stricter scorer reorders retrieved passages, cross-encoders vs bi-encoders, a Python example and the trade-offs.
- RAG Evaluation MetricsRAG evaluation metrics explained: hit rate, recall@k, MRR, faithfulness and answer relevance, with a Python example on a labelled test set and its limits.
- Hybrid Search (BM25 + Vector)Hybrid search explained: how BM25 keyword results and vector results are merged with reciprocal rank fusion, with a runnable Python runbook search example.