…
Skip to content
Topics
On this page

Build a RAG Chatbot in Python

A RAG chatbot in Python is a program that answers questions by retrieving the most relevant passage from a document collection and passing it to a language model with the question. It combines chunking, retrieval, prompt construction and generation in a loop, so every answer is grounded in the supplied documentation.

  • Document loader: Code that reads the source documentation into memory as text.
  • Chunker: A function that splits the documentation into sections that can be retrieved independently.
  • Retriever: A scoring function that compares the question with every chunk and returns the closest match.
  • Prompt construction: A template that combines an instruction, the retrieved passage and the question.
  • Generator: The language model call that writes the answer; Part 1 uses a clearly labelled stand-in.
The RAG chatbot loop in PythonThe Orders API documentation is split into chunks once with chunk(). For each question, retrieve() scores every chunk, the top chunk must reach MIN_SCORE, build_prompt() combines it with the question, and the answer comes from the stand-in template in Part 1 or from Claude in Part 2. The loop then returns to the next question. The flow is illustrative.Orders API docschunk()QuestionScore chunksretrieve()Top chunk>= MIN_SCOREPromptbuild_prompt()Answerstand-in or Claudenext question
The RAG chatbot loop in Python

For example, a developer asks the documentation chatbot how to fetch the next page of orders, and it answers from the Pagination section of the Orders API documentation.

The implementation is divided into two parts. Part 1 (Steps 1 and 2) shows how to build RAG from scratch with the standard library only, and a template quotes the passage in place of the model. Part 2 (Step 3) replaces that stand-in with a real call to Claude through the Anthropic Python SDK.

Prerequisites

  • Python 3.11: Any recent Python 3 release is compatible; the examples were tested with version 3.11.
  • RAG fundamentals: Familiarity with retrieval-augmented generation and the stages described in how RAG works.
  • Terminal familiarity: Basic use of a command-line shell to create directories, execute scripts and define environment variables.
  • API key (Part 2 only): An Anthropic API key, stored in an environment variable and never written into the code.

Setup

Create a project directory. Part 1 requires no additional installation; for Part 2, create a virtual environment and install the pinned SDK version.

Bash
python3.11 --version
mkdir rag-chatbot && cd rag-chatbot

# Part 2 only
python3.11 -m venv .venv
source .venv/bin/activate
pip install anthropic==1.0.0
export ANTHROPIC_API_KEY="paste-your-key-here"

Step 1: Load and Chunk the Orders API Docs

The documentation is stored as one string with a heading per section, and the chunker divides it at every heading. In a production project, the string would be loaded from Markdown files with open() or retrieved from a documentation repository.

Python
# Step 1: the Orders API docs, split into one chunk per "## " heading.
DOCS = """
## Authentication
Send your API key in the Authorization header as a Bearer token. A request without a valid key returns error 401.

## Rate limits
Each API key may send 100 requests per minute. A request over the limit returns error 429 with a Retry-After header.

## Pagination
GET /orders returns 50 orders per page. Pass next_cursor from the response as the cursor parameter to get the next page.

## Webhooks
When an order changes status, the API sends a POST request to your webhook URL. A failed delivery is retried three times.

## Error codes
400 means an invalid request body, 401 means a missing or invalid key, 404 means the order was not found and 429 means too many requests.
"""

def chunk(text):
    chunks = []
    for block in text.split("## ")[1:]:
        title, _, body = block.partition("\n")
        chunks.append({"title": title.strip(), "text": body.strip()})
    return chunks

chunks = chunk(DOCS)
for c in chunks:
    print(f"{c['title']:<15} {len(c['text'].split())} words")

Output:

Example
Authentication  21 words
Rate limits     21 words
Pagination      21 words
Webhooks        22 words
Error codes     26 words
  • Section boundaries: Each heading becomes one chunk, which suits reference documentation with short, self-contained sections, as explained in chunking in RAG.
  • Metadata: The section title is stored alongside the text, so every answer can identify and cite its source section.

Step 2: Build the RAG Chatbot in Python

Save the complete program as rag_chatbot.py. It repeats Step 1 so that it runs independently, then adds scoring, retrieval, prompt construction and the stand-in answer.

  • vector(): Converts text into a word-frequency vector after removing stop words and applying a crude seven-letter stem.
  • cosine(): Calculates the cosine similarity between two vectors, a value between 0 and 1 that measures lexical overlap.
  • retrieve(): Scores every chunk, including its section title, and returns the highest-scoring chunk with its similarity score.
  • build_prompt(): Combines a grounding instruction, the retrieved passage and the question into one augmented prompt.
  • answer(): Represents the language model; in Part 1 it only quotes the passage, so the output is deterministic.
Python
# rag_chatbot.py: a retrieval chatbot over the Orders API docs.
import math
import re
from collections import Counter

DOCS = """
## Authentication
Send your API key in the Authorization header as a Bearer token. A request without a valid key returns error 401.

## Rate limits
Each API key may send 100 requests per minute. A request over the limit returns error 429 with a Retry-After header.

## Pagination
GET /orders returns 50 orders per page. Pass next_cursor from the response as the cursor parameter to get the next page.

## Webhooks
When an order changes status, the API sends a POST request to your webhook URL. A failed delivery is retried three times.

## Error codes
400 means an invalid request body, 401 means a missing or invalid key, 404 means the order was not found and 429 means too many requests.
"""
STOP = {"a", "an", "the", "is", "to", "i", "do", "how", "my", "if", "of", "what", "does", "as", "from", "with"}
MIN_SCORE = 0.15  # below this, the docs do not cover the question

def chunk(text):
    chunks = []
    for block in text.split("## ")[1:]:
        title, _, body = block.partition("\n")
        chunks.append({"title": title.strip(), "text": body.strip()})
    return chunks

def vector(text):
    # Crude stemming: "authenticate" and "authentication" share "authent".
    return Counter(w.rstrip("s")[:7] for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP)

def cosine(a, b):
    dot = sum(a[w] * b[w] for w in a.keys() & b.keys())
    norm = math.sqrt(sum(v * v for v in a.values())) * math.sqrt(sum(v * v for v in b.values()))
    return dot / norm if norm else 0.0

def retrieve(question, chunks):
    q = vector(question)
    scored = [(cosine(q, vector(c["title"] + " " + c["text"])), c) for c in chunks]
    return max(scored, key=lambda s: s[0])

def build_prompt(question, passage):
    return (f"Answer only from the passage. If it does not answer the question, say so.\n"
            f"Passage ({passage['title']}): {passage['text']}\nQuestion: {question}")

def answer(prompt, passage):
    # Stand-in for the model: quotes the passage instead of generating text.
    return f"[stand-in, no model] From the {passage['title']} docs: \"{passage['text']}\""

def main():
    chunks = chunk(DOCS)
    questions = [  # in a terminal, use: question = input("You: ")
        "How do I authenticate my requests?",
        "What happens if I go over the rate limit?",
        "How do I get the next page of orders?",
        "Does the API support GraphQL?",
    ]
    for turn, question in enumerate(questions, 1):
        score, passage = retrieve(question, chunks)
        print(f"You: {question}")
        if score < MIN_SCORE:
            print(f"Bot: The Orders API docs do not cover this. (top score {score:.2f})\n")
            continue
        prompt = build_prompt(question, passage)
        if turn == 1:
            print(f"--- prompt sent to the model ---\n{prompt}\n---")
        print(f"Bot: {answer(prompt, passage)} (score {score:.2f})\n")

if __name__ == "__main__":
    main()

Output

Example
You: How do I authenticate my requests?
--- prompt sent to the model ---
Answer only from the passage. If it does not answer the question, say so.
Passage (Authentication): Send your API key in the Authorization header as a Bearer token. A request without a valid key returns error 401.
Question: How do I authenticate my requests?
---
Bot: [stand-in, no model] From the Authentication docs: "Send your API key in the Authorization header as a Bearer token. A request without a valid key returns error 401." (score 0.32)

You: What happens if I go over the rate limit?
Bot: [stand-in, no model] From the Rate limits docs: "Each API key may send 100 requests per minute. A request over the limit returns error 429 with a Retry-After header." (score 0.37)

You: How do I get the next page of orders?
Bot: [stand-in, no model] From the Pagination docs: "GET /orders returns 50 orders per page. Pass next_cursor from the response as the cursor parameter to get the next page." (score 0.77)

You: Does the API support GraphQL?
Bot: The Orders API docs do not cover this. (top score 0.13)
  • Retrieval scores: The Pagination chunk scores 0.77 because the question shares "get", "next", "page" and "orders" with it; the other questions score lower but still select the correct section.
  • Refusal threshold: The GraphQL question scores below MIN_SCORE, so the chatbot states that the documentation does not cover it instead of guessing.
  • Placeholder generation: The [stand-in, no model] label marks text produced by a template, not by a language model.
  • Similarity measure: Word-count cosine similarity is a simplified substitute for embeddings stored in a vector database.

Step 3: Replace the Stand-in with a Real LLM Call

Part 2 completes RAG with Claude: it imports the existing retrieval functions and replaces only the answer() function with an authenticated request to the Messages API. Save this as rag_chatbot_llm.py in the same folder, with the virtual environment active and ANTHROPIC_API_KEY exported.

Python
# rag_chatbot_llm.py: Part 2, the same chatbot with a real model call.
import os

import anthropic

from rag_chatbot import DOCS, MIN_SCORE, build_prompt, chunk, retrieve

MODEL = "claude-sonnet-5"  # check the current model name before running
client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])

def answer(prompt):
    response = client.messages.create(
        model=MODEL,
        max_tokens=1024,
        system="You answer questions about the Orders API. Use only the passage in the prompt.",
        messages=[{"role": "user", "content": prompt}],
    )
    return "".join(block.text for block in response.content if block.type == "text")

def main():
    chunks = chunk(DOCS)
    while True:
        question = input("You: ").strip()
        if question.lower() in {"", "quit", "exit"}:
            break
        score, passage = retrieve(question, chunks)
        if score < MIN_SCORE:
            print("Bot: The Orders API docs do not cover this.")
            continue
        print(f"Bot: {answer(build_prompt(question, passage))}")

if __name__ == "__main__":
    main()
  • Credential management: The key is read from the ANTHROPIC_API_KEY environment variable, so it never appears in the source code or in version control.
  • Grounding instruction: The system prompt and the prompt template both restrict the model to information in the retrieved passage.
  • Interactive conversation: input() replaces the fixed question list, and an empty line or quit ends the session.

Illustrative output (the model's wording varies between runs):

Example
You: What happens if I go over the rate limit?
Bot: Each API key may send 100 requests per minute. A request over that limit returns error 429 with a Retry-After header.
You: quit

Common Errors

  • ModuleNotFoundError: No module named 'anthropic': The virtual environment is inactive or the package is not installed, so activate it with source .venv/bin/activate and run pip install anthropic==1.0.0.
  • KeyError: 'ANTHROPIC_API_KEY': The environment variable is not defined in the current terminal session, so export it again in that session before executing the script.
  • anthropic.AuthenticationError: The API returned status 401 because the key is invalid or revoked, so generate a replacement key in the Anthropic Console and export it.
  • anthropic.NotFoundError: The model identifier is incorrect or unavailable to the account, so compare it with the current model list in the Anthropic documentation.
  • anthropic.RateLimitError: The account exceeded its request limit and received status 429; the SDK retries automatically, and persistent failures require a delay between questions.
  • Every question returns "do not cover this": The similarity threshold is too strict for the documentation, so print the retrieval scores, then lower MIN_SCORE or confirm that the STOP set does not remove important terms.

Next Steps

  • Quality measurement: Build a labelled test set and compute hit rate, MRR and faithfulness with RAG evaluation metrics.
  • Ranking improvement: Retrieve the top five chunks and add reranking in RAG before building the prompt.
  • Retrieval improvement: Combine keyword and vector scores with hybrid search so identifiers such as 429 still match.
  • Autonomous retrieval: Let an agent decide when to search again, as described in agentic RAG.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In the RAG chatbot in Python, what happens when the top retrieval score is below MIN_SCORE?

Frequently Asked Questions

Can you build RAG without LangChain in Python?

A RAG chatbot needs only a chunker, a retriever, a prompt template and a model call, which plain Python and one model SDK can provide. Frameworks add connectors and utilities but are not required.

Which vector database should a Python RAG chatbot use?

A small documentation set can be searched in memory, as in the example code. For thousands of chunks or search by meaning, a vector database or a search engine with vector support is the usual choice.

How does a RAG chatbot avoid answering questions outside its documents?

It sets a minimum retrieval score and refuses when no chunk reaches it. The prompt also instructs the model to answer only from the passage and to say when the passage does not contain the answer.

How should the API key be stored for a RAG chatbot?

The key belongs in an environment variable or a secrets manager, never in the source code. The script reads it at start-up, so the same code runs in development and production with different keys.