Chunking Strategies in RAG
Chunking strategies are the rules a retrieval system uses to split documents into smaller passages, called chunks, before they are embedded and indexed. The chosen strategy determines chunk boundaries, size and overlap, and therefore which text a RAG system can retrieve for a given question.
- Chunk: A contiguous passage of a document that is embedded and retrieved as a single unit.
- Size: The maximum length of a chunk, usually measured in tokens rather than characters.
- Overlap: The number of tokens repeated between consecutive chunks, so boundary sentences are not separated from their context.
- Boundary: The position where one chunk ends and the next begins, chosen by length, structure or meaning.
- Metadata: Supplementary fields stored with each chunk, such as the source file, section heading and version.
For example, a documentation assistant for the Orders API can split its Markdown reference into one chunk per section (Authentication, Rate limits, Pagination, Webhooks, Error codes) instead of cutting it every 30 words.
Common Chunking Strategies
| Strategy | Boundary rule | Typical use |
|---|---|---|
| Fixed-size | Every N tokens, with overlap | Uniform plain text, quick baselines |
| Recursive | Paragraphs, then sentences, then words, until under the size limit | General prose of mixed length |
| Heading-based | Markdown or HTML headings | API references, manuals, wikis |
| Semantic | Where embedding similarity between sentences drops | Long narrative documents without structure |
| Code-aware | Functions, classes or whole code blocks | Source code and SDK examples |
Key Characteristics of Chunking Strategies
- Granularity: Smaller chunks produce more precise embeddings, whereas larger chunks preserve more surrounding context.
- Structure: Structure-aware strategies respect headings, lists and code blocks instead of splitting at arbitrary positions.
- Redundancy: Overlap increases index size, because repeated tokens are embedded and stored more than once.
- Self-containment: An effective chunk can be understood without its neighbours, often by prefixing the section path.
- Determinism: Rule-based strategies always produce identical chunks for identical input, which simplifies incremental re-indexing.
How Chunking Strategies Work
- Parsing: The document is converted into text while its structural markers, such as headings and code fences, are preserved.
- Segmentation: The text is divided at candidate boundaries defined by the strategy, such as headings, paragraphs or a fixed token count.
- Normalisation: Segments longer than the limit are split again, and very short neighbouring segments are merged.
- Overlap: Consecutive chunks receive a shared window of tokens at their boundary. A common starting point is 200 to 500 tokens per chunk with 10 to 20 percent overlap.
- Enrichment: Each chunk is prefixed with its document title and section path, and receives metadata for filtering.
- Embedding: The finished chunks are embedded and written to a vector database.
Example: Fixed-Size vs Heading-Based Chunking in Python
The code below chunks the same Markdown Orders API reference twice, once every 30 words with an 8-word overlap and once by section heading.
DOC = """# Orders API
## Authentication
Send your API key in the Authorization header as a Bearer token. Keys are created in the dashboard and can be revoked at any time.
## Rate limits
Each key may make 100 calls per minute. Extra calls return status 429 with a Retry-After header. Wait for the number of seconds it gives before retrying.
## Pagination
GET /orders returns 50 items per page. Pass the next_cursor value from the response as the cursor parameter to fetch the next page.
## Webhooks
A POST request is sent to your URL when an order changes status. Verify the X-Signature header before trusting the payload.
## Error codes
401 means the key is missing or invalid. 404 means the order was not found. 429 means the rate limit was exceeded."""
def fixed_size(text, size=30, overlap=8):
words = text.split()
return [words[i:i + size] for i in range(0, len(words) - overlap, size - overlap)]
def by_heading(text, size=30, overlap=8):
chunks, title, section = [], "", ""
for line in text.splitlines():
if line.startswith("# "):
title = line[2:]
elif line.startswith("## "):
section = line[3:]
else:
words = line.split() # long sections are split again, with overlap
for i in range(0, max(1, len(words) - overlap), size - overlap):
chunks.append([f"[{title} > {section}]"] + words[i:i + size])
return chunks
for name, chunks in [("Fixed-size", fixed_size(DOC)), ("Heading-based", by_heading(DOC))]:
print(f"{name}: {len(chunks)} chunks")
for n, chunk in enumerate(chunks, 1):
w = " ".join(chunk).split()
print(f" {n}. {len(w):>2} words | {' '.join(w[:5])} ... {' '.join(w[-3:])}")Output:
Fixed-size: 6 chunks
1. 30 words | # Orders API ## Authentication ... at any time.
2. 30 words | dashboard and can be revoked ... header. Wait for
3. 30 words | status 429 with a Retry-After ... next_cursor value from
4. 30 words | items per page. Pass the ... your URL when
5. 30 words | POST request is sent to ... is missing or
6. 23 words | codes 401 means the key ... limit was exceeded.
Heading-based: 5 chunks
1. 29 words | [Orders API > Authentication] Send ... at any time.
2. 32 words | [Orders API > Rate limits] ... gives before retrying.
3. 27 words | [Orders API > Pagination] GET ... the next page.
4. 25 words | [Orders API > Webhooks] A ... trusting the payload.
5. 27 words | [Orders API > Error codes] ... limit was exceeded.- Misalignment: Fixed-size chunk 3 begins with "status 429" and ends inside Pagination, so it mixes two topics and omits the 100 calls per minute rule.
- Alignment: Each heading-based chunk covers exactly one topic, which yields a focused embedding and a clean citation.
- Breadcrumbs: The four-word breadcrumb makes a chunk self-contained, which is why heading chunk 2 slightly exceeds 30 words.
Applications of Chunking Strategies
- API references: Splitting endpoint documentation by heading, with request examples kept intact.
- Runbooks: Keeping each numbered remediation procedure inside a single chunk.
- Source code: Splitting repositories by function or class for code search assistants.
- Specifications: Chunking long technical specifications by numbered clause.
- Support articles: Splitting help centre pages into question and answer pairs.
Advantages
- Precision: Focused chunks match specific questions more closely than whole documents.
- Efficiency: Only relevant passages consume space in the context window.
- Citations: Answers can reference an exact section rather than an entire page.
- Maintenance: A modified section requires re-embedding only its own chunks, not the entire document.
Limitations
- Fragmentation: Pronouns and references such as "this header" can lose their meaning once separated from earlier text.
- Storage: Overlap and small chunk sizes multiply the number of stored vectors.
- Uneven structure: Documents without headings require fallback rules, such as recursive or semantic splitting.
- Tuning effort: The optimal size varies by corpus and must be measured with RAG evaluation metrics.
- Ranking noise: Many similar chunks can crowd the results, which reranking in RAG helps to correct.
Chunking is one stage of the indexing pipeline described in how RAG works.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In the Python example, why does fixed-size chunk 3 answer a rate-limit question poorly?
Frequently Asked Questions
What is chunking in RAG?
Chunking, also called document chunking, is the step that splits long files into smaller passages before they are embedded and indexed. Retrieval then returns individual chunks instead of whole documents, so the prompt contains only the relevant part.
How do you choose a RAG chunk size?
There is no universal chunk size. Short chunks give precise matches but lose context, and long chunks keep context but dilute the embedding. Teams start with a moderate size and tune it against a set of real questions.
What is chunk overlap and why is it used?
Overlap repeats the last part of one chunk at the start of the next. A sentence that falls on a boundary then appears whole in at least one chunk, which reduces lost context.
What is semantic chunking?
Semantic chunking places boundaries where the topic changes, measured by the embedding similarity between neighbouring sentences. It produces coherent chunks but costs more to compute than splitting by length or headings.
Should code blocks be split across chunks?
Code blocks are usually kept whole, because half of a function or request example is rarely useful on its own. Structure-aware chunkers treat each code block as an unbreakable unit.
Related Articles
- How RAG WorksHow RAG works in detail: the offline indexing pipeline, the online query pipeline, a traced Python example over API docs, and where each stage can fail.
- Vector Database in RAGVector database in RAG explained: how embeddings are stored, indexed and searched with cosine similarity, ANN indexes, metadata filters and a Python demo.
- Embeddings in LLMLearn what embeddings in LLMs are, how text becomes vectors, how cosine similarity compares meaning and where embeddings are used, with Python code.
- Tokens and Tokenization in LLMLearn how tokenization in LLMs splits text and code into subword tokens and why token counts drive cost and context, with a Python tokenizer example.
- Reranking in RAGReranking in RAG explained: how a second, stricter scorer reorders retrieved passages, cross-encoders vs bi-encoders, a Python example and the trade-offs.
- RAG Evaluation MetricsRAG evaluation metrics explained: hit rate, recall@k, MRR, faithfulness and answer relevance, with a Python example on a labelled test set and its limits.