…
Skip to content
Topics
On this page

Chunking Strategies in RAG

Chunking strategies are the rules a retrieval system uses to split documents into smaller passages, called chunks, before they are embedded and indexed. The chosen strategy determines chunk boundaries, size and overlap, and therefore which text a RAG system can retrieve for a given question.

  • Chunk: A contiguous passage of a document that is embedded and retrieved as a single unit.
  • Size: The maximum length of a chunk, usually measured in tokens rather than characters.
  • Overlap: The number of tokens repeated between consecutive chunks, so boundary sentences are not separated from their context.
  • Boundary: The position where one chunk ends and the next begins, chosen by length, structure or meaning.
  • Metadata: Supplementary fields stored with each chunk, such as the source file, section heading and version.
Fixed-size chunking compared with heading-based chunkingThe Orders API Markdown reference is shown as five stacked sections: Authentication, Rate limits, Pagination, Webhooks and Error codes. On the left, six fixed-size chunks with overlap cut across section boundaries, so one chunk can hold the end of one section and the start of the next. On the right, five heading-based chunks each match exactly one section. Chunk positions are illustrative.orders-api.mdFixed-size + overlapHeading-basedAuthenticationRate limitsPaginationWebhooksError codes
Fixed-size chunking compared with heading-based chunking

For example, a documentation assistant for the Orders API can split its Markdown reference into one chunk per section (Authentication, Rate limits, Pagination, Webhooks, Error codes) instead of cutting it every 30 words.

Common Chunking Strategies

StrategyBoundary ruleTypical use
Fixed-sizeEvery N tokens, with overlapUniform plain text, quick baselines
RecursiveParagraphs, then sentences, then words, until under the size limitGeneral prose of mixed length
Heading-basedMarkdown or HTML headingsAPI references, manuals, wikis
SemanticWhere embedding similarity between sentences dropsLong narrative documents without structure
Code-awareFunctions, classes or whole code blocksSource code and SDK examples

Key Characteristics of Chunking Strategies

  • Granularity: Smaller chunks produce more precise embeddings, whereas larger chunks preserve more surrounding context.
  • Structure: Structure-aware strategies respect headings, lists and code blocks instead of splitting at arbitrary positions.
  • Redundancy: Overlap increases index size, because repeated tokens are embedded and stored more than once.
  • Self-containment: An effective chunk can be understood without its neighbours, often by prefixing the section path.
  • Determinism: Rule-based strategies always produce identical chunks for identical input, which simplifies incremental re-indexing.

How Chunking Strategies Work

  1. Parsing: The document is converted into text while its structural markers, such as headings and code fences, are preserved.
  2. Segmentation: The text is divided at candidate boundaries defined by the strategy, such as headings, paragraphs or a fixed token count.
  3. Normalisation: Segments longer than the limit are split again, and very short neighbouring segments are merged.
  4. Overlap: Consecutive chunks receive a shared window of tokens at their boundary. A common starting point is 200 to 500 tokens per chunk with 10 to 20 percent overlap.
  5. Enrichment: Each chunk is prefixed with its document title and section path, and receives metadata for filtering.
  6. Embedding: The finished chunks are embedded and written to a vector database.

Example: Fixed-Size vs Heading-Based Chunking in Python

The code below chunks the same Markdown Orders API reference twice, once every 30 words with an 8-word overlap and once by section heading.

Python
DOC = """# Orders API
## Authentication
Send your API key in the Authorization header as a Bearer token. Keys are created in the dashboard and can be revoked at any time.
## Rate limits
Each key may make 100 calls per minute. Extra calls return status 429 with a Retry-After header. Wait for the number of seconds it gives before retrying.
## Pagination
GET /orders returns 50 items per page. Pass the next_cursor value from the response as the cursor parameter to fetch the next page.
## Webhooks
A POST request is sent to your URL when an order changes status. Verify the X-Signature header before trusting the payload.
## Error codes
401 means the key is missing or invalid. 404 means the order was not found. 429 means the rate limit was exceeded."""

def fixed_size(text, size=30, overlap=8):
    words = text.split()
    return [words[i:i + size] for i in range(0, len(words) - overlap, size - overlap)]

def by_heading(text, size=30, overlap=8):
    chunks, title, section = [], "", ""
    for line in text.splitlines():
        if line.startswith("# "):
            title = line[2:]
        elif line.startswith("## "):
            section = line[3:]
        else:
            words = line.split()  # long sections are split again, with overlap
            for i in range(0, max(1, len(words) - overlap), size - overlap):
                chunks.append([f"[{title} > {section}]"] + words[i:i + size])
    return chunks

for name, chunks in [("Fixed-size", fixed_size(DOC)), ("Heading-based", by_heading(DOC))]:
    print(f"{name}: {len(chunks)} chunks")
    for n, chunk in enumerate(chunks, 1):
        w = " ".join(chunk).split()
        print(f"  {n}. {len(w):>2} words | {' '.join(w[:5])} ... {' '.join(w[-3:])}")

Output:

Example
Fixed-size: 6 chunks
  1. 30 words | # Orders API ## Authentication ... at any time.
  2. 30 words | dashboard and can be revoked ... header. Wait for
  3. 30 words | status 429 with a Retry-After ... next_cursor value from
  4. 30 words | items per page. Pass the ... your URL when
  5. 30 words | POST request is sent to ... is missing or
  6. 23 words | codes 401 means the key ... limit was exceeded.
Heading-based: 5 chunks
  1. 29 words | [Orders API > Authentication] Send ... at any time.
  2. 32 words | [Orders API > Rate limits] ... gives before retrying.
  3. 27 words | [Orders API > Pagination] GET ... the next page.
  4. 25 words | [Orders API > Webhooks] A ... trusting the payload.
  5. 27 words | [Orders API > Error codes] ... limit was exceeded.
  • Misalignment: Fixed-size chunk 3 begins with "status 429" and ends inside Pagination, so it mixes two topics and omits the 100 calls per minute rule.
  • Alignment: Each heading-based chunk covers exactly one topic, which yields a focused embedding and a clean citation.
  • Breadcrumbs: The four-word breadcrumb makes a chunk self-contained, which is why heading chunk 2 slightly exceeds 30 words.

Applications of Chunking Strategies

  • API references: Splitting endpoint documentation by heading, with request examples kept intact.
  • Runbooks: Keeping each numbered remediation procedure inside a single chunk.
  • Source code: Splitting repositories by function or class for code search assistants.
  • Specifications: Chunking long technical specifications by numbered clause.
  • Support articles: Splitting help centre pages into question and answer pairs.

Advantages

  • Precision: Focused chunks match specific questions more closely than whole documents.
  • Efficiency: Only relevant passages consume space in the context window.
  • Citations: Answers can reference an exact section rather than an entire page.
  • Maintenance: A modified section requires re-embedding only its own chunks, not the entire document.

Limitations

  • Fragmentation: Pronouns and references such as "this header" can lose their meaning once separated from earlier text.
  • Storage: Overlap and small chunk sizes multiply the number of stored vectors.
  • Uneven structure: Documents without headings require fallback rules, such as recursive or semantic splitting.
  • Tuning effort: The optimal size varies by corpus and must be measured with RAG evaluation metrics.
  • Ranking noise: Many similar chunks can crowd the results, which reranking in RAG helps to correct.

Chunking is one stage of the indexing pipeline described in how RAG works.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In the Python example, why does fixed-size chunk 3 answer a rate-limit question poorly?

Frequently Asked Questions

What is chunking in RAG?

Chunking, also called document chunking, is the step that splits long files into smaller passages before they are embedded and indexed. Retrieval then returns individual chunks instead of whole documents, so the prompt contains only the relevant part.

How do you choose a RAG chunk size?

There is no universal chunk size. Short chunks give precise matches but lose context, and long chunks keep context but dilute the embedding. Teams start with a moderate size and tune it against a set of real questions.

What is chunk overlap and why is it used?

Overlap repeats the last part of one chunk at the start of the next. A sentence that falls on a boundary then appears whole in at least one chunk, which reduces lost context.

What is semantic chunking?

Semantic chunking places boundaries where the topic changes, measured by the embedding similarity between neighbouring sentences. It produces coherent chunks but costs more to compute than splitting by length or headings.

Should code blocks be split across chunks?

Code blocks are usually kept whole, because half of a function or request example is rarely useful on its own. Structure-aware chunkers treat each code block as an unbreakable unit.