Skip to main content

Text Chunking

Calculating read time…

Chunking strategies for RAG determine how a document gets split into smaller, retrievable pieces — and the strategy you pick (fixed-size, recursive, semantic, hierarchical, contextual, or one of several others) directly controls whether your retrieval system finds complete, useful answers or fragmented, useless ones. Here's every major strategy explained from first principles.

In our last deep dive, we learned that document parsing turns messy real-world files into clean, structured text. That's step one solved. But clean text alone still isn't ready for RAG — because that clean text might be 40 pages long, and no retrieval system searches "whole documents." It searches small pieces.

That's where chunking comes in — and it turns out to be one of the most underrated, highest-leverage decisions in the entire RAG stack. Pick the wrong chunking strategy, and even perfectly parsed, perfectly embedded text will retrieve poorly. Pick the right one, and mediocre embeddings suddenly perform like great ones.

This post assumes zero prior chunking knowledge. We'll build every strategy from first principles — fixed-size, recursive, sentence-based, content-aware, semantic, sliding window, hierarchical, contextual, and agentic chunking — with analogies, diagrams, code, and a clear decision framework for choosing between them.

Chunking strategies for RAG — fixed-size, recursive, semantic, and hierarchical chunking comparison diagram

✂️ Chunking decides what unit of text ever gets a chance to be retrieved
📏 Chunk size is one of the single most impactful, most under-tuned hyperparameters in RAG
🧠 Teams have moved well beyond naive fixed-size splitting toward semantic and contextual chunking
🔗 Anthropic's Contextual Retrieval technique alone has been shown to cut retrieval failure rates significantly by fixing exactly the problem this post explains
⚖️ There is no single "best" chunking strategy — only the best strategy for your document type and your questions.

Let's build the full mental model, one layer at a time.

🧠 Section 0: Quick Recap — Why Do We Even Need to Chunk?

Quick refresher from our parsing post: RAG works like an open-book exam. The system retrieves relevant pages from your documents, then hands them to an LLM to answer from. But there are two hard limits that force us to break documents into pieces before any of that can happen:

1️⃣ LLM Context Window Limits
Even models with huge context windows have a ceiling and a cost. You can't stuff every document into every prompt — you need to send only the relevant slice.
2️⃣ Embedding Granularity
An embedding model converts text into one single vector per piece of text. Embed an entire 40-page document as one vector, and it becomes a vague blur of every topic in it — useless for finding a specific fact.
💡 The Photograph Analogy

Imagine trying to describe an entire city using one single word. You couldn't — it would have to be so vague ("busy", "big") that it's useless for finding anything specific. Now imagine describing one street corner, one building, one shop — each description becomes specific and searchable. Chunking is choosing how big each "photograph" of your document should be — small enough to be specific and precise, big enough to still contain a complete idea.

Chunking, formally: splitting a long piece of parsed text into smaller, semantically meaningful segments — each of which becomes one row in your vector database, gets its own embedding, and can be retrieved independently.


🗺️ Section 1: Where Chunking Sits in the RAG Pipeline

🏗️ RAG Ingestion Pipeline — Where Chunking Fits

🔍 Document Parsing (previous post — clean Markdown/JSON output)
⬇️
✂️ Chunking Layer (today's topic)
(Splitting clean text into small, retrievable, meaningful pieces)
⬇️
🧬
Embedding
🗂️
Vector Index
⬇️
🤖 Retrieval → Reranking → LLM Generation
✅ Why This Matters: A chunk is the smallest unit your system can ever retrieve. If the right fact is split across two chunks and only one gets retrieved, the answer will be incomplete — no matter how good your embedding model, vector database, or LLM is. Chunking quality is a ceiling on everything after it.

⚖️ Section 2: The Core Trade-Off — Chunk Size Is a Balancing Act

Before learning individual strategies, you need to deeply understand the one trade-off that every chunking decision revolves around.

🔬
Chunks Too Small

Highly precise, but loses context. A chunk that says "The rate is 4.5%" with no surrounding sentence leaves you asking: 4.5% of what? Interest rate? Tax? Discount? The fact is retrievable but meaningless alone.

📦
Chunks Too Large

Full of context, but the embedding becomes a vague average of multiple unrelated ideas (back to our "describe a whole city in one word" problem). Retrieval becomes fuzzy, and you waste LLM context budget on irrelevant surrounding text.

📐 Two Numbers You'll Configure in Every Chunker

Chunk size — how much text (usually measured in tokens, roughly ¾ of a word each) goes into one chunk. Common starting points are 256–512 tokens for general Q&A, larger for summarization-style tasks.
Chunk overlap — how much text from the end of one chunk is repeated at the start of the next. This prevents a sentence that spans a chunk boundary from being cut in half and losing meaning in both pieces.

1️⃣ Section 3: Fixed-Size Chunking — The Simplest Starting Point

What it does: Cut the text every N characters or tokens, with a fixed overlap, completely ignoring sentence or paragraph boundaries. It's a ruler, not a reader — it doesn't "understand" the text at all.

📍 Visual Example (chunk size = 40 characters, overlap = 10)

Original: "The refund policy allows returns within 30 days of purchase for unused items only."

Chunk 1: "The refund policy allows returns within 3"
Chunk 2: "hin 30 days of purchase for unused items "
Chunk 3: "ms only."
🚫 The Obvious Problem: Chunk 1 cuts off mid-word ("3"), and no chunk contains the complete, meaningful sentence. This is why plain fixed-size chunking by raw character count is rarely used directly in production anymore — but it's the foundational idea every other strategy improves upon.
✅ When to still use it: Extremely uniform, structure-less text (like raw sensor logs) where sentence boundaries don't exist or don't matter.

2️⃣ Section 4: Recursive Chunking — The Practical Default

What it does: Instead of cutting blindly by character count, recursive chunking tries a list of separators, from biggest to smallest — attempting to split on natural boundaries first, and only falling back to smaller boundaries when a piece is still too big.

🔁 The Recursive Splitting Order (typical default)

1st try: Split on double newline ("\n\n") — i.e., paragraph breaks. If a resulting paragraph is still under the chunk size limit, keep it whole.
↓ still too big?
2nd try: Split on single newline ("\n") — i.e., line breaks.
↓ still too big?
3rd try: Split on sentence-ending punctuation (". ")
↓ still too big?
Last resort: Split on plain whitespace (" ") — this is the fallback, only reached if a single "paragraph" has no punctuation at all (rare, but happens with things like long unbroken code or URLs).

This is called "recursive" because the algorithm keeps recursing (retrying) with the next separator down the list, on any piece that's still too large, until every final piece fits within the target chunk size.

✅ Why it's the practical default: It respects natural language structure whenever possible (paragraphs and sentences stay intact) but still guarantees every chunk fits your size limit, unlike naive fixed-size splitting. This is the default chunker in most popular RAG frameworks (e.g. LangChain's RecursiveCharacterTextSplitter) precisely because it's a strong, reliable baseline for almost any plain-text document.

3️⃣ Section 5: Sentence & Paragraph-Based Chunking

What it does: Instead of relying on character-count separators, this approach first uses a proper sentence-boundary detector (an NLP tool that knows "Mr. Smith" isn't the end of a sentence, even though it has a period) to split text into complete sentences, then groups consecutive sentences together until a target size is reached.

📍 Why This Beats Naive Splitting

A naive splitter on periods alone would incorrectly cut "Dr. Sarah Lee, M.D., joined the board." into three broken fragments at "Dr.", "M.D.,", and the real sentence end. A real sentence-boundary tool (like spaCy's sentence segmenter or NLTK's punkt) correctly recognizes abbreviations and keeps the sentence whole.

✅ Best for: Prose-heavy documents like articles, books, and reports where individual sentences carry complete, standalone ideas.

4️⃣ Section 6: Content-Aware (Structure-Aware) Chunking

What it does: Uses the actual document structure — headings, sections, bullet lists, table boundaries — as the split points, instead of generic text separators. This is only possible because our parsing stage (previous post) preserved real Markdown headings and structure in the first place!

🏗️ How It Works, Step by Step

1. Parse the document's heading hierarchy — e.g. # Chapter 1, ## Section 1.1, ### Subsection 1.1.1.
↓
2. Split at each heading boundary — everything under "Section 1.1" until the next heading becomes one candidate chunk.
↓
3. Attach the heading path as metadata — each chunk remembers it belongs to "Chapter 1 > Section 1.1", even if that heading text isn't repeated inside the chunk body.
↓
4. If a section is still too long, fall back to recursive chunking (Section 4) within that section — best of both worlds.
🚫 What Goes Wrong Without It: Imagine a legal contract with "Section 4: Termination" and "Section 5: Liability" back to back. Generic recursive chunking might place the last sentence of Section 4 and the first sentence of Section 5 in the same chunk — mixing two legally distinct clauses into one retrievable unit, which can genuinely mislead an answer about liability.
✅ Best for: Technical documentation, legal contracts, textbooks, and any document with a clear, meaningful heading hierarchy — which is most well-formatted business and technical content.

5️⃣ Section 7: Semantic Chunking — Splitting by Meaning, Not Structure

What it does: Uses the embedding model itself (the same one used for retrieval — recall from our RAG primer that embeddings turn text into meaning-vectors) to detect where the topic actually changes, and splits there — even if there's no heading or paragraph break at that point.

🧬 How Semantic Chunking Actually Works

1. Split into sentences first (using Section 5's technique).
↓
2. Embed each sentence individually — turning every sentence into its own meaning-vector.
↓
3. Compare each sentence's vector to its neighbor's vector — calculating how "semantically similar" consecutive sentences are.
↓
4. Where similarity drops sharply — meaning the topic just shifted — insert a chunk boundary right there.
💡 A Concrete Example

Imagine a single paragraph (no line break) that starts by discussing "vacation policy" and then, mid-paragraph, shifts to discussing "sick leave policy" without a new paragraph marker. Structure-aware chunking (Section 6) would miss this shift entirely, since there's no heading. Semantic chunking would detect the meaning-vector "jump" between those sentences and correctly split the two topics into separate chunks.
✅ Best for: Unstructured or loosely structured text — meeting transcripts, chat logs, free-flowing narrative content — where topics shift without any formatting cue at all.
🚫 The Trade-Off: Semantic chunking requires embedding every sentence during ingestion (not just once per chunk), which costs more compute and time than the simpler strategies above. Reserve it for content where it genuinely earns its cost.

6️⃣ Section 8: Sliding Window Chunking — Overlap as a First-Class Strategy

We touched on "overlap" in Section 2, but it's worth calling out sliding window chunking as its own named approach, since some pipelines treat overlap as the primary design choice rather than an afterthought.

📍 Visual Example (window = 3 sentences, step = 1 sentence)

Sentences: [S1] [S2] [S3] [S4] [S5]

Chunk 1: S1, S2, S3
Chunk 2: S2, S3, S4   ← window "slides" forward by 1 sentence
Chunk 3: S3, S4, S5

Each chunk shares sentences with its neighbors. This means any single idea spanning 2–3 sentences is guaranteed to appear intact in at least one full chunk, even if it happens to sit right at what would otherwise be a boundary. The cost is redundancy — you store more total text, and near-duplicate chunks appear in your index.

✅ Best for: High-recall use cases where missing a fact is far more costly than a bit of index redundancy and storage overhead — e.g., medical or legal search where completeness matters more than efficiency.

7️⃣ Section 9: Hierarchical Chunking — Small-to-Big Retrieval

What it does: Instead of picking one single chunk size, this strategy creates chunks at two levels: small "child" chunks (great for precise embedding matches) that each link back to a larger "parent" chunk (great for full context) — and retrieves the parent once its child is matched.

🌳 How Small-to-Big Retrieval Works

1. Create small child chunks (e.g. 100–200 tokens each) — these are precise and embed cleanly, matching narrow queries well.
↓
2. Group several children under one larger parent chunk (e.g. 1,000+ tokens — perhaps the entire section they came from).
↓
3. Embed and search only the small children — this is what gets compared against the user's question for matching.
↓
4. When a child chunk matches, hand its parent chunk to the LLM — so the model sees the full surrounding context, not just the narrow matched fragment.
💡 Why This Solves the Section 2 Trade-Off Directly

Recall the core dilemma: small chunks are precise but lose context; big chunks keep context but embed vaguely. Hierarchical chunking sidesteps the trade-off entirely — you get the precision of small-chunk matching AND the context of large-chunk generation, because matching and generation are allowed to use different-sized text.
✅ Best for: Most serious production RAG systems eventually adopt some form of this — it's widely considered a best practice once you've outgrown a simple single-size chunker.

8️⃣ Section 10: Contextual Chunking — The Frontier

Here's a subtle problem every strategy above still has: once you cut a chunk out of its document, it loses the surrounding context that explained what it was about. A chunk that says "The company reported a 12% increase" doesn't say which company, or which quarter — that information was in a sentence three paragraphs earlier, now in a totally different chunk.

🧩 Contextual Retrieval (Anthropic's Approach) — How It Works

1. Chunk the document normally using any strategy above (recursive, structure-aware, etc.).
↓
2. For each chunk, ask an LLM to write a short context note — given the whole document plus this specific chunk, generate one or two sentences explaining what this chunk is about and where it sits (e.g. "This chunk is from Acme Corp's Q3 2026 earnings report, discussing revenue growth in the cloud division.").
↓
3. Prepend that context note to the chunk before embedding it. Now the embedding captures both the local detail and the document-level context it was missing.
✅ Why This Is a Big Deal: This single technique directly fixes the "chunk lost its context" problem that every earlier strategy in this post still suffers from, without needing bigger chunks (which would reintroduce the vagueness problem from Section 2). It costs extra LLM calls during ingestion (one per chunk, done once), but the retrieval accuracy gains are substantial — this is genuinely one of the most impactful chunking-adjacent techniques to reach mainstream production RAG stacks.

A closely related idea is "late chunking" — instead of chunking first and embedding each chunk separately, you embed the entire document in one pass (using a long-context embedding model), and only afterward slice that rich, context-aware representation into chunks. The chunking happens "late," after the model has already seen the whole document — so every resulting chunk's vector still carries a trace of the full document's context.


9️⃣ Section 11: Agentic (LLM-Based) Chunking

What it does: Instead of using any fixed rule or algorithm, you hand the entire document to an LLM and directly ask it to decide the chunk boundaries — essentially treating chunking as a reasoning task, not a mechanical splitting task.

📌 Example Prompt Pattern

"Here is a document. Break it into self-contained chunks, where each chunk represents one complete idea or fact that could be understood on its own without needing the rest of the document. Return each chunk with a one-line summary."

✅ Best for: Small, high-value document sets (like a legal playbook or key policy documents) where retrieval quality matters enormously and you can afford the extra LLM cost per document.
🚫 The Trade-Off: This is the slowest and most expensive strategy in this entire post, since it requires an LLM call per document (or per section) purely to decide boundaries. It doesn't scale economically to millions of pages — reserve it for your highest-value, lowest-volume content.

🧩 Section 12: Special Cases — Chunking Code, Tables & Q&A Pairs

Not all content is prose. A few content types need their own chunking rules entirely.

💻 Code

Split along function or class boundaries, never mid-function — a code-aware splitter understands syntax (like matching braces and indentation) to keep each function whole and syntactically valid.

📊 Tables

Recall from our parsing post: tables should already be Markdown-formatted. When chunking, keep the header row attached to every row-group chunk (repeat it if the table must be split across multiple chunks) — otherwise a chunk of table rows with no header becomes meaningless numbers.

❓ FAQ / Q&A Content

Each question-answer pair should be its own chunk — never split a question from its answer, and never merge two unrelated Q&A pairs into one chunk.


🧭 Section 13: Which Strategy Should You Actually Use?

Strategy Best For Cost
Fixed-size Uniform, structureless text Lowest
Recursive General-purpose default Low
Sentence-based Prose-heavy articles, books Low
Content-aware Technical docs, contracts, textbooks Low–Medium
Semantic Transcripts, chat logs, loose narrative Medium
Sliding window High-recall, medical/legal search Medium (storage)
Hierarchical (parent-child) Most serious production systems Medium
Contextual / late chunking High-stakes accuracy, best practice Medium–High
Agentic (LLM-based) Small, high-value document sets Highest
💡 A Practical Starting Recipe

Start with recursive chunking as your baseline. Layer content-aware splitting on top if your documents have real headings. Add hierarchical parent-child retrieval once you notice answers missing context. Add contextual chunking once you've measured (Section 15) that chunks are losing document-level context. Reach for semantic or agentic chunking only for the specific document types that need it.

💻 Section 14: Let's Build a Recursive + Content-Aware Chunker (With Code)

📌 What This Code Does (Read Before The Code!)

This combines Section 6 (content-aware splitting by heading) with Section 4 (recursive splitting as a fallback). It first tries to keep each document section whole; if a section is still too big, it recursively falls back to smaller separators — exactly the "best of both worlds" approach mentioned earlier.

# Content-Aware Chunker with Recursive Fallback (Pseudocode)

def chunk_document(markdown_text, max_chunk_size=500, overlap=50):

    # Step 1: Split the document at heading boundaries first
    # e.g. "## Section 1.1" starts a new candidate chunk
    sections = split_by_markdown_headings(markdown_text)

    final_chunks = []

    for section in sections:

        # Step 2: Does this whole section already fit within our limit?
        if token_count(section.text) <= max_chunk_size:
            # Keep it whole — no need to split further
            final_chunks.append({
                "text":         section.text,
                "heading_path": section.heading_path,  # e.g. "Ch.1 > Sec 1.1"
            })

        else:
            # Step 3: Section too big — fall back to recursive splitting
            # Try separators from biggest to smallest, in order
            sub_chunks = recursive_split(
                section.text,
                separators=["\n\n", "\n", ". ", " "],
                chunk_size=max_chunk_size,
                chunk_overlap=overlap
            )
            for sub in sub_chunks:
                final_chunks.append({
                    "text":         sub,
                    "heading_path": section.heading_path,  # still remembers its parent section
                })

    return final_chunks


# ─────────────────────────────────────────────
# The recursive fallback itself (Section 4's algorithm)

def recursive_split(text, separators, chunk_size, chunk_overlap):
    if not separators:
        # No separators left — hard-cut as a last resort
        return hard_cut_by_length(text, chunk_size, chunk_overlap)

    current_separator = separators[0]
    pieces = text.split(current_separator)

    chunks, buffer = [], ""
    for piece in pieces:
        if token_count(buffer + piece) <= chunk_size:
            buffer += piece + current_separator
        else:
            if buffer:
                chunks.append(buffer)
            # This single piece is still too big on its own —
            # recurse with the NEXT, smaller separator
            if token_count(piece) > chunk_size:
                chunks.extend(recursive_split(piece, separators[1:], chunk_size, chunk_overlap))
                buffer = ""
            else:
                buffer = piece + current_separator

    if buffer:
        chunks.append(buffer)

    return add_overlap(chunks, chunk_overlap)
✅ Notice the Pattern: Just like the parsing router from our last post, this chunker always tries the "better" option first (keep a whole section intact) and only reaches for a more aggressive fallback when it genuinely has to. This "prefer the least destructive option" philosophy shows up again and again in well-designed RAG systems.

🧪 Section 15: How Do You Know If Your Chunking Is Actually Good?

📏 Context Sufficiency Test

Pick 20–30 real chunks at random and ask: "if I only read this chunk, with zero other context, could I understand and act on it?" If the answer is often "no" — you need more context per chunk, or contextual chunking (Section 10).

🎯 Retrieval Recall on a Known-Answer Set

Build a small test set of questions with known correct source passages. Run retrieval and check: does the correct passage actually appear, complete and intact, in one of the top retrieved chunks? If the fact is present but split across two separately-retrieved chunks, that's a chunking problem.

📐 Chunk Size Distribution

Plot the size (in tokens) of every chunk in your index. A huge spike of very tiny chunks or very oversized chunks usually signals a bug in your splitting logic — for example, a table that failed to be recognized and got shattered sentence-by-sentence.


❓ Frequently Asked Questions

What is the best chunking strategy for RAG?

There isn't one universal best strategy — recursive chunking is the safest general-purpose default, content-aware chunking is best for well-structured documents with headings, and hierarchical or contextual chunking are worth adopting once you measure real retrieval gaps. Match the strategy to your document type and query patterns.

What chunk size should I use for RAG?

A common starting point is 256 to 512 tokens for general question answering, but the right number depends on your content and questions. Always validate with a retrieval recall test rather than trusting a default number blindly.

What is contextual chunking?

Contextual chunking prepends a short, LLM-generated summary of where a chunk sits in the overall document before embedding it, so the chunk's vector captures document-level context it would otherwise lose once cut out on its own. It's one of the highest-impact recent techniques for improving retrieval accuracy.

What is hierarchical or parent-child chunking?

It creates small child chunks for precise embedding matches and larger parent chunks for full context, searching against the children but handing the matched parent to the LLM. This resolves the core precision-versus-context trade-off that every other single-size chunking strategy has to compromise on.

How do I know if my chunking strategy is working well?

Run a context sufficiency test on random chunks, measure retrieval recall against a known-answer test set, and check the chunk size distribution for anomalies. Chunking quality should be measured independently, not assumed from a good-looking final answer.


🛡️ Section 16: Common Chunking Pitfalls to Avoid

🔢 Choosing Chunk Size by Guessing, Not Testing

"512 tokens" is a common default, not a law of nature. The right size depends heavily on your document type and the kind of questions users ask. Always validate with the recall test from Section 15.

🧾 Forgetting to Repeat Table Headers

A large table split into multiple chunks needs its header row repeated in every chunk (Section 12) — otherwise you get orphaned rows of numbers with no idea what column they belong to.

🔁 Zero Overlap Between Chunks

Skipping overlap entirely to "save space" often cuts important sentences exactly in half at chunk boundaries — a small amount of overlap (Section 2) is cheap insurance against this.

🧬 Using One Strategy for Every Document Type

Legal contracts, chat transcripts, and API documentation don't behave the same way — mixing document-type-specific rules (Section 12) with a general default (Section 4) usually beats forcing one strategy on everything.


🎓 Section 17: Cheat Sheet — Designing Your Chunking Layer

Step 1: Understand Your Content
  • Is it structured (headings) or unstructured (transcripts, chat)?
  • Does it contain tables, code, or Q&A pairs needing special handling?
  • What's the typical length of a "complete idea" in this content?
Step 2: Pick a Baseline Strategy (Section 13's Table)
  • Start with recursive chunking as the safe default
  • Layer content-aware splitting if real headings exist
  • Reserve semantic/agentic chunking for content that truly needs it
Step 3: Add Overlap & Metadata
  • Set a modest overlap (Section 2) to protect boundary sentences
  • Attach heading path, page number, and source file to every chunk
Step 4: Level Up With Hierarchical & Contextual Chunking
  • Add parent-child retrieval (Section 9) once precision vs. context becomes a real issue
  • Add contextual chunking (Section 10) once you measure chunks losing document-level context
Step 5: Measure, Don't Assume (Section 15)
  • Run the context sufficiency test on random chunks
  • Track retrieval recall against a known-answer test set
  • Watch the chunk size distribution for silent bugs

🎉 Final Summary

✂️ Chunking decides the smallest retrievable unit in your RAG system — get it wrong and nothing downstream can compensate
⚖️ The core trade-off is precision vs. context — small chunks are specific but vague alone, big chunks keep context but embed fuzzily
🔁 Recursive chunking is the reliable, practical default for most plain text
🏗️ Content-aware chunking uses real document structure (headings) to keep ideas intact
🧬 Semantic chunking splits by meaning shift, for content with no formatting cues at all
🌳 Hierarchical (parent-child) chunking resolves the core trade-off directly, and is a widely adopted production best practice
🔗 Contextual chunking is the frontier technique, fixing the "chunk lost its context" problem directly
🧪 Always measure chunking quality independently — context sufficiency, retrieval recall, and chunk size distribution
✅ The Core Lesson:

There is no universal "correct" chunk size or strategy — only the one that fits your documents and your users' actual questions. The engineers who get this right don't pick one clever technique and stop; they start simple (recursive), measure honestly (Section 15), and layer in more sophisticated strategies — structure-aware, hierarchical, contextual — exactly where the data shows they're needed. That disciplined, measured approach beats chasing the fanciest technique every single time.


Happy Building! Chunk Wisely, Retrieve Precisely. 🔥

Comments