Chunking strategies for RAG determine how a document gets split into smaller, retrievable pieces — and the strategy you pick (fixed-size, recursive, semantic, hierarchical, contextual, or one of several others) directly controls whether your retrieval system finds complete, useful answers or fragmented, useless ones. Here's every major strategy explained from first principles.
In our last deep dive, we learned that document parsing turns messy real-world files into clean, structured text. That's step one solved. But clean text alone still isn't ready for RAG — because that clean text might be 40 pages long, and no retrieval system searches "whole documents." It searches small pieces.
That's where chunking comes in — and it turns out to be one of the most underrated, highest-leverage decisions in the entire RAG stack. Pick the wrong chunking strategy, and even perfectly parsed, perfectly embedded text will retrieve poorly. Pick the right one, and mediocre embeddings suddenly perform like great ones.
This post assumes zero prior chunking knowledge. We'll build every strategy from first principles — fixed-size, recursive, sentence-based, content-aware, semantic, sliding window, hierarchical, contextual, and agentic chunking — with analogies, diagrams, code, and a clear decision framework for choosing between them.
📏 Chunk size is one of the single most impactful, most under-tuned hyperparameters in RAG
🧠 Teams have moved well beyond naive fixed-size splitting toward semantic and contextual chunking
🔗 Anthropic's Contextual Retrieval technique alone has been shown to cut retrieval failure rates significantly by fixing exactly the problem this post explains
⚖️ There is no single "best" chunking strategy — only the best strategy for your document type and your questions.
Let's build the full mental model, one layer at a time.
- Why Do We Even Need to Chunk?
- Where Chunking Sits in the RAG Pipeline
- The Core Trade-Off: Chunk Size
- 1. Fixed-Size Chunking
- 2. Recursive Chunking
- 3. Sentence & Paragraph-Based Chunking
- 4. Content-Aware (Structure-Aware) Chunking
- 5. Semantic Chunking
- 6. Sliding Window Chunking
- 7. Hierarchical (Parent-Child) Chunking
- 8. Contextual / Late Chunking
- 9. Agentic (LLM-Based) Chunking
- Special Cases: Code, Tables & Q&A
- Which Strategy Should You Use?
- Code: A Recursive + Content-Aware Chunker
- Evaluating Chunking Quality
- FAQ
- Pitfalls
- Cheat Sheet
🧠 Section 0: Quick Recap — Why Do We Even Need to Chunk?
Quick refresher from our parsing post: RAG works like an open-book exam. The system retrieves relevant pages from your documents, then hands them to an LLM to answer from. But there are two hard limits that force us to break documents into pieces before any of that can happen:
Even models with huge context windows have a ceiling and a cost. You can't stuff every document into every prompt — you need to send only the relevant slice.
An embedding model converts text into one single vector per piece of text. Embed an entire 40-page document as one vector, and it becomes a vague blur of every topic in it — useless for finding a specific fact.
Imagine trying to describe an entire city using one single word. You couldn't — it would have to be so vague ("busy", "big") that it's useless for finding anything specific. Now imagine describing one street corner, one building, one shop — each description becomes specific and searchable. Chunking is choosing how big each "photograph" of your document should be — small enough to be specific and precise, big enough to still contain a complete idea.
Chunking, formally: splitting a long piece of parsed text into smaller, semantically meaningful segments — each of which becomes one row in your vector database, gets its own embedding, and can be retrieved independently.
🗺️ Section 1: Where Chunking Sits in the RAG Pipeline
🏗️ RAG Ingestion Pipeline — Where Chunking Fits
(Splitting clean text into small, retrievable, meaningful pieces)
Embedding
Vector Index
⚖️ Section 2: The Core Trade-Off — Chunk Size Is a Balancing Act
Before learning individual strategies, you need to deeply understand the one trade-off that every chunking decision revolves around.
Highly precise, but loses context. A chunk that says "The rate is 4.5%" with no surrounding sentence leaves you asking: 4.5% of what? Interest rate? Tax? Discount? The fact is retrievable but meaningless alone.
Full of context, but the embedding becomes a vague average of multiple unrelated ideas (back to our "describe a whole city in one word" problem). Retrieval becomes fuzzy, and you waste LLM context budget on irrelevant surrounding text.
📐 Two Numbers You'll Configure in Every Chunker
1️⃣ Section 3: Fixed-Size Chunking — The Simplest Starting Point
What it does: Cut the text every N characters or tokens, with a fixed overlap, completely ignoring sentence or paragraph boundaries. It's a ruler, not a reader — it doesn't "understand" the text at all.
📍 Visual Example (chunk size = 40 characters, overlap = 10)
Original: "The refund policy allows returns within 30 days of purchase for unused items only." Chunk 1: "The refund policy allows returns within 3" Chunk 2: "hin 30 days of purchase for unused items " Chunk 3: "ms only."
2️⃣ Section 4: Recursive Chunking — The Practical Default
What it does: Instead of cutting blindly by character count, recursive chunking tries a list of separators, from biggest to smallest — attempting to split on natural boundaries first, and only falling back to smaller boundaries when a piece is still too big.
🔁 The Recursive Splitting Order (typical default)
This is called "recursive" because the algorithm keeps recursing (retrying) with the next separator down the list, on any piece that's still too large, until every final piece fits within the target chunk size.
RecursiveCharacterTextSplitter) precisely because it's a strong,
reliable baseline for almost any plain-text document.
3️⃣ Section 5: Sentence & Paragraph-Based Chunking
What it does: Instead of relying on character-count separators, this approach first uses a proper sentence-boundary detector (an NLP tool that knows "Mr. Smith" isn't the end of a sentence, even though it has a period) to split text into complete sentences, then groups consecutive sentences together until a target size is reached.
📍 Why This Beats Naive Splitting
A naive splitter on periods alone would incorrectly cut "Dr. Sarah Lee, M.D.,
joined the board." into three broken fragments at "Dr.", "M.D.,", and the
real sentence end. A real sentence-boundary tool (like spaCy's sentence
segmenter or NLTK's punkt) correctly recognizes abbreviations and
keeps the sentence whole.
4️⃣ Section 6: Content-Aware (Structure-Aware) Chunking
What it does: Uses the actual document structure — headings, sections, bullet lists, table boundaries — as the split points, instead of generic text separators. This is only possible because our parsing stage (previous post) preserved real Markdown headings and structure in the first place!
🏗️ How It Works, Step by Step
# Chapter 1,
## Section 1.1,
### Subsection 1.1.1.
5️⃣ Section 7: Semantic Chunking — Splitting by Meaning, Not Structure
What it does: Uses the embedding model itself (the same one used for retrieval — recall from our RAG primer that embeddings turn text into meaning-vectors) to detect where the topic actually changes, and splits there — even if there's no heading or paragraph break at that point.
🧬 How Semantic Chunking Actually Works
Imagine a single paragraph (no line break) that starts by discussing "vacation policy" and then, mid-paragraph, shifts to discussing "sick leave policy" without a new paragraph marker. Structure-aware chunking (Section 6) would miss this shift entirely, since there's no heading. Semantic chunking would detect the meaning-vector "jump" between those sentences and correctly split the two topics into separate chunks.
6️⃣ Section 8: Sliding Window Chunking — Overlap as a First-Class Strategy
We touched on "overlap" in Section 2, but it's worth calling out sliding window chunking as its own named approach, since some pipelines treat overlap as the primary design choice rather than an afterthought.
📍 Visual Example (window = 3 sentences, step = 1 sentence)
Sentences: [S1] [S2] [S3] [S4] [S5] Chunk 1: S1, S2, S3 Chunk 2: S2, S3, S4 ← window "slides" forward by 1 sentence Chunk 3: S3, S4, S5
Each chunk shares sentences with its neighbors. This means any single idea spanning 2–3 sentences is guaranteed to appear intact in at least one full chunk, even if it happens to sit right at what would otherwise be a boundary. The cost is redundancy — you store more total text, and near-duplicate chunks appear in your index.
7️⃣ Section 9: Hierarchical Chunking — Small-to-Big Retrieval
What it does: Instead of picking one single chunk size, this strategy creates chunks at two levels: small "child" chunks (great for precise embedding matches) that each link back to a larger "parent" chunk (great for full context) — and retrieves the parent once its child is matched.
🌳 How Small-to-Big Retrieval Works
Recall the core dilemma: small chunks are precise but lose context; big chunks keep context but embed vaguely. Hierarchical chunking sidesteps the trade-off entirely — you get the precision of small-chunk matching AND the context of large-chunk generation, because matching and generation are allowed to use different-sized text.
8️⃣ Section 10: Contextual Chunking — The Frontier
Here's a subtle problem every strategy above still has: once you cut a chunk out of its document, it loses the surrounding context that explained what it was about. A chunk that says "The company reported a 12% increase" doesn't say which company, or which quarter — that information was in a sentence three paragraphs earlier, now in a totally different chunk.
🧩 Contextual Retrieval (Anthropic's Approach) — How It Works
A closely related idea is "late chunking" — instead of chunking first and embedding each chunk separately, you embed the entire document in one pass (using a long-context embedding model), and only afterward slice that rich, context-aware representation into chunks. The chunking happens "late," after the model has already seen the whole document — so every resulting chunk's vector still carries a trace of the full document's context.
9️⃣ Section 11: Agentic (LLM-Based) Chunking
What it does: Instead of using any fixed rule or algorithm, you hand the entire document to an LLM and directly ask it to decide the chunk boundaries — essentially treating chunking as a reasoning task, not a mechanical splitting task.
"Here is a document. Break it into self-contained chunks, where each chunk represents one complete idea or fact that could be understood on its own without needing the rest of the document. Return each chunk with a one-line summary."
🧩 Section 12: Special Cases — Chunking Code, Tables & Q&A Pairs
Not all content is prose. A few content types need their own chunking rules entirely.
Split along function or class boundaries, never mid-function — a code-aware splitter understands syntax (like matching braces and indentation) to keep each function whole and syntactically valid.
Recall from our parsing post: tables should already be Markdown-formatted. When chunking, keep the header row attached to every row-group chunk (repeat it if the table must be split across multiple chunks) — otherwise a chunk of table rows with no header becomes meaningless numbers.
Each question-answer pair should be its own chunk — never split a question from its answer, and never merge two unrelated Q&A pairs into one chunk.
🧭 Section 13: Which Strategy Should You Actually Use?
| Strategy | Best For | Cost |
|---|---|---|
| Fixed-size | Uniform, structureless text | Lowest |
| Recursive | General-purpose default | Low |
| Sentence-based | Prose-heavy articles, books | Low |
| Content-aware | Technical docs, contracts, textbooks | Low–Medium |
| Semantic | Transcripts, chat logs, loose narrative | Medium |
| Sliding window | High-recall, medical/legal search | Medium (storage) |
| Hierarchical (parent-child) | Most serious production systems | Medium |
| Contextual / late chunking | High-stakes accuracy, best practice | Medium–High |
| Agentic (LLM-based) | Small, high-value document sets | Highest |
Start with recursive chunking as your baseline. Layer content-aware splitting on top if your documents have real headings. Add hierarchical parent-child retrieval once you notice answers missing context. Add contextual chunking once you've measured (Section 15) that chunks are losing document-level context. Reach for semantic or agentic chunking only for the specific document types that need it.
💻 Section 14: Let's Build a Recursive + Content-Aware Chunker (With Code)
This combines Section 6 (content-aware splitting by heading) with Section 4 (recursive splitting as a fallback). It first tries to keep each document section whole; if a section is still too big, it recursively falls back to smaller separators — exactly the "best of both worlds" approach mentioned earlier.
# Content-Aware Chunker with Recursive Fallback (Pseudocode) def chunk_document(markdown_text, max_chunk_size=500, overlap=50): # Step 1: Split the document at heading boundaries first # e.g. "## Section 1.1" starts a new candidate chunk sections = split_by_markdown_headings(markdown_text) final_chunks = [] for section in sections: # Step 2: Does this whole section already fit within our limit? if token_count(section.text) <= max_chunk_size: # Keep it whole — no need to split further final_chunks.append({ "text": section.text, "heading_path": section.heading_path, # e.g. "Ch.1 > Sec 1.1" }) else: # Step 3: Section too big — fall back to recursive splitting # Try separators from biggest to smallest, in order sub_chunks = recursive_split( section.text, separators=["\n\n", "\n", ". ", " "], chunk_size=max_chunk_size, chunk_overlap=overlap ) for sub in sub_chunks: final_chunks.append({ "text": sub, "heading_path": section.heading_path, # still remembers its parent section }) return final_chunks # ───────────────────────────────────────────── # The recursive fallback itself (Section 4's algorithm) def recursive_split(text, separators, chunk_size, chunk_overlap): if not separators: # No separators left — hard-cut as a last resort return hard_cut_by_length(text, chunk_size, chunk_overlap) current_separator = separators[0] pieces = text.split(current_separator) chunks, buffer = [], "" for piece in pieces: if token_count(buffer + piece) <= chunk_size: buffer += piece + current_separator else: if buffer: chunks.append(buffer) # This single piece is still too big on its own — # recurse with the NEXT, smaller separator if token_count(piece) > chunk_size: chunks.extend(recursive_split(piece, separators[1:], chunk_size, chunk_overlap)) buffer = "" else: buffer = piece + current_separator if buffer: chunks.append(buffer) return add_overlap(chunks, chunk_overlap)
🧪 Section 15: How Do You Know If Your Chunking Is Actually Good?
Pick 20–30 real chunks at random and ask: "if I only read this chunk, with zero other context, could I understand and act on it?" If the answer is often "no" — you need more context per chunk, or contextual chunking (Section 10).
Build a small test set of questions with known correct source passages. Run retrieval and check: does the correct passage actually appear, complete and intact, in one of the top retrieved chunks? If the fact is present but split across two separately-retrieved chunks, that's a chunking problem.
Plot the size (in tokens) of every chunk in your index. A huge spike of very tiny chunks or very oversized chunks usually signals a bug in your splitting logic — for example, a table that failed to be recognized and got shattered sentence-by-sentence.
❓ Frequently Asked Questions
There isn't one universal best strategy — recursive chunking is the safest general-purpose default, content-aware chunking is best for well-structured documents with headings, and hierarchical or contextual chunking are worth adopting once you measure real retrieval gaps. Match the strategy to your document type and query patterns.
A common starting point is 256 to 512 tokens for general question answering, but the right number depends on your content and questions. Always validate with a retrieval recall test rather than trusting a default number blindly.
Contextual chunking prepends a short, LLM-generated summary of where a chunk sits in the overall document before embedding it, so the chunk's vector captures document-level context it would otherwise lose once cut out on its own. It's one of the highest-impact recent techniques for improving retrieval accuracy.
It creates small child chunks for precise embedding matches and larger parent chunks for full context, searching against the children but handing the matched parent to the LLM. This resolves the core precision-versus-context trade-off that every other single-size chunking strategy has to compromise on.
Run a context sufficiency test on random chunks, measure retrieval recall against a known-answer test set, and check the chunk size distribution for anomalies. Chunking quality should be measured independently, not assumed from a good-looking final answer.
🛡️ Section 16: Common Chunking Pitfalls to Avoid
"512 tokens" is a common default, not a law of nature. The right size depends heavily on your document type and the kind of questions users ask. Always validate with the recall test from Section 15.
A large table split into multiple chunks needs its header row repeated in every chunk (Section 12) — otherwise you get orphaned rows of numbers with no idea what column they belong to.
Skipping overlap entirely to "save space" often cuts important sentences exactly in half at chunk boundaries — a small amount of overlap (Section 2) is cheap insurance against this.
Legal contracts, chat transcripts, and API documentation don't behave the same way — mixing document-type-specific rules (Section 12) with a general default (Section 4) usually beats forcing one strategy on everything.
🎓 Section 17: Cheat Sheet — Designing Your Chunking Layer
- Is it structured (headings) or unstructured (transcripts, chat)?
- Does it contain tables, code, or Q&A pairs needing special handling?
- What's the typical length of a "complete idea" in this content?
- Start with recursive chunking as the safe default
- Layer content-aware splitting if real headings exist
- Reserve semantic/agentic chunking for content that truly needs it
- Set a modest overlap (Section 2) to protect boundary sentences
- Attach heading path, page number, and source file to every chunk
- Add parent-child retrieval (Section 9) once precision vs. context becomes a real issue
- Add contextual chunking (Section 10) once you measure chunks losing document-level context
- Run the context sufficiency test on random chunks
- Track retrieval recall against a known-answer test set
- Watch the chunk size distribution for silent bugs
🎉 Final Summary
There is no universal "correct" chunk size or strategy — only the one that fits your documents and your users' actual questions. The engineers who get this right don't pick one clever technique and stop; they start simple (recursive), measure honestly (Section 15), and layer in more sophisticated strategies — structure-aware, hierarchical, contextual — exactly where the data shows they're needed. That disciplined, measured approach beats chasing the fanciest technique every single time.
Happy Building! Chunk Wisely, Retrieve Precisely. 🔥
Comments
Post a Comment