Skip to main content

Fixed-Size vs Recursive vs Semantic Chunking

Calculating read time…

How should you split documents so that a RAG system can actually find the right information?

One of the first things you discover when building a Retrieval-Augmented Generation (RAG) system is that you cannot simply throw an entire document into the retrieval system and expect good results.

A 100-page policy, technical manual, contract, or support knowledge base contains far more information than most individual questions need.

So we divide the document into smaller pieces called chunks.

But that creates an important question:

How should we decide where one chunk ends and the next chunk begins?

That is where different chunking strategies come into play.

In this article, we will understand three common approaches:

1. Fixed-Size Chunking
“Cut the document into pieces of roughly the same size.”
2. Recursive Chunking
“Try to cut at natural document boundaries before making smaller cuts.”
3. Semantic Chunking
“Try to keep text together when it is talking about the same idea.”

The goal is not to find the fanciest technique.

The goal is to create chunks that give the retrieval system the right amount of information, in the right context, at the right time.

What We Will Learn

1. First, Build the Right Mental Model

Before talking about algorithms, forget RAG for a moment.

Imagine that you own a very large library.

A customer walks in and asks:

“Can I claim a hotel expense when travelling internationally?”

You could hand the customer the entire library.

Technically, the answer is somewhere inside it.

But that is not useful.

A better librarian would quickly locate the relevant section and bring only the useful pages.

📌 This is the basic idea behind chunking.

Instead of asking the AI model to work with an entire document, we create smaller searchable pieces so the retrieval system has a better chance of finding the evidence needed for the question.

The important part is this:

Chunking does not create knowledge. It creates better retrieval units.

2. Why Does Chunking Matter in RAG?

A simplified RAG pipeline looks like this:

Documents
↓
Parsing
↓
Chunking
↓
Embeddings + Index
↓
User Question
↓
Retrieval
↓
Relevant Context
↓
LLM Answer

The retrieval system searches over the chunks you created.

That means the quality of your chunks influences what the retrieval system can find.

Consider this simple example:

Rule: Employees travelling internationally may claim hotel expenses up to the applicable accommodation limit.

Condition: The trip must be approved before the booking is made.

If your chunk contains only the first sentence, the retrieval system may find the hotel rule but miss the approval condition.

The information existed in the document.

The problem was how the information was packaged for retrieval.

3. Fixed-Size Chunking

📌 Imagine cutting a chocolate bar into equal pieces.

You decide how large each piece should be, and then keep cutting at approximately the same interval.

Fixed-size chunking is the simplest approach to understand.

You define a target size and divide the document into chunks around that size.

The size might be measured using:

  • Tokens
  • Characters
  • Words
  • Another defined text unit

For example, you might experiment with chunks of approximately 600–900 tokens.

That number is not a magic setting. It is simply an experimental starting point.

What does fixed-size chunking actually do?

Original document:
A very long stream of text...

↓

Chunk 1: first section of text
Chunk 2: next section of text
Chunk 3: next section of text
Chunk 4: next section of text

Why is it attractive?

  • Very easy to understand.
  • Easy to implement.
  • Predictable chunk sizes.
  • Useful as a baseline for RAG experiments.
  • Easy to process across large document collections.

Where can it go wrong?

The document does not know about your chunk-size setting.

Suppose the document says:

Chunk A:
“International travel must be approved by the employee's manager before booking. Hotel expenses are...”

Chunk B:
“...reimbursable up to the applicable accommodation limit.”

The rule has been split between two chunks.

A retrieval system may retrieve Chunk B because it contains “hotel expenses” but miss the approval requirement sitting just before the boundary.

This is the biggest idea to remember about fixed-size chunking:

Fixed-size chunking is simple, predictable, and sometimes surprisingly effective—but it does not understand the meaning of the document.

4. Recursive Chunking

📌 Think about cutting a book.

You would probably prefer to separate chapters and paragraphs before cutting through the middle of a sentence.

Recursive chunking tries to do something similar.

Instead of immediately cutting text at a fixed position, it tries a hierarchy of possible boundaries.

A simplified version might think like this:

Can this section fit?
↓
Yes → Keep it together.

No → Can its paragraphs fit?
↓
Yes → Split by paragraphs.

No → Can the sentences fit?
↓
Split further.

The exact hierarchy can vary, but the central idea is the same:

Try to preserve natural structure before making smaller cuts.

Why is this useful?

Consider a technical document with:

  • Authentication
  • API Configuration
  • Error Handling
  • Examples

If the Authentication section fits inside your target size, there is little reason to cut it into arbitrary pieces.

If the section is too large, you can progressively break it into paragraphs and smaller units.

This makes recursive chunking a useful bridge between two extremes:

Fixed-size: “Size first.”

Recursive: “Structure first, then size.”

But recursive chunking still has a limitation

It understands structure better than simple fixed-size splitting, but structure is not the same thing as meaning.

Two consecutive paragraphs can belong to completely different topics.

And two paragraphs in different locations can sometimes be closely related.

That leads us to semantic chunking.

5. Semantic Chunking

📌 Imagine organizing your study notes.

You would not necessarily group notes just because they were written next to each other. You would group them because they are about the same subject.

Semantic chunking takes this idea into the RAG pipeline.

Instead of asking only:

“How much text should go into this chunk?”

we also ask:

“Are these pieces of text still talking about the same idea?”

A simple example

Paragraph 1
Employees can submit travel expenses through the expense portal.

Paragraph 2
Receipts are required for hotel and transportation claims.

Paragraph 3
International travel requires additional approval before booking.

The first two paragraphs are closely connected.

The third paragraph introduces a more specific topic: international approval.

A semantic strategy can attempt to detect that change in meaning and use it as a chunk boundary.

How does it know that the topic changed?

At a high level, semantic chunking can represent nearby pieces of text and compare how similar their meanings are.

If neighboring text remains strongly related, it can stay together.

If the relationship changes significantly, that location becomes a candidate boundary.

Semantic chunking tries to make the chunk boundary follow the meaning of the document rather than merely its physical position.

Does that make semantic chunking better?

Not automatically.

This is one of the most important points in this entire article.

Semantic chunking can be useful, but it also introduces more processing, more decisions, and more opportunities for tuning.

If a simple recursive strategy already gives excellent retrieval for your documents, adding semantic processing may not provide enough additional value.

6. One Document, Three Different Chunking Results

Let's make the difference very concrete.

Imagine a travel-policy document containing this simplified content:

Travel Eligibility
Employees travelling for approved business activities may claim eligible expenses.

Hotel Expenses
Hotel expenses are reimbursable within the applicable accommodation limit.

Documentation
A valid hotel invoice must be submitted with the expense claim.

International Travel
International trips require additional approval before booking.

Now imagine the user asks:

“Can I claim my hotel expense for an international business trip, and what do I need to submit?”

Fixed-size approach

The document may be divided according to position. The hotel rule could end up in one chunk while the required documentation appears in another.

Recursive approach

The section headings and paragraphs are more likely to remain intact, making each policy area easier to retrieve as a meaningful unit.

Semantic approach

The system may try to group related information around concepts such as hotel reimbursement, documentation, and international travel.

But even semantic chunking may not put every required fact into one chunk.

And that is perfectly fine.

A good RAG system can retrieve multiple complementary chunks.

Good chunking does not mean “one question = one chunk.”

It means the chunks are useful building blocks for retrieving the evidence needed to answer the question.

7. Fixed-Size vs Recursive vs Semantic

Approach What it cares about Best starting point for Main trade-off
Fixed-Size Chunk length Simple, consistent documents and baselines Can cut through meaning
Recursive Document structure + chunk size Policies, manuals, guides, technical documents Depends on good document structure
Semantic Meaning and topic changes Topic-heavy or irregular documents More processing and tuning

The simplest way to remember the difference:

Fixed: “Cut by size.”

Recursive: “Respect structure, then cut by size.”

Semantic: “Follow changes in meaning.”

8. Which One Should You Start With?

This is where beginners often look for a single answer.

There isn't one.

Instead, look at your documents.

If your documents look like... A reasonable starting point Why
Fairly consistent plain text Fixed-size Simple baseline and easy to test
Policies, manuals, technical guides Recursive / structure-aware Headings and paragraphs carry useful meaning
Documents with frequent topic changes Semantic experimentation Meaning may be more useful than physical position
Mixed enterprise corpus Hybrid approach Different document types may need different treatment

A practical progression

If you are building your first RAG system, you do not need to jump directly into sophisticated semantic chunking.

Start
Build a simple fixed-size baseline.

↓

Improve
Try recursive or structure-aware chunking.

↓

Investigate
If retrieval still struggles because of topic boundaries, test semantic chunking.

↓

Measure
Keep the approach that produces better results for your actual questions and documents.

This progression keeps your RAG system understandable while you learn.

9. How Do You Know Which Chunking Strategy Works?

This is the most important engineering question.

Do not decide based only on how impressive the technique sounds.

Create a small evaluation dataset containing realistic questions.

For each question, know what evidence should be retrieved.

Then test:

  • Fixed-size chunks
  • Recursive chunks
  • Semantic chunks

Keep the rest of the experiment as consistent as possible.

Look at three levels

Level 1 — Retrieval
Did the system retrieve the information needed?

Level 2 — Context
Was the retrieved information complete and understandable?

Level 3 — Answer
Did the model produce an answer supported by that evidence?

This distinction is extremely useful when debugging RAG.

If the correct information never reaches the model, changing the prompt may not solve the underlying problem.

You need to inspect the retrieval pipeline.

Inspect the actual chunks

Do not evaluate only numbers.

Open several generated chunks and read them like a human.

Ask:

  • Does this chunk represent one understandable piece of information?
  • Is the heading still connected to its content?
  • Did a rule become separated from its exception?
  • Did a table get destroyed during splitting?
  • Are unrelated topics mixed together?
  • Would a human understand what this chunk is about without seeing the entire document?

That last question is particularly valuable.

10. What About Chunk Overlap?

Chunking strategy and chunk overlap are related, but they are not the same thing.

Overlap means repeating some content between neighboring chunks.

Chunk 1:
A B C D E F

Chunk 2:
E F G H I J

Here, E and F appear in both chunks.

The purpose is to reduce the chance that important information is lost exactly at a boundary.

But overlap also creates duplication.

More overlap can mean more indexed text, more embeddings, more candidates, and potentially more duplicated context during retrieval.

So overlap should also be treated as something to test—not something to increase automatically whenever retrieval performs poorly.

11. A Very Important Production Lesson

Your chunking strategy is part of your retrieval architecture.

Changing the chunking method changes the actual records being indexed.

That can affect:

  • How many chunks exist.
  • What each embedding represents.
  • Which chunks are retrieved.
  • How much context reaches the model.
  • Storage requirements.
  • Embedding processing.
  • Retrieval latency.
  • Answer quality.

So in a production environment, changing chunking should be treated as a meaningful retrieval-system change rather than a tiny configuration tweak.

12. Common Beginner Mistakes

❌ Mistake 1: “Semantic must be better because it is smarter.”
More complexity does not automatically mean better retrieval.

❌ Mistake 2: “I found a chunk size online, so I will use it everywhere.”
Chunk size depends on your documents and questions.

❌ Mistake 3: “Every question should be answered by one chunk.”
Complex questions often require multiple pieces of evidence.

❌ Mistake 4: “If retrieval is bad, increase overlap.”
The problem may actually be parsing, metadata, embeddings, ranking, or document structure.

❌ Mistake 5: “The final answer is all I need to evaluate.”
Always inspect what was retrieved before blaming the model.

13. Beginner FAQ

Q1. Which chunking method should a beginner learn first?

Start with fixed-size chunking so you understand the basic mechanics. Then learn recursive chunking because it introduces document structure. After that, explore semantic chunking to understand meaning-aware boundaries.

Q2. Is recursive chunking always better than fixed-size chunking?

No. Recursive chunking can preserve document structure better, but the actual benefit depends on your corpus and evaluation results.

Q3. Is semantic chunking the most advanced approach?

It can involve more sophisticated processing, but “more sophisticated” does not mean “best for every application.”

Q4. Can I use different chunking strategies for different documents?

Yes. In an enterprise RAG system, different document families may benefit from different processing rules.

Q5. Can chunking solve hallucinations?

Good chunking can improve the evidence available to the model, but it does not eliminate hallucinations by itself. Grounding also depends on retrieval, context assembly, prompting, model behavior, and evaluation.

Q6. Should I always use overlap?

Not necessarily. Use overlap when it solves a real boundary problem, and measure whether the additional duplication improves retrieval enough to justify the cost.

Q7. What is the biggest lesson?

Think about chunking from the perspective of the question the user will ask, not just the document you are processing.

14. Final Takeaway

Fixed-size, recursive, and semantic chunking are three different ways of answering one fundamental question:

“What information should travel together when this document is searched?”

Fixed-size chunking gives you simplicity and predictability.

Recursive chunking gives more respect to the structure of the document.

Semantic chunking tries to follow changes in meaning and topic.

None of them is a magic button.

A production RAG system should choose chunking based on its documents, its users' questions, its retrieval architecture, and measurable results.

If you remember only three lines, remember these:

Fixed: Cut by size.

Recursive: Respect structure, then control size.

Semantic: Follow the meaning.

And the most important engineering principle:

Don't choose the chunking method that sounds smartest. Choose the one that helps your retrieval system find the right evidence consistently.

Comments