Skip to main content

The Two Pipelines of RAG — Indexing vs Query-Time Retrieval

Calculating read time…

RAG is not one pipeline. It is two connected pipelines: one prepares your knowledge for search, and the other finds the right knowledge when a user asks a question. 🔎

That distinction is one of the most important ideas to understand before designing or troubleshooting a production Retrieval-Augmented Generation system. A retrieval problem can originate months earlier during ingestion, parsing, chunking, embedding, or indexing—or it can appear milliseconds before generation because the query was rewritten poorly, filters were wrong, retrieval was too broad, or reranking selected the wrong evidence. Understanding which pipeline owns which responsibility makes debugging dramatically easier. 🧭

Original RAG architecture diagram showing the separate indexing pipeline and query-time retrieval pipeline

💡 The key mental model

Think of indexing as building and maintaining the library catalogue. Query-time retrieval is the librarian finding the right books for one question. A perfect librarian cannot find a book that was never catalogued correctly—and a beautifully indexed library can still produce a poor answer if the librarian asks the wrong search question.

🔀 Quick Comparison

Dimension Indexing Pipeline Query-Time Retrieval
When?Before user queries, and whenever knowledge changes.For each user request or retrieval operation.
Main jobTurn source material into searchable units.Select useful evidence for the current query.
Typical workParse, clean, chunk, enrich, embed, index.Rewrite, filter, search, merge, rerank, select.
Main failureThe right information never becomes retrievable in the right form.The right information exists but the query path fails to select it.
OptimizationContent quality, chunking, metadata, embeddings, index design.Query formulation, filters, retrieval depth, ranking, context budget.
Operational concernFreshness, versioning, re-indexing, pipeline failures.Latency, cost, access control, query spikes, regressions.

1. The Two-Pipeline Mental Model

Imagine an enormous company library. Before anyone asks a question, somebody must collect documents, understand their structure, divide them into useful pieces, label them, and place them where a search system can find them. That is indexing.

Now imagine an employee asks: “What is the current approval process for an international business trip?” The system does not rebuild the library. It interprets the question, searches the prepared index, applies access and metadata constraints, ranks candidate passages, and sends a selected context to the language model. That is query-time retrieval.

Production RAG systems therefore have a temporal boundary:

Before the question

Prepare the searchable representation of knowledge.

After the question

Select the most useful evidence for this specific request.

AWS describes RAG as a sequence in which documents are prepared and embedded into a vector database, followed by query-time similarity search and generation. Microsoft similarly separates document preparation, chunking, indexing, search strategy, and later evaluation of retrieval and generation.

Important: “indexing” does not necessarily mean only writing vectors into a vector database. In a production system, it is better understood as the entire preparation path that transforms source material into searchable records.

🎯 Use this when... you are debugging a RAG system and need to decide whether the problem exists in knowledge preparation or in query-time search.

2. Pipeline 1 — Indexing: Build the Searchable Knowledge Layer

A simple analogy: indexing is preparing a warehouse before customers arrive. If boxes are mixed together, unlabeled, broken apart incorrectly, or stored in the wrong aisle, the warehouse worker will struggle later—even if the worker is extremely capable.

Technically, indexing transforms source content into searchable units. Depending on the architecture, those units may contain text, embeddings, metadata, identifiers, document locations, permissions, timestamps, and other fields required by retrieval.

2.1 Step 1 — Acquire the source documents

Start with authoritative sources: policies, product documentation, procedures, contracts, knowledge articles, tickets, structured records, or other approved enterprise data. The source itself should have an identity and lifecycle.

A production record should ideally allow you to answer questions such as:

  • Where did this content come from?
  • Which source version produced this chunk?
  • When was it ingested?
  • Who is allowed to see it?
  • What document and section does the chunk belong to?
  • Which index generation contains it?

✅ Practical example

Suppose an HR policy document is replaced by a new approved version. A production index should not silently mix chunks from the old and new versions. Store enough provenance to identify which source version produced each searchable record.

2.2 Step 2 — Parse the content

Parsing converts files or records into information that the indexing process can reason about. The difficult part is that a document's meaning is not always represented by plain text alone.

Consider a PDF containing:

  • a heading;
  • three paragraphs;
  • a table of approval limits;
  • a footnote;
  • a page number;
  • an image containing a process diagram.

A parser that extracts only the visible paragraph text may lose relationships that matter during retrieval. Production ingestion therefore needs to understand the document types and the information structures that users actually query.

Modern managed RAG services also expose parsing and chunking as explicit stages rather than treating documents as an undifferentiated blob. Google Cloud's RAG interfaces, for example, represent chunks together with metadata such as source information and page spans.

2.3 Step 3 — Clean and normalize

Cleaning is not simply “remove everything that looks strange.” Some apparently repetitive text is actually valuable. A document title, product name, policy identifier, or section heading may provide context that improves retrieval.

Useful normalization can include removing extraction artifacts, repairing broken whitespace, preserving headings, identifying tables, normalizing metadata, and separating boilerplate from meaningful content.

The objective is not to make text aesthetically perfect. The objective is to create searchable units whose content and metadata preserve the information users will need later.

2.4 Step 4 — Chunk the documents

Child-friendly analogy: if a textbook is too large to hand to someone looking for one fact, you divide it into meaningful pages or sections.

Chunking divides a document into smaller retrieval units. The important word is meaningful. A chunk should contain enough surrounding information to make its content understandable, while remaining focused enough that retrieval can select it without bringing a large amount of irrelevant material.

Microsoft's RAG guidance explicitly treats chunking as a major design phase because oversized chunks can increase token consumption and introduce irrelevant information, while poor chunking can reduce retrieval quality.

Common strategies include:

  • Fixed-size: split content according to a defined size.
  • Structure-aware: use headings, sections, paragraphs, tables, or document boundaries.
  • Semantic: attempt to keep conceptually coherent material together.
  • Hierarchical: maintain relationships between larger parent sections and smaller child passages.

There is no universal chunk size that is correct for every corpus. The appropriate design depends on document structure, query patterns, embedding behavior, retrieval method, and the amount of context the downstream model can use effectively.

2.5 Step 5 — Attach metadata

Think of metadata as the labels on a library shelf. The text tells you what the passage says; metadata can tell you where it came from, who owns it, what version it belongs to, and which business context it applies to.

Useful metadata can include:

  • document ID and chunk ID;
  • source URI or repository location;
  • document title and section;
  • business unit or product;
  • language;
  • document version;
  • publication or effective date;
  • security classification;
  • authorization attributes;
  • content type.

Metadata can be used at retrieval time for filtering. Google Cloud documents metadata filtering for RAG retrieval, while OpenAI's vector-store API supports attributes attached to vector-store files. These are product-specific implementations of a broader architectural idea: searchable content can carry structured information that retrieval uses alongside semantic similarity.

2.6 Step 6 — Generate embeddings

An embedding model converts text into a numerical representation designed to capture useful relationships between pieces of content.

For example, a policy passage containing “employees must obtain manager approval before booking international travel” and a user query asking “Do I need my manager's approval for overseas travel?” may use different words while referring to related concepts.

The embedding representation gives a semantic retrieval system a way to compare such content mathematically.

A crucial production rule follows: the query representation and indexed representation must be compatible with the retrieval design. Changing an embedding model is therefore not merely a configuration tweak. It can require re-embedding the corpus and reevaluating retrieval quality.

2.7 Step 7 — Build the searchable index

The final stage places searchable representations into an index optimized for the selected retrieval strategy. Vector search is one approach; keyword search and hybrid search are others.

AWS describes vector retrieval as a process where chunks are embedded, vectors are indexed, and a query is converted into a compatible vector for similarity search. It also notes that index configuration influences retrieval speed and accuracy.

A production index may therefore contain more than one searchable representation:

  • dense vector representation;
  • lexical or keyword representation;
  • metadata fields;
  • document and chunk identifiers;
  • authorization attributes;
  • version and freshness information.

🎯 Use this when... your RAG answers are repeatedly missing information that definitely exists in your source corpus. Inspect the indexing pipeline before changing the LLM prompt.

3. Pipeline 2 — Query-Time Retrieval: Find Evidence for One Question

The second pipeline starts when the user asks something. A useful analogy is a librarian receiving a specific request and deciding which shelves, books, pages, and paragraphs are relevant.

Unlike indexing, this pipeline runs repeatedly. That makes latency, availability, authorization, query variability, and operational cost particularly important.

3.1 Step 1 — Understand the query

The literal user question is not always the best search query.

Suppose the user asks:

“What happens if I need to travel overseas for a customer meeting?”

A retrieval system may need to identify concepts such as international travel, business travel, customer meetings, approvals, booking rules, and expense policy. Query rewriting can transform conversational wording into search-oriented representations.

However, query rewriting creates a new failure mode: the rewritten query can accidentally remove an important constraint. Therefore, the original query should remain observable and, where appropriate, available to downstream logic.

3.2 Step 2 — Apply authorization and metadata filters

This is one of the most important enterprise distinctions: relevance is not authorization.

A document can be semantically perfect for a question and still be forbidden for the requesting user.

Therefore, retrieval should enforce access constraints rather than hoping the language model will ignore unauthorized content after it has already been retrieved.

Filters can also narrow the search by product, geography, document type, effective date, department, language, or other business attributes. Metadata filtering is explicitly supported in some current managed RAG systems.

💡 Security warning

Do not treat retrieval scores as permission checks. A high similarity score only says that content appears relevant according to the search mechanism. Authorization must come from an explicit access-control design.

3.3 Step 3 — Search the index

The retriever searches the prepared index. There are three broad patterns:

  • Keyword or lexical search: useful when exact terms, identifiers, product codes, or names matter.
  • Dense semantic search: useful when meaning matters more than exact wording.
  • Hybrid retrieval: combines lexical and semantic signals.

The choice should follow query behavior rather than fashion. A support system searching for an exact error code may benefit greatly from lexical matching. A knowledge assistant answering conceptual questions may benefit from semantic similarity. Many enterprise corpora contain both kinds of queries.

AWS documentation describes vector similarity using measures such as cosine distance, Euclidean distance, or dot product, depending on the system. The specific metric and score semantics must be understood before comparing scores across systems.

3.4 Step 4 — Retrieve candidates

The first search does not necessarily need to produce the final context. It can produce a candidate pool that later stages inspect.

This separation is useful because the first retrieval stage can prioritize recall—bringing potentially useful evidence into the candidate set—while a later ranking stage can improve precision.

This is also where top-k becomes important. A very small candidate set can miss relevant evidence. A very large candidate set can introduce noise, increase processing, and consume more context.

3.5 Step 5 — Rerank the candidates

Reranking is like asking a second librarian: “Of these potentially useful books, which ones actually answer this particular question?”

A reranker can inspect the query and candidate passages together and produce a more query-specific ordering than the initial retrieval stage.

This creates a useful architecture:

  1. Retrieve a candidate set efficiently.
  2. Apply authorization and business filters.
  3. Rerank candidates using a more discriminating signal when justified.
  4. Select the final context set.

Not every RAG system requires every stage. Additional stages are valuable only when they solve a measured problem without creating unacceptable latency or operational complexity.

3.6 Step 6 — Assemble context

The retriever's job is not simply “return the top documents.” It must produce context that the generation stage can interpret correctly.

Context assembly may involve:

  • removing duplicate passages;
  • preserving source identifiers;
  • grouping related chunks;
  • placing relevant passages in a deliberate order;
  • compressing redundant content;
  • respecting a context budget.

Microsoft's RAG guidance highlights the practical tension between retrieval quality, latency, and token usage: retrieved passages become model input, so retrieving more material is not automatically better.

3.7 Step 7 — Generate the answer

Only after retrieval has selected evidence does the generation layer receive the user request and retrieved context.

This leads to a powerful debugging principle:

If the correct evidence is not in the retrieved context, prompt engineering alone cannot reliably fix the retrieval failure.

The generation model can only work with the evidence and instructions it receives. Retrieval quality is therefore an upstream dependency of grounded generation.

🎯 Use this when... the correct documents exist in your index but the application still produces poor answers. Inspect query transformation, filters, retrieval, ranking, and context assembly.

4. Worked Enterprise Example — An Internal Travel Policy Assistant

The following is a hypothetical enterprise example, used to demonstrate the architecture rather than claim a specific company's implementation.

Imagine an organization with thousands of documents covering travel, expenses, approvals, customer visits, regional policies, and reimbursement rules.

4.1 What happens during indexing?

  1. Approved policy documents are collected from the authoritative repository.
  2. Each document is parsed into structured text and document elements.
  3. Headings, sections, tables, and important metadata are preserved.
  4. Large documents are divided into coherent chunks.
  5. Each chunk receives identifiers and metadata such as policy version and effective date.
  6. Embeddings are generated for the searchable text.
  7. The chunks, metadata, and searchable representations are written to the index.
  8. The pipeline records the index version and ingestion outcome.

4.2 A user asks a question

“I am visiting a customer in Germany. What approvals do I need before booking the trip?”

4.3 What happens at query time?

  1. The application receives the question and authenticated user identity.
  2. The retrieval layer identifies relevant concepts such as international travel, customer visit, booking, and approval.
  3. Security and business filters restrict the searchable corpus.
  4. The system performs lexical, semantic, or hybrid retrieval.
  5. A candidate set is produced.
  6. Candidate passages are reranked if the architecture uses a reranking stage.
  7. The system selects the final evidence and preserves source identifiers.
  8. The selected context is passed to the generation model.
  9. The response can cite or otherwise expose the underlying sources, depending on the application design.

Notice what did not happen: the system did not re-parse every policy document when the user asked the question. The expensive preparation work was performed earlier.

✅ Practical debugging example

Suppose the assistant cannot find the approval rule. First inspect whether the current policy version was indexed. Then inspect the chunk containing the rule. Then inspect metadata and authorization filters. Only after confirming that the correct evidence is retrievable should you investigate query rewriting, ranking, context assembly, or generation.

🎯 Use this when... you need a concrete mental model for explaining RAG architecture to application developers, data engineers, security teams, or business stakeholders.

5. Implementation Blueprint

A production implementation can be designed as two independently observable workflows connected by a shared searchable data contract.

5.1 Indexing workflow

SOURCE
  ↓
PARSE
  ↓
NORMALIZE
  ↓
CHUNK
  ↓
ENRICH METADATA
  ↓
EMBED
  ↓
INDEX
  ↓
VALIDATE + PUBLISH INDEX VERSION

The final “validate + publish” stage is important. A failed indexing job should not automatically become the new production knowledge base.

5.2 Query-time workflow

USER QUESTION
  ↓
IDENTITY + POLICY
  ↓
QUERY UNDERSTANDING
  ↓
FILTERS
  ↓
RETRIEVE CANDIDATES
  ↓
RERANK / DEDUPLICATE
  ↓
SELECT CONTEXT
  ↓
GENERATE
  ↓
CITATIONS / RESPONSE

5.3 Example record design

{
  "document_id": "policy-1042",
  "chunk_id": "policy-1042-sec-4",
  "title": "International Travel Policy",
  "section": "Pre-Trip Approval",
  "version": "approved-version",
  "effective_date": "YYYY-MM-DD",
  "access_scope": "travel-policy",
  "text": "Illustrative chunk content",
  "embedding": "vector representation"
}

The values above are illustrative. The important design idea is that the searchable unit carries both semantic content and enough provenance to support retrieval, filtering, debugging, auditing, and lifecycle management.

5.4 Keep pipeline contracts explicit

A common architectural mistake is allowing indexing and retrieval teams to evolve independently without agreeing on a shared contract.

At minimum, define:

  • what a chunk is;
  • which metadata fields are mandatory;
  • how document versions are represented;
  • how deleted content is removed from the index;
  • which fields participate in filtering;
  • how authorization attributes are represented;
  • which embedding representation is expected;
  • how index versions are published and retired.

🎯 Use this when... multiple teams own ingestion, search, application, and model layers and you need clear ownership boundaries.

6. Evaluate the Two Pipelines Separately

One of the most useful production practices is to avoid evaluating only the final answer. If a RAG response is wrong, you need evidence about where the failure occurred.

Microsoft's current RAG evaluation guidance explicitly separates document retrieval evaluation from system-level measures such as groundedness, relevance, and response completeness.

6.1 Indexing-quality questions

  • Was the source document ingested?
  • Was the correct version ingested?
  • Did parsing preserve important information?
  • Did chunking keep related facts together?
  • Were important metadata fields populated?
  • Were embeddings generated successfully?
  • Was the chunk actually published to the active index?
  • Was stale content removed or superseded correctly?

6.2 Retrieval-quality questions

  • Did the correct source appear in the candidate set?
  • Did filtering remove relevant evidence?
  • Did semantic or lexical search miss an exact term?
  • Did reranking move the best evidence downward?
  • Was the final context unnecessarily noisy?
  • Was relevant evidence truncated because of a context budget?

When ground-truth relevance labels exist, retrieval-specific metrics can measure ranking quality. Microsoft documents measures such as NDCG and other retrieval metrics for evaluating whether relevant documents appear in appropriate positions.

6.3 Generation-quality questions

  • Did the answer accurately use the retrieved evidence?
  • Did the response introduce unsupported information?
  • Did it answer the actual user question?
  • Did it omit an important retrieved fact?
  • Were citations associated with the correct evidence?

This separation creates a diagnostic chain:

Source exists? → Indexed correctly? → Retrieved? → Ranked correctly? → Included in context? → Used correctly by the model?

That chain is far more actionable than a single “RAG accuracy” number.

🎯 Use this when... a team is arguing about whether the embedding model, vector database, retriever, prompt, or LLM is responsible for poor answers.

7. Enterprise Rollout: Treat Indexing as a Product, Not a Script

A prototype can index a folder once. An enterprise RAG platform cannot assume that knowledge stays still.

7.1 Ownership

Define who owns source quality, ingestion, index infrastructure, retrieval configuration, application prompts, evaluation, security, and incident response. “The AI team owns RAG” is usually too broad to be operationally useful.

7.2 Version the knowledge layer

Treat index generations as deployable artifacts. A useful model is:

  1. Build a new index generation.
  2. Run validation and retrieval tests.
  3. Compare against the current production index.
  4. Publish the new generation.
  5. Monitor production behavior.
  6. Retain a rollback path.

This is particularly important when chunking logic, embeddings, metadata schema, or source content changes together. A seemingly small indexing change can alter retrieval behavior across many queries.

7.3 Freshness and re-indexing

Not every source needs the same refresh frequency. A legal or HR policy may change infrequently but require strong version controls. A support knowledge base may change more frequently. A transactional source may need a different architecture altogether.

Design freshness policies around business requirements rather than using one universal schedule.

7.4 CI gates

Indexing changes should be testable in CI or an equivalent controlled evaluation process. Useful gates can include:

  • parsing success rate;
  • required metadata completeness;
  • duplicate detection;
  • retrieval regression tests;
  • security-filter tests;
  • representative query-set evaluation;
  • latency checks;
  • cost checks.

7.5 Privacy and production-derived data

Evaluation data can itself contain sensitive information. Production-derived questions, retrieved passages, feedback, and traces should be handled according to the organization's data-protection requirements.

Do not assume that moving content into an evaluation dataset makes the content non-sensitive.

7.6 Observability

A useful RAG trace should allow engineers to reconstruct the retrieval decision without exposing unnecessary sensitive content.

Depending on the environment, capture structured telemetry for:

  • index version;
  • retrieval strategy;
  • filter decisions;
  • candidate counts;
  • ranking stages;
  • retrieved source identifiers;
  • latency by stage;
  • model and configuration identifiers;
  • evaluation outcomes;
  • error categories.

7.7 Incident response

When a RAG system suddenly gives poor answers, the first operational question should not automatically be “Did the LLM change?” Check the pipeline boundary.

  1. Did the source system change?
  2. Did ingestion fail?
  3. Did the active index version change?
  4. Did metadata filtering change?
  5. Did retrieval configuration change?
  6. Did reranking change?
  7. Did the prompt or generation model change?

This turns a vague AI incident into a sequence of testable hypotheses.

🎯 Use this when... RAG moves from an experiment into a business-critical application with security, governance, freshness, and rollback requirements.

8. Common Mistakes — And Why They Hurt

Mistake 1: Treating RAG as “vector database + LLM”

This hides the ingestion, metadata, authorization, ranking, evaluation, and lifecycle problems that dominate production reliability. The result is often an architecture that works on a small demo corpus but becomes difficult to diagnose at scale.

Mistake 2: Fixing retrieval failures with prompt changes

If the required evidence never reaches the context, a better instruction cannot manufacture missing evidence. Prompt changes may alter behavior temporarily while leaving the actual retrieval defect untouched.

Mistake 3: Choosing chunk size once and never evaluating it

Chunking changes what the retriever can see as an independent unit. If chunks are too broad, relevant facts can become diluted by unrelated content. If they are too narrow, essential context can be separated. Microsoft and AWS both identify chunking as an important RAG design consideration.

Mistake 4: Ignoring metadata

Without useful metadata, retrieval has fewer ways to constrain the search space and enforce business context. This becomes especially problematic when one corpus contains multiple products, regions, versions, or access scopes.

Mistake 5: Assuming semantic search replaces keyword search

Exact identifiers, error codes, legal phrases, SKUs, product names, and version numbers may behave differently from natural-language concepts. The retrieval strategy should reflect the actual query distribution.

Mistake 6: Evaluating only the final answer

A final-answer score cannot tell you whether the problem was missing evidence, poor ranking, context assembly, or generation. Process-level retrieval evaluation provides the diagnostic signal needed to isolate upstream problems.

Mistake 7: Mixing document versions

A user can receive a response assembled from passages that were individually relevant but collectively inconsistent because they came from different policy versions. Version metadata and controlled index publication reduce this class of failure.

Mistake 8: Using retrieval scores as universal truth

A similarity score is meaningful within the context of its search implementation. Different metrics, indexes, models, and systems can produce scores with different meanings and ranges. Google Cloud's retrieval documentation explicitly notes that score interpretation depends on the underlying vector database and metric type.

🎯 Use this when... reviewing a RAG design and looking for failure modes before they become production incidents.

9. ❓ FAQ

Q1. Is indexing part of RAG?

Yes. In a production architecture, indexing is the preparation path that turns source knowledge into searchable representations. It normally happens before individual user queries and is repeated when the source knowledge changes.

Q2. Does query-time retrieval create embeddings every time?

The query may be embedded at retrieval time when dense vector search is used, while document embeddings are generally prepared during indexing. The exact implementation depends on the retrieval architecture.

Q3. Why not retrieve the whole document?

Whole-document retrieval can introduce irrelevant information and increase the amount of context passed downstream. Chunking allows the system to select smaller evidence units that are more closely related to the query.

Q4. Should every RAG system use reranking?

No. Reranking can improve candidate ordering in some workloads, but it adds processing and operational complexity. It should be introduced when evaluation shows that the additional stage solves a meaningful retrieval problem.

Q5. How do I know whether my RAG problem is indexing or retrieval?

Trace the evidence backward. First verify that the correct source and version were indexed, then check the relevant chunk and metadata, then inspect candidate retrieval and ranking, and finally inspect context assembly and generation. This isolates the failing stage instead of guessing.

10. 🔗 References & Further Reading

11. 📝 Summary

  • The Two-Pipeline Mental Model: indexing prepares the searchable knowledge world; query-time retrieval selects evidence from that world.
  • Indexing: ingestion, parsing, cleaning, chunking, metadata, embeddings, and index publication determine what can be found later.
  • Query-Time Retrieval: query understanding, authorization, filtering, retrieval, reranking, and context selection determine what evidence reaches the model.
  • Worked Example: a travel-policy assistant shows why indexing and retrieval must be debugged as separate stages.
  • Implementation: explicit contracts between indexing and retrieval make production systems easier to operate.
  • Evaluation: retrieval quality should be measured separately from final-answer quality so failures can be localized.
  • Enterprise Rollout: index versioning, ownership, freshness, security, CI gates, observability, budgets, and rollback are part of RAG engineering.
  • Common Mistakes: most serious failures come from treating RAG as a single black box rather than a chain of independently testable stages.
  • FAQ: indexing and retrieval answer different operational questions and should be designed accordingly.


Comments