Skip to main content

Build Your First RAG Application — A Simple End-to-End Example

Calculating read time…

A practical guide to understanding what actually happens between a user's question and a grounded RAG answer.

Retrieval-Augmented Generation (RAG) is an application pattern that lets a language model answer using information retrieved from an external knowledge collection instead of depending only on what the model already knows. A useful mental model is simple: first prepare your knowledge so it can be searched, then retrieve the most relevant evidence for each question, and finally give that evidence to the language model to formulate the response. 📚

That sounds straightforward, but production RAG is not simply “put documents into a vector database and call an LLM.” Chunk boundaries, metadata, retrieval strategy, access control, context selection, prompt design, evaluation, freshness, latency, cost, and failure handling can all change the quality of the final answer. The goal of this tutorial is to build the mental model first and then connect it to a small end-to-end application. 🚀

The whole application can be understood as two connected pipelines: knowledge preparation and question answering.

Original RAG architecture showing indexing from documents to chunks and vectors, followed by query retrieval and grounded generation.

🔀 Quick Comparison

Approach Where knowledge lives Best fit Main trade-off
Prompt-only The supplied prompt Small, stable context Manual context selection
Fine-tuning Model parameters Changing model behavior or task style Updating factual knowledge is not the primary mechanism
RAG External searchable knowledge Document-grounded answers Retrieval quality becomes part of application quality
Long-context prompting Large supplied context When the relevant corpus can be supplied directly Context selection, cost and attention remain engineering concerns

1. RAG Foundations

👦 Simple analogy: Imagine asking a librarian a question. The librarian does not recite every book in the building. They find a few useful pages and hand those pages to you.

RAG follows the same basic idea. A retrieval system searches a collection of external information and selects material that may help answer the user's question. The language model then receives the question together with the selected context and generates the response.

The important distinction is that retrieval and generation are separate responsibilities. Retrieval is primarily an information-retrieval problem. Generation is primarily a language-model problem. A good RAG application has to make the handoff between them reliable.

✅ Practical example: Suppose a company has an internal travel policy. A user asks, “Can I claim a hotel booked two days before the trip?” A RAG application can retrieve the relevant policy section and ask the language model to answer from that evidence rather than expecting the model's training data to contain the company's current policy.

RAG does not magically make an answer true. If the retriever returns the wrong evidence, the generator may produce a polished answer based on irrelevant material. If the source documents are outdated, retrieval can faithfully return outdated information. If access control is missing, retrieval can expose information the user should never have seen.

🎯 Use this when you need an LLM to answer questions using a knowledge collection that should remain outside the model's parameters and can change independently.

2. The End-to-End Architecture

👦 Simple analogy: Think of a library. Before anyone asks a question, books have to be organized. When someone asks something, the librarian searches the organized collection and gives the useful pages to the person answering.

A typical RAG application therefore has two paths.

  1. Indexing path: source documents are loaded, parsed, cleaned, split into useful chunks, enriched with metadata, converted into embeddings, and stored in a searchable index.
  2. Query path: a user question is analyzed, relevant chunks are retrieved, optional filtering and reranking are applied, the selected context is assembled, and a language model generates the response.

Official Microsoft RAG architecture guidance describes the same separation between preparing grounding data and answering queries. Its current guidance also emphasizes that chunking, enrichment, embedding, retrieval, prompt design and evaluation should be treated as separate engineering decisions rather than one opaque operation.

The separation matters operationally. You should be able to rebuild an index without changing the generation prompt. You should be able to test a different retriever without changing the source documents. You should also be able to compare generation models while keeping the retrieval experiment fixed.

💡 Trade-off: A beginner implementation can look like one short script. Production architecture should resist that temptation. Separating ingestion, indexing, retrieval, generation and evaluation makes failures easier to isolate and experiments easier to reproduce.

🎯 Use this when designing the first version of a RAG system: draw the indexing path and query path separately before choosing tools.

3. Preparing the Knowledge Base

👦 Simple analogy: If you cut a textbook into random pieces, finding one complete explanation becomes difficult. Good chunking is like cutting the textbook along meaningful topics instead of cutting every page in half.

Real-world example: Consider a collection of product-support documents containing installation instructions, troubleshooting steps, limitations and configuration notes. A useful RAG index should preserve enough surrounding meaning that a retrieved chunk can explain the relevant instruction without requiring the model to reconstruct the entire document.

Microsoft's current RAG guidance explicitly treats chunking as an important design phase. Oversized chunks can introduce irrelevant material, while chunks that are too small can lose the surrounding information required to answer a question.

A practical ingestion sequence is:

  1. Identify the source: retain the document identifier, location, version and ownership information.
  2. Extract content: convert PDFs, HTML, Word documents or other formats into usable text or structured content.
  3. Preserve structure: keep headings, sections, tables, page information and other signals when they matter.
  4. Clean the content: remove accidental duplication, navigation noise and extraction artifacts without destroying useful meaning.
  5. Chunk the content: split around semantic or structural boundaries rather than choosing a universal character count blindly.
  6. Add metadata: record information such as document ID, title, section, source, version, date, language, permissions and content type where applicable.
  7. Create embeddings: transform searchable text into numerical representations that support semantic similarity.
  8. Persist the record: store the chunk, embedding and metadata together so retrieval can return both evidence and provenance.

An important production principle is that the chunk should not be treated as “just text.” A useful record is closer to:

{
  "chunk_text": "Example policy content...",
  "document_id": "policy-2026-04",
  "section": "Hotel expenses",
  "source": "travel-policy.pdf",
  "version": "2026-04",
  "access_scope": "employees"
}

Illustrative schema only. The fields and values should be designed around the application's actual governance and retrieval requirements.

Metadata becomes particularly important when semantic similarity alone is insufficient. A user may ask for information from a particular country, product version, department, language or effective-date range. Filtering can prevent semantically similar but unauthorized or outdated chunks from entering the candidate set.

Embeddings provide another useful capability. They represent content numerically so that semantically related pieces can be compared in vector space. But embeddings do not eliminate traditional information retrieval. Exact product codes, policy IDs, error strings and names can behave differently from conceptual questions.

That is why hybrid retrieval can be valuable: lexical search can capture exact terms while semantic retrieval can capture conceptual similarity. Research from Google has described hybrid retrieval as a way to combine complementary lexical and semantic signals.

🎯 Use this when your knowledge source contains structured sections, identifiers, versions or permissions that a pure semantic representation may not capture reliably.

4. Retrieval: Finding the Right Evidence

👦 Simple analogy: Imagine asking a librarian for “the page about cancelling a hotel.” The librarian searches many shelves and returns a small group of likely pages. They do not hand you the entire library.

Hypothetical example: A user asks, “What happens if I cancel a hotel after the free-cancellation period?” The retriever might search thousands of chunks and return several candidates from the hotel-cancellation section.

The retrieval stage can be decomposed into several decisions:

  1. Understand the query: determine what the user is actually asking and whether conversation history changes its meaning.
  2. Apply security filters: restrict candidate documents according to the user's authorization before sensitive content reaches generation.
  3. Run retrieval: use semantic, lexical, hybrid or another appropriate search strategy.
  4. Collect candidates: retrieve enough candidates to give the later ranking stage room to select useful evidence.
  5. Rerank when justified: apply a more expensive relevance model to a smaller candidate set if the application needs additional precision.
  6. Apply context selection: remove duplicates, conflicting fragments and low-value material before constructing the final prompt.

A common beginner mistake is to treat the first few vector-search results as “the answer.” Retrieval returns candidates. The application still has to decide which candidates deserve to become model context.

This is also where query rewriting can help. A conversational question such as “What about international travel?” may be meaningless without the previous turn. A query transformation layer can turn it into something closer to “What are the company's reimbursement rules for international travel?” before retrieval.

For more difficult workloads, retrieval may use multiple searches, metadata filters, reranking or iterative retrieval. The goal is not to retrieve as much as possible. The goal is to retrieve sufficient evidence with manageable noise.

✅ Practical example: If a user asks about “version 4.2,” a good retrieval design can combine semantic similarity with a metadata filter for version 4.2. This reduces the chance that a highly similar passage from version 3.9 becomes the evidence used for the answer.

🎯 Use this when retrieval quality is becoming the bottleneck: inspect retrieved candidates directly instead of changing the generation prompt first.

5. Generation: Turning Evidence into an Answer

👦 Simple analogy: The librarian has found three useful pages. Now a teacher reads those pages and explains the answer in simple language. The teacher should not invent a fourth page.

The generation layer receives at least two conceptual inputs: the user's question and selected retrieved context. A grounding instruction then tells the model how to use that context.

A minimal illustrative prompt structure might look like this:

System:
Answer using the supplied context.
If the context is insufficient, say so.

Context:
[retrieved evidence]

Question:
[user question]

This is intentionally simple. Production prompts often need explicit rules for conflicting sources, citations, uncertainty, output structure, prohibited actions and cases where the retrieved context does not contain an answer.

The crucial concept is context sufficiency. A language model can fail because the retrieved material is insufficient, because the model fails to use sufficient material correctly, or because both problems occur together. Recent Google Research work has explicitly studied this distinction in RAG systems.

💡 Warning: “The LLM hallucinated” is often too vague for debugging. First ask: Did retrieval contain the evidence needed to answer the question? If not, investigate ingestion, chunking, metadata, query formulation and retrieval. If the evidence was present, investigate context selection and generation behavior.

🎯 Use this when debugging answers: separate retrieval failure from generation failure before changing the model or prompt.

6. A Simple Worked Example

👦 Simple analogy: Imagine creating a tiny “company handbook assistant.” You give it three documents, teach its search system where each useful paragraph lives, and then ask it questions.

Hypothetical application: Build a small “IT Policy Assistant” using three internal documents:

  • Password policy
  • Remote-work policy
  • Expense policy

Assume the expense policy contains a section explaining hotel reimbursement. The document is parsed and divided into chunks. Each chunk receives metadata such as document name, section and policy version. The chunks are embedded and stored in a search index.

Now the user asks:

“What is the hotel reimbursement rule for domestic travel?”

The application performs the following sequence:

  1. The application receives the question.
  2. The question is converted into the representation required by the retrieval strategy.
  3. The search layer finds candidate chunks related to hotel reimbursement and domestic travel.
  4. Metadata rules remove documents the user cannot access or versions that should not be used.
  5. The strongest candidates are selected.
  6. The application builds a context containing the selected evidence.
  7. The language model receives the question and context.
  8. The response is generated and, ideally, linked back to the supporting document or section.

Now consider a failure. Suppose the question is “What is the hotel reimbursement rule for international travel?” but the index contains only domestic policy information. A responsible system should not manufacture an international rule simply because the question sounds similar. It should indicate that the available evidence is insufficient.

That single example captures a central RAG principle: retrieval determines what evidence is available; generation determines how that evidence is communicated.

🎯 Use this when teaching RAG to a team: start with three small documents and trace one question through every stage before introducing advanced infrastructure.

7. Implementation Blueprint

👦 Simple analogy: Building a RAG app is like assembling a restaurant kitchen. You need ingredients, preparation, storage, order handling and a chef. Buying a better chef does not fix spoiled ingredients.

A framework such as LangChain can provide reusable building blocks for retrieval and application orchestration. Its current learning material includes semantic search and RAG-agent tutorials. A managed search service such as Azure AI Search can provide search and vector-retrieval capabilities. These are implementation choices, not requirements of RAG itself.

A framework-neutral implementation can be organized around these interfaces:

documents = load_documents()
chunks = split_into_chunks(documents)
records = enrich_with_metadata(chunks)
index(records)

question = receive_question()
candidates = retrieve(question)
context = select_context(candidates)
answer = generate(question, context)

Illustrative pseudocode. It intentionally avoids provider-specific API names and version assumptions.

For a first working prototype, keep the architecture intentionally small:

  1. Use a small, representative document set.
  2. Make the chunk records inspectable.
  3. Store source metadata with every chunk.
  4. Implement one retrieval strategy first.
  5. Log retrieved chunks for every test question.
  6. Use a conservative grounding prompt.
  7. Build a small evaluation set before tuning aggressively.
  8. Only then experiment with hybrid retrieval, reranking, query rewriting or more sophisticated orchestration.

This sequence makes debugging much easier. If the application gives a wrong answer, you can inspect the original document, chunk, metadata, retrieval result, selected context and final prompt rather than staring only at the generated response.

🎯 Use this when building your first prototype: optimize for observability and understanding before optimizing for architectural sophistication.

8. Evaluating the RAG Application

👦 Simple analogy: If a student gets an answer wrong, the teacher asks whether the student read the wrong page or misunderstood the right page. RAG evaluation needs the same separation.

Do not evaluate a RAG system only by looking at final answers. A useful evaluation process examines at least two layers: retrieval and generation.

Layer Question Useful evidence
Retrieval Did we retrieve relevant evidence? Relevant-document coverage, ranking quality, metadata correctness
Context Did useful evidence survive selection? Selected chunks, duplicates, conflicts, context completeness
Generation Did the model answer from the evidence? Correctness, groundedness, completeness, unsupported claims
Operations Can the system meet production requirements? Latency, cost, failures, throughput, access-control incidents

Create a representative test set before making major changes. Include easy questions, ambiguous questions, questions requiring multiple chunks, questions whose answers are absent, questions involving exact identifiers and questions that test permissions.

For every test case, retain the question, expected evidence, acceptable answer characteristics and observed retrieved context. This creates a regression harness. When chunking or retrieval changes, you can determine whether an improvement helped one class of questions while damaging another.

LLM-as-judge evaluation can accelerate review, but it should not automatically become the only source of truth. Judge prompts can be biased, models can systematically prefer certain answer styles, and evaluation can become circular when the same model family is used both to generate and judge responses. Human review remains useful for calibration and difficult cases.

A mature evaluation loop therefore looks like this:

  1. Define representative questions.
  2. Record expected evidence or answer criteria.
  3. Measure retrieval independently.
  4. Measure final answer quality.
  5. Inspect disagreements manually.
  6. Change one major variable at a time.
  7. Run regression tests before deployment.
  8. Monitor production behavior after release.

🎯 Use this when a team says “the RAG is hallucinating”: reproduce the question and inspect retrieval before changing the model.

9. Enterprise Rollout

👦 Simple analogy: A classroom project can survive if someone remembers where the files are. An enterprise system cannot depend on memory. Someone must own the documents, indexes, permissions, deployments and incidents.

Moving from prototype to production changes the problem. The question becomes not only “Can it answer?” but also “Can we operate it safely, repeatedly and economically?”

Ownership: define who owns source content, ingestion, retrieval configuration, generation prompts, evaluation datasets and production incidents.

Versioning: version source documents, chunking logic, embedding configuration, retrieval settings, prompts and evaluation datasets. Otherwise a quality regression can become impossible to reproduce.

Access control: enforce authorization before sensitive chunks are supplied to the generation model. A RAG system should not treat “the user can ask about it” as proof that the user is allowed to see it.

Freshness: define how documents enter the system, how updates are detected, when indexes are rebuilt and how obsolete versions are removed or excluded.

Security: retrieved documents are data, not automatically trusted instructions. A malicious document can contain text attempting to manipulate the model. Treat retrieved content as potentially adversarial and design prompts, permissions and tool access accordingly.

Observability: record safe operational telemetry around retrieval latency, candidate counts, selected sources, model latency, failures, token consumption where applicable, and evaluation outcomes. Avoid logging sensitive content unnecessarily.

Budget controls: generation is not the only cost. Document parsing, OCR, embedding, indexing, retrieval infrastructure, reranking and repeated experimentation can also consume resources. Measure the entire pipeline.

Deployment: use controlled releases. A retrieval change can alter the evidence presented to the model even when the generation prompt remains unchanged. Regression testing and canary releases therefore apply to retrieval configuration as well as application code.

Rollback: define objective rollback signals before production deployment. Examples include a regression in critical evaluation cases, elevated retrieval failures, increased latency, unacceptable authorization failures or unexpected cost growth.

🎯 Use this when moving beyond a demo: treat the index, retrieval configuration and evaluation set as production assets—not disposable implementation details.

10. Common Mistakes

👦 Simple analogy: If a restaurant serves a bad meal, replacing the waiter does not fix spoiled ingredients. Debug the stage that caused the problem.

Mistake 1 — Sending entire documents to the model. Large documents often contain unrelated information. The model has to process more context, while the relevant passage may receive less useful attention. Better document segmentation and context selection usually provide a more controllable pipeline.

Mistake 2 — Choosing chunk size by folklore. A universal chunk size cannot account for contracts, manuals, source code, FAQs and tables having different structures. Chunking should be tested against representative questions and documents.

Mistake 3 — Using vector similarity for everything. Semantic similarity is powerful, but exact identifiers, error codes and product names can benefit from lexical matching or metadata filters. Retrieval strategy should follow query characteristics.

Mistake 4 — Measuring only the final answer. A final answer can be wrong because retrieval failed or because generation failed. Without intermediate evidence, teams often tune the wrong component.

Mistake 5 — Ignoring metadata. Without version, source, permissions or effective-date information, the system may retrieve something semantically relevant but operationally incorrect.

Mistake 6 — Assuming retrieved text is trustworthy. Retrieval makes external text available to the model. It does not make that text safe instructions. Treat external content as data and apply security controls around what the model can do with it.

Mistake 7 — Building without a regression set. RAG systems have many interacting variables. Changing chunking, embeddings or ranking can improve one query while damaging another. A stable test set provides the feedback loop required to see those changes.

Mistake 8 — Assuming RAG automatically eliminates hallucinations. RAG supplies evidence; it does not guarantee that the generator will use the evidence correctly. Grounding instructions, context selection, evaluation and appropriate refusal behavior remain necessary.

🎯 Use this when debugging: trace the question from source document to chunk to retrieval result to final context to generated answer.

11. ❓ FAQ

What is the simplest way to understand RAG?

Think of RAG as search plus generation. The application searches an external knowledge collection, selects relevant evidence, and supplies that evidence to a language model before generating the answer.

Do I need a vector database to build RAG?

Not conceptually. RAG requires a mechanism for finding relevant external information. Vector retrieval is one important approach, but lexical, hybrid and other retrieval strategies can also be part of a RAG architecture.

Why does chunking matter so much?

A chunk is the unit that retrieval can return. If it is too broad, it can contain unnecessary material; if it is too narrow, it can lose the context needed to answer the question. The appropriate strategy depends on document structure and the application's questions.

How do I know whether a RAG problem is retrieval or generation?

Inspect the retrieved evidence. If the required information was never retrieved, investigate ingestion, chunking, metadata, query formulation or retrieval. If the required evidence was retrieved but the answer ignores or misinterprets it, investigate context selection and generation.

What should I build first when learning RAG?

Start with a small document collection and one question-answering flow. Make chunks and retrieved results visible, preserve source metadata, and create a small evaluation set before adding advanced retrieval or agentic behavior.

12. 🔗 References & Further Reading

13. 📝 Summary

  • Foundations: RAG combines external retrieval with language-model generation.
  • Architecture: Separate the indexing pipeline from the query pipeline.
  • Knowledge preparation: Good parsing, chunking, metadata and embeddings create the foundation for retrieval.
  • Retrieval: The goal is not maximum context; it is useful, authorized and sufficient evidence.
  • Generation: The model should transform retrieved evidence into an answer without inventing unsupported information.
  • Worked example: A small policy assistant is enough to understand the complete RAG loop.
  • Implementation: Start with simple, inspectable components before adding sophisticated orchestration.
  • Evaluation: Measure retrieval and generation separately and maintain regression tests.
  • Enterprise rollout: Treat documents, indexes, permissions, evaluation and observability as production assets.
  • Common mistakes: Most RAG failures become easier to diagnose when the pipeline is inspected stage by stage.
  • FAQ: RAG is fundamentally an evidence-retrieval and grounded-generation pattern, not merely a vector database feature.

If you understand the journey from document → chunk → embedding → retrieval → context → generation → evaluation, you already have the foundation needed to start reasoning about much more advanced RAG systems.

Comments