Skip to main content

The Complete RAG Terminology Reference for Engineering & Product Teams

Calculating read time…

This glossary explains the 52 core terms behind Retrieval-Augmented Generation (RAG), in plain, beginner-friendly language.

📌 A. Data Preparation

1. RAG (Retrieval-Augmented Generation)

A technique where an AI model looks up real documents before answering, instead of relying only on what it memorized during training. This makes answers more accurate, current, and traceable to a source.

2. Knowledge Base / Corpus

The full collection of documents a RAG system can search through — manuals, wikis, tickets, contracts, and so on. It's the "library" the system draws its answers from.

3. Document Ingestion

The process of pulling documents into the system from wherever they live — files, websites, databases — so they can be prepared for search. It's the very first step before anything else can happen.

4. Document Loader / Parser

A tool that opens a file (PDF, Word doc, HTML page, etc.) and extracts the readable text from it. Different file types need different parsers because they store text differently.

5. Document Cleaning

Removing junk from extracted text — headers, footers, broken formatting, duplicate whitespace — so the content is clean and readable. Messy input leads to messy search results later.

6. Chunking

Splitting a long document into smaller pieces so the system can search and retrieve at a manageable size. Whole documents are usually too big and unfocused to hand to an AI model directly.

7. Chunk Size

How large each split piece of text is, usually measured in words or tokens. Too small loses context; too large dilutes relevance and wastes space.

8. Chunk Overlap

A small amount of shared text repeated between consecutive chunks, so an idea that spans a chunk boundary isn't cut in half and lost. It's a safety margin for continuity.

9. Metadata

Extra descriptive information attached to a chunk, like its source, date, author, or category. It lets the system filter and label results without re-reading the text itself.

📌 B. Embeddings & Storage

10. Embeddings

A way of turning text into a list of numbers (a vector) that captures its meaning. Texts with similar meaning end up with similar numbers.

11. Embedding Model

The specific AI model used to convert text into embeddings. Different embedding models produce different vector spaces, so the same model must be used consistently for indexing and searching.

12. Vectors

The numeric list produced by an embedding model, representing a piece of text's meaning across many dimensions. Vectors are what the system actually compares when judging similarity.

13. Vector Database / Vector Store

A specialized database built to store and quickly search through millions of vectors. It's the searchable "shelf" where all your embedded chunks live.

14. Vector Index

The internal data structure a vector database builds to make similarity search fast, instead of comparing a query against every vector one by one. It trades a little accuracy for a lot of speed.

📌 C. Search & Retrieval

15. Similarity Search

The process of finding vectors that are mathematically closest to a query vector. Closeness in vector space usually means closeness in meaning.

16. Cosine Similarity

A common way to measure how similar two vectors are, based on the angle between them rather than their length. A score near 1 means very similar meaning; near 0 means unrelated.

17. Semantic Search

Searching by meaning instead of exact words, powered by embeddings and similarity search. It can find a relevant passage even if it uses completely different wording than the question.

18. Keyword / Lexical Search

Searching by matching the exact words or phrases in a query against the exact words in documents. It's fast and precise for exact terms like IDs or codes, but misses different phrasing.

19. BM25

A well-known keyword search algorithm that ranks documents by how often and how rarely a search term appears. It's the classic baseline for lexical search in most search engines.

20. Dense Retrieval

Retrieval based on comparing dense embedding vectors to find semantically similar text. It's the "meaning-based" half of most modern search systems.

21. Sparse Retrieval

Retrieval based on keyword matching, where most values in the representation are zero except for the specific words present. BM25 is the most common sparse retrieval method.

22. Hybrid Search

Combining dense (semantic) and sparse (keyword) retrieval and merging their results. It catches both meaning-based matches and exact-term matches that either method alone would miss.

23. Retriever

The component responsible for searching the knowledge base and pulling back the most relevant chunks for a given query. It's the "search engine" part of a RAG system.

24. Query

The question or request a user types in, which the system turns into a search to find relevant information. Everything in retrieval starts from the query.

25. Query Embedding

The vector representation of the user's query, created with the same embedding model used on the documents. It's what actually gets compared against the stored document vectors.

26. Top-K Retrieval

Retrieving only the K most relevant chunks (like the top 5 or top 10) instead of everything that matches. It keeps the amount of information manageable for the model to read.

27. Similarity Score

A number showing how closely a retrieved chunk matches the query, based on the retrieval method used. Higher scores generally mean the chunk is more likely to be relevant.

28. Relevance Score

A score, often from a reranker, showing how well a chunk actually answers the specific question rather than just resembling it. It's usually more accurate than a plain similarity score.

29. Metadata Filtering

Narrowing search results using metadata, such as only searching documents from a certain date range or department. It's how a system avoids searching (or leaking) content it shouldn't touch.

30. Query Rewriting

Automatically rephrasing a user's question into a clearer or more search-friendly form before retrieval runs. It helps when the original question is vague, short, or conversational.

31. Query Expansion

Adding related terms or synonyms to a query to widen the search net. It helps catch relevant documents that use different wording than the original question.

32. Reranking

Re-scoring and reordering an initial batch of retrieved chunks using a more precise (but slower) model. It moves the truly best chunks to the top after a fast first search.

33. Reranker

The model that performs reranking, typically comparing the query and each candidate chunk together for a more accurate relevance judgment. It's slower than the retriever but far more precise.

34. Cross-Encoder

A model architecture, often used for reranking, that reads the query and a chunk together at the same time rather than separately. This joint reading makes it more accurate but too slow to use on an entire index.

📌 D. Context & Generation

35. Context

The specific text (retrieved chunks) that is handed to the language model along with the question, for it to base its answer on. It's the "evidence" the model is allowed to use.

36. Context Window

The maximum amount of text a language model can read at once, including the question, instructions, and retrieved context. Anything beyond that limit simply doesn't fit and gets cut off.

37. Context Selection

Deciding which retrieved chunks actually make it into the model's prompt, since there's rarely room for everything retrieved. Good selection keeps only the most useful, non-redundant passages.

38. Prompt / Prompt Template

The instructions and structure given to the language model, often a reusable template that inserts the question and retrieved context into a fixed format. It tells the model exactly how to use what it's been given.

39. Grounding

Making sure the model's answer is actually based on the supplied context rather than its own guesses. A well-grounded answer can be traced back to specific retrieved text.

40. Generation / Generator

The part of a RAG system where the language model actually writes the final answer, using the query and retrieved context. It's the "writer" that turns evidence into a readable response.

41. LLM

Large Language Model — an AI model trained on huge amounts of text that can understand and generate human-like language. In RAG, the LLM is usually the generator that produces the final answer.

42. Citations / Source Attribution

Showing which retrieved document or passage supports each part of the answer, so the user can verify it themselves. This is one of RAG's biggest advantages over a plain chatbot.

43. Hallucination

When a language model generates something that sounds confident and fluent but is factually wrong or made up. RAG reduces this risk but doesn't fully eliminate it.

📌 E. Quality, Safety & Operations

44. Retrieval Quality

How well the retriever finds the actually relevant chunks for a given question, rather than just similar-looking ones. Since generation can only be as good as what it's given, this is the ceiling on the whole system's accuracy.

45. Precision

Of the chunks retrieved, the fraction that are actually relevant to the question. High precision means the system isn't wasting context space on irrelevant material.

46. Recall

Of all the truly relevant chunks that exist in the knowledge base, the fraction the system actually managed to retrieve. High recall means the system isn't missing the passage that has the real answer.

47. RAG Evaluation

The practice of systematically testing a RAG system's retrieval and generation quality using metrics and test questions, rather than just eyeballing a few answers. It's how teams catch problems before real users do.

48. Guardrails

Rules and checks placed around a RAG system to prevent unsafe, off-topic, or policy-violating outputs. They act as a safety net around both what goes in and what comes out.

49. Prompt Injection

An attack where hidden or malicious instructions are placed inside a document or user input, trying to trick the model into ignoring its real instructions. RAG systems are especially exposed to this because they read untrusted external documents.

50. Access Control / Authorization

Making sure a user can only retrieve and see documents they're actually permitted to access, matching the same permissions as the original source system. Without this, RAG can accidentally leak private information.

51. Latency

The time it takes for the system to return an answer, from the moment a question is asked. RAG adds extra steps (retrieval, reranking) compared to a plain LLM, so latency usually needs careful optimization.

52. RAG Pipeline

The full end-to-end sequence of steps — ingestion, chunking, embedding, retrieval, reranking, and generation — that together turn a raw question into a grounded answer. It's the complete system, not just any one piece of it.


Comments