Skip to main content

Text Embeddings

Calculating read time…

Embeddings in RAG are the translation layer that turns text into the numbers a computer can actually compare — and getting this layer wrong quietly wrecks retrieval quality no matter how good everything else is. In our last two deep dives, we turned messy files into clean text (parsing), then cut that text into small, meaningful pieces (chunking). We now have a pile of perfect little text chunks. But a chunk of text is still just a string of characters — and computers can't compare "meaning" between two strings. They can only compare numbers.

That's the entire reason embeddings exist. An embedding model is the translator that converts a chunk of human language into a list of numbers — a vector — positioned in space so that texts with similar meaning end up physically close together, no matter how differently they're worded. Get this translation step wrong, and no amount of clever chunking, reranking, or prompt engineering downstream can save your RAG system.

This post assumes zero prior embedding knowledge. We'll build the full mental model from first principles — what a vector space actually is, how embedding models are trained, bi-encoders vs. cross-encoders vs. late-interaction models, dimensions and Matryoshka compression, similarity math, the real embedding model landscape, fine-tuning, evaluation, and the pitfalls that quietly wreck production retrieval quality.

Embeddings in RAG: vector space, bi-encoders, and similarity search explained
🧬 An embedding is a coordinate — it places meaning in space, not just words in a list
📐 Cosine similarity, not word overlap, is how RAG decides "these two texts are related"
🏗️ Almost every embedding model in production is a bi-encoder — that single architectural choice is why RAG can scale to billions of documents
🪆 Matryoshka Representation Learning now lets you shrink vectors 4x with barely any accuracy loss — a genuine shift in RAG economics
⚖️ Open-source embedding models like Qwen3-Embedding have closed the gap with — and in cases overtaken — closed APIs on retrieval benchmarks

Let's build the full mental model, one layer at a time.

🧠 Section 0: What Is an Embedding, Really?

Formally: an embedding is a fixed-length array of numbers (a vector) that represents the meaning of a piece of text. A sentence like "The cat sat on the mat" might become something like [0.021, -0.184, 0.097, ..., 0.045] — often 768, 1024, or 3072 numbers long. That's it. No magic — just a long list of coordinates in a very high-dimensional space.

💡 The City Map Analogy

Think of a city map. Two addresses that are physically close to each other are easy to describe as "nearby." An embedding model does exactly this, but for meaning instead of geography — it builds an imaginary map where "dog" and "puppy" are neighbors, "dog" and "car" live in different districts entirely, and "refund policy" sits close to "return policy" even though the two phrases don't share a single common word. Embedding a piece of text is simply plotting its address on this meaning-map.

This is what makes embeddings the backbone of retrieval: instead of matching keywords (which fails the moment your user's wording differs from the document's wording), you match coordinates — and coordinates capture meaning, synonyms, paraphrases, and even cross-lingual equivalence, all at once.


🗺️ Section 1: Where Embedding Sits in the RAG Pipeline

Diagram showing where embedding sits in the RAG ingestion pipeline: after parsing and chunking, before the vector index, query embedding, similarity search, reranking, and LLM generation

🏗️ RAG Ingestion Pipeline — Where Embedding Fits (Text Version)

🔍 Document Parsing (clean Markdown/JSON output)
⬇️
✂️ Chunking (clean, sized, self-contained text pieces)
⬇️
🧬 Embedding Layer (today's topic)
(Converting each chunk into a meaning-vector)
⬇️
🗂️
Vector Index (ANN)
🔎
Query Embedding
⬇️
🤖 Similarity Search → Reranking → LLM Generation
✅ Why This Matters: Embedding happens twice in every RAG query — once (up front, in bulk) for every chunk during ingestion, and once (live, per request) for the user's question at query time. Both must use the exact same embedding model, or the two vectors won't share the same coordinate system, and similarity search silently breaks.

🎓 Section 2: How Does a Model Learn to Do This? (Contrastive Learning)

Embedding models aren't hand-coded with a dictionary of "similar words." They learn the meaning-map through a training technique called contrastive learning — the same core idea behind almost every modern embedding model, from open-source BGE models to closed APIs.

🏋️ Contrastive Training, Step by Step

1. Feed the model triplets — an anchor text, a positive (a text that truly matches it in meaning, e.g. a question and its correct answer passage), and one or more negatives (unrelated or superficially similar-but-wrong texts).
↓
2. Embed all three using the model being trained — anchor, positive, and negatives each become a vector.
↓
3. Push the positive vector closer, and the negative vectors farther, from the anchor — adjusting the model's internal weights after every batch so the geometry improves slightly.
↓
4. Repeat across millions of triplets — search-click logs, question–answer pairs, translated sentence pairs, paraphrase datasets — until the geometry generalizes to text the model has never seen.
💡 Why "Negatives" Are the Hard Part

A weak negative (e.g. "banana" vs. "spaceship") teaches the model almost nothing — it's trivially easy to separate them. The real skill in building a great embedding model comes from mining hard negatives: text that looks superficially similar to the correct answer (same keywords, same topic) but is actually wrong. This is exactly what separates a mediocre embedding model from a state-of-the-art one.

🏗️ Section 3: Bi-Encoders vs. Cross-Encoders vs. Late Interaction

Not all "text similarity" models work the same way internally. Understanding the three architectural families explains why embedding models can search billions of documents in milliseconds, while a related-but-different model (the reranker) can't scale past a few hundred candidates.

🧬
Bi-Encoder

Embeds the query and each document completely separately — the two never see each other during embedding. This means every document's vector can be pre-computed once, stored, and reused for every future query. This is what makes it possible to search millions of vectors instantly. Almost every embedding model used in RAG is a bi-encoder.

🔬
Cross-Encoder

Feeds the query and document together, in one pass, letting the model directly compare every word of one against every word of the other. Far more accurate — but there's no vector to pre-compute, so it must run fresh for every query-document pair. This is what powers rerankers, applied to only the top 20–100 candidates a bi-encoder already narrowed down.

🧩
Late Interaction (e.g. ColBERT)

A middle ground: instead of one vector per document, it keeps one vector per token, and compares query tokens against document tokens at search time. More precise than a single-vector bi-encoder, cheaper than a full cross-encoder — at the cost of much more storage per document.

✅ The Production Pattern: Use a fast bi-encoder to retrieve the top 50–100 candidates from millions of chunks, then use a cross-encoder reranker to re-sort just those few candidates with far higher precision. This "retrieve cheap, rerank precise" two-stage design is the standard architecture across serious RAG stacks.

📐 Section 4: The Math — How "Similar" Gets Measured

Once two pieces of text are vectors, "how similar are they" becomes a geometry question with an exact numeric answer. Three metrics dominate RAG:

Cosine Similarity — measures the angle between two vectors, ignoring their length. Returns a value from -1 (opposite meaning) to 1 (identical meaning). This is the default metric for almost every embedding model, because it's insensitive to text length.
Dot Product — cosine similarity's faster cousin; it factors in vector magnitude too. If a model's vectors are already length-normalized during training (most modern ones are), dot product and cosine similarity give identical rankings — and dot product is cheaper to compute at scale, which is why most vector databases default to it.
Euclidean (L2) Distance — straight-line distance between two points. Less common for text, since it's sensitive to vector magnitude in ways that don't always track meaning — but some models are explicitly trained for it.
🚫 The Silent Bug: Using cosine similarity math on vectors from a model that was trained with dot product in mind (or vice versa) can quietly degrade retrieval quality. Always check your embedding model's documentation for its intended similarity metric — and configure your vector database to match it.

🪆 Section 5: Dimensions & Matryoshka Representation Learning

Every embedding model outputs a fixed number of dimensions — common sizes are 384, 768, 1024, or 3072. More dimensions can capture more nuance, but at a real cost: bigger vectors mean more storage, more memory for your approximate nearest-neighbor (ANN) index, and slightly higher search latency.

🪆 What Matryoshka Representation Learning (MRL) Solves

Named after Russian nesting dolls, MRL trains a model so that the first N numbers of its full-length vector are, on their own, already a usable, meaningful embedding — like a smaller doll nested inside a bigger one. This means you can truncate a 3072-dimension vector down to 1024 or even 256 dimensions after the fact, with only a small accuracy trade-off, instead of needing an entirely separate smaller model.

✅ Why This Is a Big Deal: MRL-trained models let you shrink vectors roughly 4x — cutting storage and ANN memory costs by a similar factor — while losing only a few points of retrieval accuracy. For most production RAG workloads, 768–1024 dimensions is the practical sweet spot: going much higher gives diminishing returns relative to the extra storage and latency it costs.

🔀 Section 6: Dense vs. Sparse vs. Hybrid Embeddings

So far we've described dense embeddings — every one of the hundreds of numbers in the vector is (usually) non-zero, densely packing meaning across the whole array. But dense vectors have a known weakness: they're good at fuzzy conceptual matches, but sometimes miss exact keyword, ID, or code-token matches that a human would consider an obvious hit.

🧬 Dense Vectors

Capture semantic meaning; great for paraphrases and conceptual questions ("How do I get my money back?" ↔ "refund policy"). Weak at rare exact terms like SKU numbers, product codes, or acronyms the model rarely saw in training.

🔑 Sparse Vectors

The modern, learned successor to keyword search (think BM25, but ML-weighted). Mostly zeros, with a few strongly-weighted terms — excellent at exact matches, poor at paraphrase and synonym understanding.

⚖️ Hybrid Search

Runs dense and sparse retrieval in parallel, then merges the ranked results (commonly via Reciprocal Rank Fusion). Combines conceptual understanding with exact-match precision — the standard architecture for serious production RAG systems.


🌍 Section 7: The Embedding Model Landscape

The embedding leaderboard has shifted meaningfully. The Massive Text Embedding Benchmark (MTEB) remains the closest thing to a standard reference, and open-source models had genuinely closed — and in places overtaken — the gap with closed APIs, while multimodal embedding (text + image + audio + video in one shared vector space) moved from research demo to production reality.

Model Type Notable For
Qwen3-Embedding-8B Open-source (Apache 2.0) Top-ranking multilingual MTEB score; strong self-hosting option
Gemini Embedding 2 API, multimodal Text, image, video, audio, and PDF in one shared vector space; native MRL
Cohere Embed v4 API, multimodal Early production-grade multimodal embedding for text + image + interleaved docs
OpenAI text-embedding-3-large API, text-only Reliable general-purpose default; weaker on some non-English languages
Voyage AI (legal / finance / code) API, domain-specialized Outperforms generic models notably within its specialized domain
BGE-M3 / GTE / E5 family Open-source Mature, well-documented, strong community fine-tuning ecosystem
💡 An Important Caveat

MTEB is text-retrieval-only — it doesn't test cross-modal search (text query against an image collection), cross-lingual retrieval quality, long-document behavior, or how much accuracy you lose when truncating dimensions via MRL. Public leaderboards are a useful starting filter, not a final answer — the only benchmark that truly matters is one built from your own documents and your own users' real questions.

🛠️ Section 8: When (and How) to Fine-Tune Your Own Embedding Model

Generic embedding models are trained on broad internet-scale text. That's great for general knowledge, but domain-specific language — legal clauses, medical terminology, internal product codenames, proprietary jargon — often sits outside what the model learned well.

📍 Signal #1: Domain-Specific Vocabulary

Your users' queries use internal terminology, abbreviations, or jargon a general-purpose model has rarely, if ever, seen in training.

📍 Signal #2: Measured Recall Gap

You've run the retrieval recall evaluation (Section 9) against a known-answer test set, and a generic off-the-shelf model consistently underperforms even after trying several candidates.

📍 The Practical Recipe

Start from a strong open-source base model (not from scratch), collect real query–passage pairs from your own logs or SMEs, mine hard negatives from your own corpus, and fine-tune with a contrastive loss (Section 2) for a few epochs — evaluating against your held-out test set after each pass.

✅ Best Practice: Don't fine-tune until a generic model has genuinely and measurably failed you. Fine-tuning is engineering effort and ongoing maintenance cost — most teams get further, faster, by first trying two or three strong generic models and only reaching for fine-tuning once the gap is proven, not assumed.

🧪 Section 9: How Do You Know If Your Embedding Model Is Actually Good?

🎯 Retrieval Recall@K

Build a test set of real questions with known correct source chunks. For each question, check whether the correct chunk appears in the top K (commonly K=5 or K=10) retrieved results. This is the single most important embedding-quality metric for RAG.

📊 MRR & nDCG

Mean Reciprocal Rank rewards getting the right answer near the top, not just "somewhere in the top K." Normalized Discounted Cumulative Gain (nDCG) generalizes this further when multiple results can be partially relevant, not just right-or-wrong.

🌐 Cross-Lingual & Cross-Modal Spot Checks

If your content spans languages or includes images/PDFs, test those cases explicitly — general text-only benchmarks like MTEB don't cover them, and quality can differ sharply between languages or modalities even within the same model.


💻 Section 10: Let's Build an Embedding + Retrieval Flow (With Code)

📌 What This Code Does (Read Before The Code!)

This shows the two moments embedding actually happens in RAG: once in bulk during ingestion (every chunk becomes a stored vector), and once per request at query time — followed by a cosine-similarity comparison against every stored vector to find the closest matches.

# Embedding + Cosine Similarity Retrieval (Pseudocode)

import numpy as np

# ── Step 1: Ingestion-time embedding (runs once per chunk) ──
def embed_and_store(chunks, embedding_model, vector_db):
    for chunk in chunks:
        vector = embedding_model.embed(chunk.text)   # e.g. 1024 floats
        vector_db.upsert(
            id=chunk.id,
            vector=vector,
            metadata={"text": chunk.text, "heading_path": chunk.heading_path}
        )

# ── Step 2: Query-time embedding (runs once per user question) ──
def retrieve(question, embedding_model, vector_db, top_k=5):
    # CRITICAL: must be the SAME model used during ingestion
    query_vector = embedding_model.embed(question)

    # Vector DB does this via an ANN index (e.g. HNSW) — shown here as brute force
    scored = []
    for item in vector_db.all():
        score = cosine_similarity(query_vector, item.vector)
        scored.append((score, item))

    scored.sort(reverse=True, key=lambda x: x[0])
    return scored[:top_k]

# ── The similarity math itself ──
def cosine_similarity(vec_a, vec_b):
    dot = np.dot(vec_a, vec_b)
    norm = np.linalg.norm(vec_a) * np.linalg.norm(vec_b)
    return dot / norm   # range: -1 (opposite) to 1 (identical meaning)
✅ Notice the Pattern: Ingestion-time embedding runs once, in bulk, offline — it can afford to be slow and thorough. Query-time embedding runs live, on every single user question — it must be fast. This asymmetry is exactly why bi-encoders (Section 3) dominate production RAG: the expensive part happens ahead of time, and the live part is just one lightweight embed call plus a vector lookup.

🛡️ Section 11: Common Embedding Pitfalls to Avoid

🔀 Mixing Embedding Models Between Ingestion and Query

Swapping to a "better" model for new documents while old chunks stay embedded with the previous model silently corrupts your vector space — the two sets of vectors don't share the same coordinate system. Any model change requires re-embedding the entire corpus.

📏 Forgetting Query/Passage Prefixes

Many embedding models are trained with different instruction prefixes for queries versus documents (e.g. "search_query: ..." vs. "search_document: ..."). Skipping this asymmetric prefixing on a model that expects it can quietly cut retrieval accuracy without throwing any error.

📐 Trusting MTEB Alone

A model topping the general leaderboard can still underperform on your specific domain, language mix, or document length. Always validate with your own retrieval recall test (Section 9) before committing.

💸 Over-Provisioning Dimensions

Defaulting to the largest available vector size "to be safe" multiplies storage, memory, and latency costs for gains that flatten out quickly past roughly 768–1024 dimensions for most tasks. Use MRL truncation (Section 5) to find your real sweet spot.


🎓 Section 12: Cheat Sheet — Choosing & Deploying Your Embedding Layer

Step 1: Understand Your Content & Users
  • Single language, multilingual, or cross-lingual queries?
  • Text-only, or does it include images, PDFs, or tables that need multimodal embedding?
  • How domain-specific is the vocabulary?
Step 2: Pick a Baseline Model (Section 7's Table)
  • Start with a strong generic model — open-source or API
  • Check MTEB as a starting filter, not a final verdict
  • Reach for domain-specialized models only where a real gap is proven
Step 3: Choose Dimensions & Metric
  • 768–1024 dimensions for most workloads; use MRL to compress further
  • Match your vector database's similarity metric to the model's training setup
Step 4: Consider Hybrid Search
  • Add sparse retrieval alongside dense for exact-match precision (Section 6)
  • Add a cross-encoder reranker as a second stage (Section 3)
Step 5: Measure, Don't Assume (Section 9)
  • Build a real recall@K test set from your own domain
  • Track MRR / nDCG, not just top-1 accuracy
  • Re-validate after every model or chunking change

❓ Frequently Asked Questions

What is an embedding in RAG?

A fixed-length list of numbers representing the meaning of a piece of text, positioned so that texts with similar meaning end up close together in vector space, regardless of how differently they're worded.

Why are almost all embedding models bi-encoders?

Because a bi-encoder embeds documents independently of any query, every document's vector can be pre-computed once and reused for every future search — the only architecture that makes searching millions or billions of chunks in milliseconds possible.

How many dimensions should my embedding vectors have?

768–1024 dimensions is the practical sweet spot for most production RAG workloads (Section 5). Matryoshka Representation Learning lets you compress higher-dimension vectors down to this range with minimal accuracy loss.

Should I use dense, sparse, or hybrid search?

Hybrid search (Section 6) is the standard for serious production systems — it combines dense embeddings' conceptual understanding with sparse retrieval's exact-match precision for IDs, codes, and rare terms.

Do I need to fine-tune my own embedding model?

Only after a measured recall gap (Section 8) — don't fine-tune until two or three strong generic models have genuinely and measurably underperformed on your own retrieval recall test set.


🎉 Final Summary

🧬 An embedding is a coordinate — it places meaning in space so similar ideas end up physically close, no matter how they're worded
🎓 Contrastive learning with hard negatives is how models learn this meaning-map, from millions of anchor–positive–negative triplets
🏗️ Bi-encoders pre-compute document vectors independently of the query, which is exactly what makes billion-scale retrieval possible
📐 Cosine similarity and dot product turn "how related are these two texts" into simple, fast vector math
🪆 Matryoshka Representation Learning lets you compress vectors 4x with minimal accuracy loss — a real shift in RAG economics
🔀 Hybrid search (dense + sparse) combines conceptual understanding with exact-match precision
🌍 Open-source models like Qwen3-Embedding now compete directly with — and sometimes beat — closed APIs on retrieval quality
🧪 Always measure retrieval recall against your own domain — public benchmarks are a starting filter, not a final verdict
✅ The Core Lesson:

Embedding is not a commodity checkbox you pick once and forget — it's the coordinate system your entire RAG system reasons within. Every retrieval failure, every "the answer was in the docs but the bot missed it" complaint, traces back either to chunking or to embedding. Treat model choice, dimension size, similarity metric, and evaluation as first-class architecture decisions — not defaults you inherited from a tutorial — and your retrieval quality will show it.


Happy Building! Embed Meaning, Retrieve Precisely. 🔥

Comments