Skip to main content

Re-Ranking

Calculating read time…

Re-ranking in RAG is a second-pass retrieval step that uses a cross-encoder model to re-score and reorder the chunks a vector database already retrieved — putting the genuinely most relevant results first before they ever reach the LLM. Here's exactly why vector search alone isn't enough to guarantee that, and how re-ranking closes the gap.

In this series so far, we've parsed messy documents into clean text, chunked that text into small pieces, and turned each piece into a meaning-vector through embedding. Now imagine your vector database just retrieved the top 20 chunks for a user's question by comparing those vectors. Job done, right?

Not quite. Vector search is fast but approximate — it's built to scan millions of chunks in milliseconds, which means it trades away some accuracy to get that speed. Often the truly best chunk is sitting at position #7 or #12 in that list, not #1.

Re-ranking is the step that fixes this — a second, smarter, slower pass that re-examines those top candidates and puts the genuinely best ones first. Let's understand exactly how it works, from first principles.

Re-ranking in RAG — vector search versus cross-encoder reranker pipeline diagram

🔍 Vector search retrieves candidates; re-ranking finds the right ones among them
⚡ Vector search compares against millions of chunks in milliseconds; re-rankers compare against just the top 20–100
🎯 Adding a reranker to an existing RAG pipeline is one of the highest ROI upgrades — often a bigger accuracy jump than switching embedding models
🧠 The whole trick: re-rankers read the question and the chunk together, while vector search reads them separately.

Let's see exactly why that difference matters so much.

🕵️ Section 1: Why Vector Search Alone Isn't Enough

💡 The Resume Screening Analogy

Imagine an HR system with 10,000 resumes. Round 1: a keyword scanner quickly shortlists 50 resumes that roughly mention "Python" and "5 years experience" — fast, but crude. It might rank a resume that just lists "Python" in a skills section above one that describes 6 years of deep, relevant Python project work, just because of how the words happen to appear.

Round 2: a senior hiring manager actually reads all 50 shortlisted resumes carefully, comparing each one directly against the actual job requirements, and re-orders them by true fit. That's re-ranking — a slower, smarter second pass over a small, already-narrowed list.

Vector search (from our earlier posts: converting text into meaning-vectors and comparing them) plays the role of Round 1. It's remarkably good at fast, approximate filtering — but "approximate" is the key word. It embeds the question and the chunk completely separately, then just measures how close their two vectors are. That separation is exactly what re-ranking fixes.


🗺️ Section 2: Where Re-Ranking Sits in the RAG Pipeline

🏗️ Full RAG Pipeline — Where Re-Ranking Fits

📄
PDF
🖼️
Scans
📊
Slides
📝
Word/HTML
⬇️
🔍 Document Parsing — raw files → clean Markdown/JSON
⬇️
✂️ Chunking — clean text → small, self-contained pieces
⬇️
🧬 Embedding — every chunk → a meaning-vector, stored in a vector index
⬇️
❓ User's Question
⬇️
🗂️ Vector Search — retrieves Top 50 candidate chunks (fast, approximate)
⬇️
🎯 Re-Ranker (today's topic)
(Reads question + chunk together, scores each precisely, re-sorts)
⬇️
🤖 Top 3–5 Re-Ranked Chunks → LLM Generation

Parsing, chunking, and embedding built the map. Re-ranking is what makes sure we trust the right pin on it before showing it to the LLM.

✅ Why Fewer, Better Chunks Matter: Feeding an LLM 50 loosely relevant chunks wastes context budget and can actually confuse the model with noise. Re-ranking narrows those 50 down to the 3–5 chunks that are genuinely the most useful — a cleaner, cheaper, more accurate final prompt.

⚙️ Section 3: The Core Idea — Bi-Encoders vs. Cross-Encoders

This single distinction explains almost everything about how re-ranking works.

🔀
Bi-Encoder (what your embedding model is)

Encodes the question and each chunk completely separately into two independent vectors, then compares them with simple math (cosine similarity). Fast — the chunk vectors can even be pre-computed and stored ahead of time. But the model never actually "sees" the question and chunk side-by-side.

🔗
Cross-Encoder (what a re-ranker is)

Feeds the question and the chunk into the model together, at the same time, and the model directly outputs one relevance score. It can notice fine-grained interactions between specific words in the question and specific words in the chunk — far more accurate, but must be run fresh for every question/chunk pair, so it can't be pre-computed.

📍 A Concrete Example of the Difference

Question: "What voltage does the device require?" — Chunk: "The device operates at 220V but should not exceed 240V."

A bi-encoder sees these as two separate vectors that happen to be about similar topics (voltage, device) and gives a decent-but-generic similarity score. A cross-encoder reads both together and can directly notice: the question asks "what voltage," and this chunk explicitly states a voltage value — producing a sharper, more confident relevance score.

🚫 Why We Don't Just Use Cross-Encoders for Everything: A cross-encoder must run once per question/chunk pair — comparing a question against 1 million chunks would mean 1 million model runs per query, far too slow. Bi-encoders (vector search) exist precisely to cheaply narrow 1 million down to ~50 first, so the expensive cross-encoder only has to score that small shortlist.

🎓 Section 3.1: How Does a Cross-Encoder Actually Learn to Score Relevance?

It's not enough to know a cross-encoder "reads both together" — let's open the hood on how it learns to turn that joint reading into a trustworthy number, because this explains both its strengths and its blind spots.

🧬 Training a Cross-Encoder, Step by Step

1. Start from a language model (typically a BERT-family encoder) that already understands grammar and word relationships from general pretraining.
↓
2. Feed it labeled (query, passage, relevant?) triples — real search logs or human-annotated datasets where each pair is marked relevant or not relevant.
↓
3. The model outputs a single number (via one small classification layer on top of the joint encoding) and is penalized when that number disagrees with the true label — nudged up for true relevant pairs, down for true irrelevant ones.
↓
4. Repeated across millions of triples, the model learns general patterns of relevance — not just memorized examples — including subtle cues like negation ("does NOT support X") or numeric/unit matching that a bi-encoder's separate-encoding approach tends to miss.
🚫 The Blind Spot This Creates: Because training data rarely includes many domain-specific triples (say, aerospace maintenance manuals or pharmaceutical labeling), an off-the-shelf reranker can under-perform on narrow technical domains until it's fine-tuned on in-domain (query, passage, relevant?) examples — the same validation discipline we recommended for embedding models in our previous post applies here too.

📋 Section 3.2: Pointwise vs. Listwise Re-Ranking

There's a second, less-discussed design choice inside re-ranking: does the model score one chunk at a time, or does it look at the whole shortlist together before deciding an order?

🎯 Pointwise Scoring

The classic cross-encoder approach described above — every (question, chunk) pair is scored completely independently of every other candidate. Simple, parallelizable, and what most rerankers (Cohere Rerank, BGE-Reranker) do by default.

📊 Listwise Scoring

The model (usually an LLM prompted for this specific job) sees all shortlisted candidates at once and directly outputs a ranked order — allowing it to reason about candidates relative to each other ("chunk A answers the question more completely than chunk B, even though B mentions more matching keywords").

✅ Why Listwise Can Win on Nuance: Imagine two chunks that both mention "refund policy" — one gives the full policy, the other only references it in passing. Scored independently (pointwise), both might get a similar relevance score. Seen together (listwise), it becomes obvious which one is the clearly better answer — this relative judgment is exactly what listwise scoring is built to capture.
🚫 The Trade-Off: Listwise scoring needs a much larger prompt (all candidates at once) and can't be parallelized the way pointwise scoring can — meaningfully slower and pricier per query, so it's typically reserved for the final handful of candidates, not the full 50-item shortlist.

🧭 Section 4: Step-by-Step — How a Re-Ranker Processes a Query

1
Vector search returns the Top-K shortlist
Typically 20–100 candidate chunks — fast, approximate, cast a wide net.
⬇️
2
Each (question, chunk) pair is fed to the cross-encoder
The reranker model receives both texts joined together and outputs a single relevance score — often between 0 and 1 — for that exact pair.
⬇️
3
All candidates are re-sorted by this new score
The original vector-search order is discarded — the cross-encoder's score is now the authority on ranking.
⬇️
4
Only the top 3–5 re-ranked chunks are sent to the LLM
A dramatically smaller, higher-precision set than what vector search alone would have handed over.

☁️
Cohere Rerank

A managed API — send a query and a list of candidate texts, get back relevance scores. Widely used because it's simple to integrate and needs no self-hosting.

🧩
Open-Source Cross-Encoders (BGE-Reranker, MS MARCO MiniLM, mxbai-rerank)

Self-hostable models you run on your own infrastructure — good when data can't leave your environment, or to avoid per-call API costs at scale.

🧮
ColBERT-Style "Late Interaction" Rerankers

A clever middle ground: instead of one vector per chunk (bi-encoder) or full joint encoding (cross-encoder), it keeps a vector per token and matches question-tokens against chunk-tokens directly — much of a cross-encoder's precision, closer to a bi-encoder's speed.

🤖
LLM-Based Reranking

Directly prompting an LLM to rate or rank each candidate chunk's relevance. Highest reasoning quality (can weigh subtle context), but the slowest and most expensive option — used selectively for the final few candidates, not the whole shortlist.


🏗️ Section 5.1: Multi-Stage Cascades — Combining Speed and Precision

Production systems rarely stop at a single reranking pass. Instead, they often chain multiple stages, each narrower and more expensive than the last — the same funnel principle from Section 2, just extended one level further.

🔺 A Typical Three-Stage Cascade

Stage 1 — Vector/Hybrid Search: 1M+ chunks → Top 100. Cheapest, fastest, widest net.
Stage 2 — Lightweight Pointwise Cross-Encoder: Top 100 → Top 10. Fast enough to run on every candidate, precise enough to cut 90% of the noise.
Stage 3 — Listwise LLM Reranking: Top 10 → Top 3–5. The most expensive stage, but now only running on a genuinely small, already-strong set of candidates.
✅ Why Cascade Instead of Jumping Straight to the Best Model: Running the most powerful (and most expensive) technique — listwise LLM reranking — directly on 100 or 1,000 candidates would be prohibitively slow and costly. Cascading lets each stage "pay" only for the volume of candidates it can handle economically, while the final, priciest stage only ever touches a handful of already-strong finalists.

💻 Section 6: A Simple Re-Ranking Pipeline (With Code)

📌 What This Code Does (Read Before The Code!)

This shows the two-stage retrieval pattern from Section 4: fetch a wide candidate list cheaply with vector search, then narrow it down precisely with a cross-encoder reranker, and only pass the final few chunks to the LLM.

# Two-Stage Retrieval: Vector Search + Re-Ranking (Pseudocode)

def retrieve_and_rerank(user_question, top_k_vector=50, top_n_final=5):

    # Stage 1: Cheap, fast bi-encoder vector search over the whole index
    question_vector = embedding_model.encode(user_question)
    candidates = vector_db.search(question_vector, top_k=top_k_vector)

    # Stage 2: Expensive, precise cross-encoder scoring — but only
    # on this small shortlist, never the whole database
    scored = []
    for chunk in candidates:
        relevance_score = reranker_model.score(
            query=user_question,
            document=chunk.text   # fed together, not separately!
        )
        scored.append((chunk, relevance_score))

    # Stage 3: Sort by the new, more trustworthy score
    scored.sort(key=lambda pair: pair[1], reverse=True)

    # Only the best few chunks make it into the final LLM prompt
    return [ chunk for chunk, score in scored[:top_n_final] ]
✅ Notice the Pattern: The expensive model (reranker) only ever touches top_k_vector candidates (50), never the full database (which could be millions). This "cheap filter, then expensive precision pass" funnel shape is the same design principle we saw in the parsing router and the chunking fallback in earlier posts.

🧪 Section 7: How to Know If Re-Ranking Is Actually Helping

📈 NDCG / MRR on a Labeled Test Set

With a set of questions and known correct chunks, measure whether the correct chunk moves closer to position #1 after re-ranking compared to raw vector search order — standard information-retrieval metrics like NDCG and Mean Reciprocal Rank quantify exactly this.

⏱️ Latency Budget Check

Re-ranking adds real latency — measure the added milliseconds per query and confirm it still fits your response-time requirements, especially as you tune top_k_vector.

🎚️ Score Calibration

A reranker's raw scores aren't automatically comparable across different questions — a "0.6" for one query might represent strong relevance, while a "0.6" for another barely clears the noise floor. If your application needs an absolute relevance threshold (e.g. "don't answer if nothing scores above X"), calibrate that threshold empirically against a labeled test set rather than picking a number that merely feels reasonable.

🚫 A Common Pitfall: Setting top_k_vector too small (say, 5) defeats the purpose — if the truly best chunk wasn't even in vector search's initial shortlist, no reranker can rescue it. Give vector search a wide enough net (typically 20–100) before handing off to the reranker.

❓ Frequently Asked Questions

What is re-ranking in a RAG pipeline?

Re-ranking is a second retrieval pass that takes the top candidate chunks returned by vector search and re-scores them with a more precise model — typically a cross-encoder — that reads the question and each chunk together, then reorders the list before it's sent to the LLM.

What's the difference between a bi-encoder and a cross-encoder?

A bi-encoder (used in vector search) encodes the question and each chunk separately into independent vectors and compares them mathematically. A cross-encoder (used in re-ranking) reads the question and chunk together in a single pass, letting it notice fine-grained interactions between them — more accurate, but too slow to run against an entire database.

Why not just use a cross-encoder for retrieval instead of vector search?

A cross-encoder has to run once per question/chunk pair, so comparing a question against a million chunks would mean a million model runs. Vector search cheaply narrows that down to a small shortlist first, so the expensive cross-encoder only has to score a manageable number of candidates.

What's the difference between pointwise and listwise re-ranking?

Pointwise re-ranking scores each candidate chunk independently of the others. Listwise re-ranking looks at all shortlisted candidates together and outputs a relative ranking, which can catch nuance pointwise scoring misses, at a higher computational cost.

How many candidates should vector search return before re-ranking?

Typically 20 to 100. Too few defeats the purpose — if the truly best chunk isn't even in the initial shortlist, no reranker can recover it — so vector search needs a wide enough net before the reranker narrows it down.


🎉 Final Summary

🕵️ Vector search retrieves candidates; re-ranking finds the right ones among them — a fast, approximate first pass followed by a slower, precise second pass
⚙️ Bi-encoders (vector search) encode question and chunk separately; cross-encoders (re-rankers) read them together — that togetherness is the entire source of the accuracy gain
🧭 The pipeline is a funnel: cast a wide net with cheap vector search, then apply the expensive reranker only to that small shortlist
📋 Pointwise scoring judges each chunk alone; listwise scoring compares candidates against each other — listwise catches relative nuance that pointwise misses, at higher cost
🏗️ Production systems often cascade multiple stages — cheap vector search, then a fast cross-encoder, then an expensive listwise pass — each narrower than the last
🧰 Options range from managed APIs (Cohere Rerank) to self-hosted cross-encoders, ColBERT-style late-interaction models, and LLM-based reranking
🧪 Always measure with NDCG/MRR, latency budgets, and calibrated score thresholds — and make sure vector search's initial net is wide enough for the reranker to have good candidates to work with
✅ The Core Lesson:

Vector search and re-ranking aren't competitors — they're a team, each covering the other's weakness. Vector search brings speed at scale; re-ranking brings precision on a shortlist. Skip re-ranking, and you're trusting a fast approximation to make your final decision. Add it, and you're letting a fast system propose, and a smart system decide.


Happy Building! Retrieve Wide, Rank Precisely. 🔥

Comments