Re-ranking in RAG is a second-pass retrieval step that uses a cross-encoder model to re-score and reorder the chunks a vector database already retrieved — putting the genuinely most relevant results first before they ever reach the LLM. Here's exactly why vector search alone isn't enough to guarantee that, and how re-ranking closes the gap.
In this series so far, we've parsed messy documents into clean text, chunked that text into small pieces, and turned each piece into a meaning-vector through embedding. Now imagine your vector database just retrieved the top 20 chunks for a user's question by comparing those vectors. Job done, right?
Not quite. Vector search is fast but approximate — it's built to scan millions of chunks in milliseconds, which means it trades away some accuracy to get that speed. Often the truly best chunk is sitting at position #7 or #12 in that list, not #1.
Re-ranking is the step that fixes this — a second, smarter, slower pass that re-examines those top candidates and puts the genuinely best ones first. Let's understand exactly how it works, from first principles.
⚡ Vector search compares against millions of chunks in milliseconds; re-rankers compare against just the top 20–100
🎯 Adding a reranker to an existing RAG pipeline is one of the highest ROI upgrades — often a bigger accuracy jump than switching embedding models
🧠 The whole trick: re-rankers read the question and the chunk together, while vector search reads them separately.
Let's see exactly why that difference matters so much.
- Why Vector Search Alone Isn't Enough
- Where Re-Ranking Sits in the RAG Pipeline
- Bi-Encoders vs. Cross-Encoders
- How Cross-Encoders Are Actually Trained
- Pointwise vs. Listwise Re-Ranking
- Step-by-Step: How a Re-Ranker Processes a Query
- Popular Re-Ranking Models
- Multi-Stage Cascades
- Code: A Re-Ranking Pipeline
- How to Know If Re-Ranking Is Helping
- FAQ
🕵️ Section 1: Why Vector Search Alone Isn't Enough
Imagine an HR system with 10,000 resumes. Round 1: a keyword scanner quickly shortlists 50 resumes that roughly mention "Python" and "5 years experience" — fast, but crude. It might rank a resume that just lists "Python" in a skills section above one that describes 6 years of deep, relevant Python project work, just because of how the words happen to appear.
Round 2: a senior hiring manager actually reads all 50 shortlisted resumes carefully, comparing each one directly against the actual job requirements, and re-orders them by true fit. That's re-ranking — a slower, smarter second pass over a small, already-narrowed list.
Vector search (from our earlier posts: converting text into meaning-vectors and comparing them) plays the role of Round 1. It's remarkably good at fast, approximate filtering — but "approximate" is the key word. It embeds the question and the chunk completely separately, then just measures how close their two vectors are. That separation is exactly what re-ranking fixes.
🗺️ Section 2: Where Re-Ranking Sits in the RAG Pipeline
🏗️ Full RAG Pipeline — Where Re-Ranking Fits
Scans
Slides
Word/HTML
(Reads question + chunk together, scores each precisely, re-sorts)
Parsing, chunking, and embedding built the map. Re-ranking is what makes sure we trust the right pin on it before showing it to the LLM.
⚙️ Section 3: The Core Idea — Bi-Encoders vs. Cross-Encoders
This single distinction explains almost everything about how re-ranking works.
Encodes the question and each chunk completely separately into two independent vectors, then compares them with simple math (cosine similarity). Fast — the chunk vectors can even be pre-computed and stored ahead of time. But the model never actually "sees" the question and chunk side-by-side.
Feeds the question and the chunk into the model together, at the same time, and the model directly outputs one relevance score. It can notice fine-grained interactions between specific words in the question and specific words in the chunk — far more accurate, but must be run fresh for every question/chunk pair, so it can't be pre-computed.
📍 A Concrete Example of the Difference
Question: "What voltage does the device require?" — Chunk: "The
device operates at 220V but should not exceed 240V."
A bi-encoder sees these as two separate vectors that happen to be about
similar topics (voltage, device) and gives a decent-but-generic similarity
score. A cross-encoder reads both together and can directly notice: the
question asks "what voltage," and this chunk explicitly states a voltage
value — producing a sharper, more confident relevance score.
🎓 Section 3.1: How Does a Cross-Encoder Actually Learn to Score Relevance?
It's not enough to know a cross-encoder "reads both together" — let's open the hood on how it learns to turn that joint reading into a trustworthy number, because this explains both its strengths and its blind spots.
🧬 Training a Cross-Encoder, Step by Step
📋 Section 3.2: Pointwise vs. Listwise Re-Ranking
There's a second, less-discussed design choice inside re-ranking: does the model score one chunk at a time, or does it look at the whole shortlist together before deciding an order?
The classic cross-encoder approach described above — every (question, chunk) pair is scored completely independently of every other candidate. Simple, parallelizable, and what most rerankers (Cohere Rerank, BGE-Reranker) do by default.
The model (usually an LLM prompted for this specific job) sees all shortlisted candidates at once and directly outputs a ranked order — allowing it to reason about candidates relative to each other ("chunk A answers the question more completely than chunk B, even though B mentions more matching keywords").
🧭 Section 4: Step-by-Step — How a Re-Ranker Processes a Query
Typically 20–100 candidate chunks — fast, approximate, cast a wide net.
The reranker model receives both texts joined together and outputs a single relevance score — often between 0 and 1 — for that exact pair.
The original vector-search order is discarded — the cross-encoder's score is now the authority on ranking.
A dramatically smaller, higher-precision set than what vector search alone would have handed over.
🧰 Section 5: Popular Re-Ranking Models in 2026
A managed API — send a query and a list of candidate texts, get back relevance scores. Widely used because it's simple to integrate and needs no self-hosting.
Self-hostable models you run on your own infrastructure — good when data can't leave your environment, or to avoid per-call API costs at scale.
A clever middle ground: instead of one vector per chunk (bi-encoder) or full joint encoding (cross-encoder), it keeps a vector per token and matches question-tokens against chunk-tokens directly — much of a cross-encoder's precision, closer to a bi-encoder's speed.
Directly prompting an LLM to rate or rank each candidate chunk's relevance. Highest reasoning quality (can weigh subtle context), but the slowest and most expensive option — used selectively for the final few candidates, not the whole shortlist.
🏗️ Section 5.1: Multi-Stage Cascades — Combining Speed and Precision
Production systems rarely stop at a single reranking pass. Instead, they often chain multiple stages, each narrower and more expensive than the last — the same funnel principle from Section 2, just extended one level further.
🔺 A Typical Three-Stage Cascade
💻 Section 6: A Simple Re-Ranking Pipeline (With Code)
This shows the two-stage retrieval pattern from Section 4: fetch a wide candidate list cheaply with vector search, then narrow it down precisely with a cross-encoder reranker, and only pass the final few chunks to the LLM.
# Two-Stage Retrieval: Vector Search + Re-Ranking (Pseudocode) def retrieve_and_rerank(user_question, top_k_vector=50, top_n_final=5): # Stage 1: Cheap, fast bi-encoder vector search over the whole index question_vector = embedding_model.encode(user_question) candidates = vector_db.search(question_vector, top_k=top_k_vector) # Stage 2: Expensive, precise cross-encoder scoring — but only # on this small shortlist, never the whole database scored = [] for chunk in candidates: relevance_score = reranker_model.score( query=user_question, document=chunk.text # fed together, not separately! ) scored.append((chunk, relevance_score)) # Stage 3: Sort by the new, more trustworthy score scored.sort(key=lambda pair: pair[1], reverse=True) # Only the best few chunks make it into the final LLM prompt return [ chunk for chunk, score in scored[:top_n_final] ]
top_k_vector
candidates (50), never the full database (which could be millions). This
"cheap filter, then expensive precision pass" funnel shape is the same design
principle we saw in the parsing router and the chunking fallback in earlier posts.
🧪 Section 7: How to Know If Re-Ranking Is Actually Helping
With a set of questions and known correct chunks, measure whether the correct chunk moves closer to position #1 after re-ranking compared to raw vector search order — standard information-retrieval metrics like NDCG and Mean Reciprocal Rank quantify exactly this.
Re-ranking adds real latency — measure the added milliseconds per query
and confirm it still fits your response-time requirements, especially
as you tune top_k_vector.
A reranker's raw scores aren't automatically comparable across different questions — a "0.6" for one query might represent strong relevance, while a "0.6" for another barely clears the noise floor. If your application needs an absolute relevance threshold (e.g. "don't answer if nothing scores above X"), calibrate that threshold empirically against a labeled test set rather than picking a number that merely feels reasonable.
top_k_vector
too small (say, 5) defeats the purpose — if the truly best chunk wasn't even
in vector search's initial shortlist, no reranker can rescue it. Give vector
search a wide enough net (typically 20–100) before handing off to the reranker.
❓ Frequently Asked Questions
Re-ranking is a second retrieval pass that takes the top candidate chunks returned by vector search and re-scores them with a more precise model — typically a cross-encoder — that reads the question and each chunk together, then reorders the list before it's sent to the LLM.
A bi-encoder (used in vector search) encodes the question and each chunk separately into independent vectors and compares them mathematically. A cross-encoder (used in re-ranking) reads the question and chunk together in a single pass, letting it notice fine-grained interactions between them — more accurate, but too slow to run against an entire database.
A cross-encoder has to run once per question/chunk pair, so comparing a question against a million chunks would mean a million model runs. Vector search cheaply narrows that down to a small shortlist first, so the expensive cross-encoder only has to score a manageable number of candidates.
Pointwise re-ranking scores each candidate chunk independently of the others. Listwise re-ranking looks at all shortlisted candidates together and outputs a relative ranking, which can catch nuance pointwise scoring misses, at a higher computational cost.
Typically 20 to 100. Too few defeats the purpose — if the truly best chunk isn't even in the initial shortlist, no reranker can recover it — so vector search needs a wide enough net before the reranker narrows it down.
🎉 Final Summary
Vector search and re-ranking aren't competitors — they're a team, each covering the other's weakness. Vector search brings speed at scale; re-ranking brings precision on a shortlist. Skip re-ranking, and you're trusting a fast approximation to make your final decision. Add it, and you're letting a fast system propose, and a smart system decide.
Happy Building! Retrieve Wide, Rank Precisely. 🔥
Comments
Post a Comment