Evaluating LLMs and prompt templates for RAG means testing which model and which prompt structure, paired together, produce the most grounded, accurate answers from your retrieved context — a distinct, often-skipped decision separate from evaluating retrieval itself. Here's exactly why teams get this wrong, and how to get it right with evidence, not guesswork.
Here's a scene that plays out in almost every RAG project: the team spends weeks perfecting parsing, chunking, embeddings, and re-ranking. Then, at the very last mile, someone just... picks a model. "Let's use GPT-whatever, it's popular." They write a prompt in five minutes. Ship it. Three weeks later, users start complaining the bot "makes stuff up" or "ignores half the document." Nobody thought to actually test this last, tiny-seeming decision — and it turns out to be one of the biggest levers in the entire system.
This post is not about full RAG evaluation — we're not measuring retrieval recall or chunk quality here, that's a much bigger topic for another day. We're staying laser-focused on one narrow, practical, often-skipped question: once you have good retrieved context in hand, which LLM should generate the answer, and with exactly what prompt template? By the end, you'll have a repeatable, evidence-based way to answer that — no guessing, no "it feels better" hand-waving.
🧩 Two variables are secretly tangled together here: the model and the prompt template
⚖️ A single sentence added to a prompt template can lift accuracy more than switching to a pricier model
🧪 A tiny, carefully-built test set of 30 examples reveals more truth than 300 randomly grabbed ones.
Let's slow down and build real intuition for why this "last mile" decision matters so much.
- The Recipe Analogy: Why "Which LLM?" Isn't the Whole Question
- Scope: This Is Not Full RAG Evaluation
- The Two Tangled Variables: Model and Template
- What Does "Hallucination" Actually Mean?
- Building a Small, Honest Golden Test Set
- What "Evaluating an LLM" Actually Means for RAG
- What "Evaluating a Prompt Template" Actually Means
- Combining Both Variables Into One Matrix
- How Do You Actually Produce Those Scores?
- Code: A Simple Evaluation Harness
- FAQ
- Pitfalls
- Cheat Sheet
🍳 Section 0: The Recipe Analogy — Why "Which LLM?" Isn't the Whole Question
Imagine you've done all the hard work: you've bought the freshest ingredients (great retrieval), you've prepped and portioned everything perfectly (great chunking), and you've organized your pantry so anything can be found instantly (great embeddings and re-ranking). Now you hand those ingredients to a chef and say "make dinner."
A brilliant chef handed a vague recipe card ("make something with these") might improvise something inconsistent — sometimes great, sometimes odd. A mediocre cook handed a precise, well-written recipe card (exact steps, exact quantities, exact plating instructions) often produces a more consistent, reliable dish every single time.
The LLM is the chef. The prompt template is the recipe card. Picking "the best chef" without also fixing "the best recipe card" is only half the decision — and it's the half most teams stop at.
This is why teams are so often surprised when a "top-tier" model still gives wrong or ungrounded answers — they never tested whether the recipe card (prompt template) they were using was actually any good.
🚧 Section 1: Drawing the Boundary — This Is Not Full RAG Evaluation
Let's place this precisely in the pipeline we've been building across this series, so it's clear exactly what we're deciding here versus what's already been handled by earlier stages.
🗺️ Where This Decision Lives in the Pipeline
— all covered earlier in this series; assume these are already solid —
Given good context, which model + which recipe card turns it into the best answer?
🧩 Section 2: The Two Tangled Variables — Model and Template
Its raw ability to reason, follow instructions carefully, stay honest about what the context actually says, and produce well-formatted output.
How the instructions are structured, where the retrieved context sits, whether examples are included, and what output format is demanded.
Here's a real, concrete example of the same model producing two very different answers, purely because of the recipe card it was handed:
Context: {context}
Question: {question}
Answer the question.
Result: model confidently adds outside knowledge the context never mentioned, because nothing told it not to.
Context: {context}
Question: {question}
Answer using ONLY the context above.
If the answer isn't there, say
"I don't have enough information."
Result: same model, dramatically fewer confident wrong answers.
👻 Section 3: Quick Definition — What Does "Hallucination" Actually Mean Here?
We're about to use this word a lot, so let's nail it down precisely, since beginners often hear it thrown around vaguely.
🔍 Hallucination, in Plain English
A hallucination is when an LLM states something confidently that isn't actually supported by the context it was given — it's not "lying" on purpose; the model is simply pattern-completing text, and sometimes the most "natural-sounding" continuation happens to be false or invented. In RAG specifically, the entire point of retrieval was to give the model real facts to stand on — a hallucination means the model wandered off that ground and made something up instead.
This is exactly why groundedness — sticking strictly to the provided context — becomes the single most important thing we're about to measure in both the model and the template.
🏅 Section 4: Building a Small, Honest "Golden" Test Set
Before comparing anything, you need fixed, realistic test cases — otherwise "Model A felt better than Model B" is just an opinion, not evidence. Each test case is a simple triple:
{
"question": "What's the refund window for unopened items?",
"retrieved_context": "...(the actual chunks your pipeline would return)...",
"ideal_answer": "30 days from the purchase date, for unused items only."
}
A genuinely useful golden set doesn't just contain easy questions — it deliberately includes the situations where models tend to trip:
The answer requires combining facts from two or three different retrieved chunks, not just one.
Prices, dates, and quantities — easy to accidentally invent or mix up.
Where the retrieved context simply doesn't contain the answer — the correct response is an honest "I don't know," and this is the single best way to catch hallucination directly.
🤖 Section 5: What "Evaluating an LLM" Actually Means for RAG
Ignore general leaderboards for a moment — the ones ranking models on broad trivia, coding, or general chat quality. Those measure a completely different skill. Being generally brilliant doesn't automatically mean a model is good at the very specific job of "read this handful of paragraphs, and only answer from what's actually written there." Score each candidate model against four practical criteria instead:
Does the model stick strictly to the provided chunks, or does it quietly reach for outside knowledge it learned during training? This is the single most important trait for RAG — a model that "sounds smart" but ignores the context defeats the entire point of retrieval.
When your "unanswerable question" test cases (Section 4) come up, does the model honestly say it doesn't know, or does it confidently invent something plausible-sounding? This single behavior separates a trustworthy assistant from a dangerous one.
If you asked for JSON, three bullet points, or a specific citation style — did it comply exactly, every single time? Some models are noticeably more reliable at rigid formatting than others, which matters enormously if your output feeds another system downstream.
A model that's marginally more accurate but costs five times more and answers three times slower may not be worth it for a high-traffic support chatbot, even if it wins every quality metric on paper.
📝 Section 6: What "Evaluating a Prompt Template" Actually Means
Just like the model, a prompt template has structural choices that measurably swing output quality — test these deliberately, one at a time.
Does the retrieved context come before or after the instructions? Many models pay closer attention to text placed nearer the actual question — worth testing directly rather than assuming one order is obviously right.
A template that plainly says "If the answer isn't in the context, say you don't know" measurably reduces confident wrong answers compared to one that doesn't bother saying this at all — this single line is one of the highest-leverage additions you can test, exactly as shown in Section 2's side-by-side example.
Including 1–2 example (context, question, ideal-answer-format) demonstrations directly in the prompt often stabilizes formatting and tone — at the cost of extra tokens and latency on every single call.
For multi-chunk or numeric questions (Section 4), explicitly asking the model to reason through the pieces before giving a final answer can genuinely improve accuracy — but it adds latency, and can leak unwanted "thinking out loud" text into the final output if you don't clearly separate reasoning from the answer.
🧮 Section 7: Combining Both Variables Into One Matrix
Now bring the model (Section 5) and the template (Section 6) together in one grid — every candidate LLM, run against every candidate template, scored against the exact same golden test set from Section 4.
| Model ↓ / Template → | Template A (minimal) | Template B (+ "don't guess") | Template C (+ few-shot) |
|---|---|---|---|
| Model X | Groundedness: 72% | Groundedness: 89% | Groundedness: 91% |
| Model Y | Groundedness: 81% | Groundedness: 93% | Groundedness: 94% |
| Model Z (cheaper) | Groundedness: 65% | Groundedness: 84% | Groundedness: 88% |
🧪 Section 8: How Do You Actually Produce Those Scores?
Prompt a separate, strong LLM to compare each generated answer against the ideal answer and the original context, scoring groundedness and correctness on a simple rubric. It's fast enough to run across an entire matrix, but it has known biases — for instance, favoring longer or more confident-sounding answers regardless of whether they're actually more correct. Never treat it as the sole, final signal.
Have a person manually review the top 2–3 candidate combinations' outputs on a subset of the golden set. This is what actually validates whether the LLM-judge's scores can be trusted at all — always do this before finalizing a decision based purely on automated scores.
For format compliance — valid JSON, required fields present, citation format followed — plain code can check this perfectly. A JSON parser either succeeds or it doesn't; there's no need to spend an expensive LLM call on something a simple script verifies instantly and exactly.
💻 Section 9: A Simple Evaluation Harness (With Code)
This runs every combination of model × template against the golden test set from Section 4, scores each output with a mix of cheap deterministic checks and an LLM judge, and produces the matrix from Section 7 automatically.
# Model x Template Evaluation Harness (Pseudocode) def run_evaluation_matrix(golden_set, candidate_models, candidate_templates): results = {} for model in candidate_models: for template in candidate_templates: scores = [] for case in golden_set: prompt = template.render( context=case.retrieved_context, question=case.question ) answer = model.generate(prompt) # Cheap deterministic check first — no LLM call needed format_ok = validate_format(answer, expected_format=template.output_format) # Only spend an LLM-judge call if the format is even valid if format_ok: judge_score = llm_judge.score( question=case.question, context=case.retrieved_context, ideal_answer=case.ideal_answer, candidate_answer=answer ) else: judge_score = 0 # invalid format auto-fails scores.append(judge_score) results[(model.name, template.name)] = average(scores) return results # → the matrix from Section 7
❓ Frequently Asked Questions
Score it against RAG-specific criteria rather than general leaderboards: groundedness to the retrieved context, honest refusal on unanswerable questions, instruction/format compliance, and latency and cost per query (Section 5) — all measured on a small, deliberately hard test set (Section 4).
Because the same model can produce dramatically different results depending on how the prompt is structured (Section 2) — testing one while assuming the other is fixed can lead to the wrong conclusion about which model is actually best, as the Model Z example in Section 7 shows directly.
A small, fixed set of realistic (question, retrieved context, ideal answer) triples — usually 30–50 cases — deliberately including multi-chunk, numeric, and unanswerable questions, used as a consistent benchmark for comparing model and template combinations (Section 4).
Yes, and it scales well across a full model × template matrix, but LLM-as-judge has known biases — like favoring longer or more confident answers — so its scores should always be cross-checked with a human spot-check on the top contenders (Section 8), never trusted alone.
Adding an explicit instruction telling the model to answer only from the provided context and say "I don't know" otherwise (Section 6) — in our evaluation matrix example, this one change lifted every candidate model's groundedness score more than any other single tweak.
🛡️ Section 10: Pitfalls Beginners Commonly Fall Into
Always cross-check a sample of judge scores against real human judgment (Section 8) — an uncorrected judge bias can quietly crown the wrong winner.
If you tweak both in the same test run without the full matrix (Section 7), you can never tell which change actually caused the improvement — don't shortcut the grid.
If every test case is an easy, single-chunk lookup, every model and template will score near-perfectly and you'll learn nothing useful — the hard cases from Section 4 are exactly where real differences show up.
Bring cost and latency into the same decision from the start (Section 5) — a "winning" combination that's unaffordable at your real traffic volume isn't actually the winner, as our Model Z example in Section 7 showed directly.
🎓 Section 11: Cheat Sheet — Picking Your LLM + Prompt Template
- 30–50 cases, including multi-chunk, numeric, and no-answer-in-context questions
- Vary templates on context placement, "don't guess" instructions, and few-shot examples
- Every model × every template, scored on the same golden set — never one variable at a time
- Deterministic format checks → LLM-as-judge → human spot-check the top contenders
- Pick the combination that's best on the metrics that matter for your actual use case, not just the top raw score
🎉 Final Summary
You don't need a sprawling RAG evaluation framework to make this decision well — you need a small, honest test set, a matrix that tests model and template together, and layered scoring you actually trust. Get this one decision right, and the clean, precise context your earlier pipeline stages worked so hard to produce finally gets the thoughtful, grounded answer it deserves.
Happy Building! Test the Pair, Not Just the Part. 🔥
Comments
Post a Comment