Skip to main content

Evaluating LLM & Prompt Templates

Calculating read time…

Evaluating LLMs and prompt templates for RAG means testing which model and which prompt structure, paired together, produce the most grounded, accurate answers from your retrieved context — a distinct, often-skipped decision separate from evaluating retrieval itself. Here's exactly why teams get this wrong, and how to get it right with evidence, not guesswork.

Here's a scene that plays out in almost every RAG project: the team spends weeks perfecting parsing, chunking, embeddings, and re-ranking. Then, at the very last mile, someone just... picks a model. "Let's use GPT-whatever, it's popular." They write a prompt in five minutes. Ship it. Three weeks later, users start complaining the bot "makes stuff up" or "ignores half the document." Nobody thought to actually test this last, tiny-seeming decision — and it turns out to be one of the biggest levers in the entire system.

This post is not about full RAG evaluation — we're not measuring retrieval recall or chunk quality here, that's a much bigger topic for another day. We're staying laser-focused on one narrow, practical, often-skipped question: once you have good retrieved context in hand, which LLM should generate the answer, and with exactly what prompt template? By the end, you'll have a repeatable, evidence-based way to answer that — no guessing, no "it feels better" hand-waving.

Evaluating LLM and prompt template choices in a RAG pipeline — model versus template comparison diagram

🎯 This is a generation-stage decision — the very last mile of the RAG pipeline
🧩 Two variables are secretly tangled together here: the model and the prompt template
⚖️ A single sentence added to a prompt template can lift accuracy more than switching to a pricier model
🧪 A tiny, carefully-built test set of 30 examples reveals more truth than 300 randomly grabbed ones.

Let's slow down and build real intuition for why this "last mile" decision matters so much.

🍳 Section 0: The Recipe Analogy — Why "Which LLM?" Isn't the Whole Question

Imagine you've done all the hard work: you've bought the freshest ingredients (great retrieval), you've prepped and portioned everything perfectly (great chunking), and you've organized your pantry so anything can be found instantly (great embeddings and re-ranking). Now you hand those ingredients to a chef and say "make dinner."

💡 The Recipe Card Analogy

A brilliant chef handed a vague recipe card ("make something with these") might improvise something inconsistent — sometimes great, sometimes odd. A mediocre cook handed a precise, well-written recipe card (exact steps, exact quantities, exact plating instructions) often produces a more consistent, reliable dish every single time.

The LLM is the chef. The prompt template is the recipe card. Picking "the best chef" without also fixing "the best recipe card" is only half the decision — and it's the half most teams stop at.

This is why teams are so often surprised when a "top-tier" model still gives wrong or ungrounded answers — they never tested whether the recipe card (prompt template) they were using was actually any good.


🚧 Section 1: Drawing the Boundary — This Is Not Full RAG Evaluation

Let's place this precisely in the pipeline we've been building across this series, so it's clear exactly what we're deciding here versus what's already been handled by earlier stages.

🗺️ Where This Decision Lives in the Pipeline

🔍 Parsing
✂️ Chunking
🧬 Embedding
🎯 Re-Ranking

— all covered earlier in this series; assume these are already solid —

⬇️
🤖 LLM + Prompt Template
Given good context, which model + which recipe card turns it into the best answer?
🚫 What We're Explicitly Not Doing Here: We're not measuring retrieval recall, chunk quality, or embedding fit — assume those already work well (or are being tracked separately, as in our earlier posts). We're isolating one narrow pair of decisions and giving it the careful attention it deserves.

🧩 Section 2: The Two Tangled Variables — Model and Template

🤖
The LLM (the Chef)

Its raw ability to reason, follow instructions carefully, stay honest about what the context actually says, and produce well-formatted output.

📝
The Prompt Template (the Recipe Card)

How the instructions are structured, where the retrieved context sits, whether examples are included, and what output format is demanded.

Here's a real, concrete example of the same model producing two very different answers, purely because of the recipe card it was handed:

❌ Weak Template
Context: {context}
Question: {question}
Answer the question.
Result: model confidently adds outside knowledge the context never mentioned, because nothing told it not to.
✅ Stronger Template
Context: {context}
Question: {question}

Answer using ONLY the context above.
If the answer isn't there, say
"I don't have enough information."
Result: same model, dramatically fewer confident wrong answers.
💡 The Insight Most Teams Miss: Nobody changed the model between these two boxes above — the exact same "chef" produced both results. That's why you can never evaluate "which LLM is best" in isolation. You have to test the pairing of model and template together, because each model responds differently to structure, tone, and explicit instructions.

👻 Section 3: Quick Definition — What Does "Hallucination" Actually Mean Here?

We're about to use this word a lot, so let's nail it down precisely, since beginners often hear it thrown around vaguely.

🔍 Hallucination, in Plain English

A hallucination is when an LLM states something confidently that isn't actually supported by the context it was given — it's not "lying" on purpose; the model is simply pattern-completing text, and sometimes the most "natural-sounding" continuation happens to be false or invented. In RAG specifically, the entire point of retrieval was to give the model real facts to stand on — a hallucination means the model wandered off that ground and made something up instead.

This is exactly why groundedness — sticking strictly to the provided context — becomes the single most important thing we're about to measure in both the model and the template.


🏅 Section 4: Building a Small, Honest "Golden" Test Set

Before comparing anything, you need fixed, realistic test cases — otherwise "Model A felt better than Model B" is just an opinion, not evidence. Each test case is a simple triple:

{
  "question":          "What's the refund window for unopened items?",
  "retrieved_context":  "...(the actual chunks your pipeline would return)...",
  "ideal_answer":       "30 days from the purchase date, for unused items only."
}

A genuinely useful golden set doesn't just contain easy questions — it deliberately includes the situations where models tend to trip:

🧩 Multi-chunk questions

The answer requires combining facts from two or three different retrieved chunks, not just one.

🔢 Numeric or table-based facts

Prices, dates, and quantities — easy to accidentally invent or mix up.

🕳️ Genuinely unanswerable questions

Where the retrieved context simply doesn't contain the answer — the correct response is an honest "I don't know," and this is the single best way to catch hallucination directly.

✅ Why 30–50 cases is genuinely enough: A test set this size is small enough that a human can actually sit down and read every single generated answer during review — and that human judgment is worth far more than a bigger pile of test cases nobody ever actually looks at closely.

🤖 Section 5: What "Evaluating an LLM" Actually Means for RAG

Ignore general leaderboards for a moment — the ones ranking models on broad trivia, coding, or general chat quality. Those measure a completely different skill. Being generally brilliant doesn't automatically mean a model is good at the very specific job of "read this handful of paragraphs, and only answer from what's actually written there." Score each candidate model against four practical criteria instead:

🔒 Groundedness (Faithfulness to Context)

Does the model stick strictly to the provided chunks, or does it quietly reach for outside knowledge it learned during training? This is the single most important trait for RAG — a model that "sounds smart" but ignores the context defeats the entire point of retrieval.

🙅 Correct Refusal Behavior

When your "unanswerable question" test cases (Section 4) come up, does the model honestly say it doesn't know, or does it confidently invent something plausible-sounding? This single behavior separates a trustworthy assistant from a dangerous one.

📐 Instruction & Format Compliance

If you asked for JSON, three bullet points, or a specific citation style — did it comply exactly, every single time? Some models are noticeably more reliable at rigid formatting than others, which matters enormously if your output feeds another system downstream.

⚡ Latency & Cost Per Query

A model that's marginally more accurate but costs five times more and answers three times slower may not be worth it for a high-traffic support chatbot, even if it wins every quality metric on paper.


📝 Section 6: What "Evaluating a Prompt Template" Actually Means

Just like the model, a prompt template has structural choices that measurably swing output quality — test these deliberately, one at a time.

📍 Where the Context Sits

Does the retrieved context come before or after the instructions? Many models pay closer attention to text placed nearer the actual question — worth testing directly rather than assuming one order is obviously right.

🚫 An Explicit "Don't Guess" Instruction

A template that plainly says "If the answer isn't in the context, say you don't know" measurably reduces confident wrong answers compared to one that doesn't bother saying this at all — this single line is one of the highest-leverage additions you can test, exactly as shown in Section 2's side-by-side example.

🧾 Few-Shot Examples vs. Zero-Shot

Including 1–2 example (context, question, ideal-answer-format) demonstrations directly in the prompt often stabilizes formatting and tone — at the cost of extra tokens and latency on every single call.

🧠 Asking It to Reason Step-by-Step First

For multi-chunk or numeric questions (Section 4), explicitly asking the model to reason through the pieces before giving a final answer can genuinely improve accuracy — but it adds latency, and can leak unwanted "thinking out loud" text into the final output if you don't clearly separate reasoning from the answer.


🧮 Section 7: Combining Both Variables Into One Matrix

Now bring the model (Section 5) and the template (Section 6) together in one grid — every candidate LLM, run against every candidate template, scored against the exact same golden test set from Section 4.

Model ↓ / Template → Template A (minimal) Template B (+ "don't guess") Template C (+ few-shot)
Model X Groundedness: 72% Groundedness: 89% Groundedness: 91%
Model Y Groundedness: 81% Groundedness: 93% Groundedness: 94%
Model Z (cheaper) Groundedness: 65% Groundedness: 84% Groundedness: 88%
✅ Read This Grid Slowly — the Real Insight Is Hiding in Plain Sight: Every single model jumps significantly from Template A to Template B, just from adding one honesty instruction. Notice something else: Model Z (the cheaper option) with Template C actually beats Model X with Template A — meaning the "cheap model" could have looked like the wrong choice, purely because it was tested with a weak recipe card. This is exactly why testing model and template separately, instead of together in one matrix, quietly leads teams to the wrong conclusion.

🧪 Section 8: How Do You Actually Produce Those Scores?

🤖 LLM-as-Judge (Fast, Scalable — Use With Caution)

Prompt a separate, strong LLM to compare each generated answer against the ideal answer and the original context, scoring groundedness and correctness on a simple rubric. It's fast enough to run across an entire matrix, but it has known biases — for instance, favoring longer or more confident-sounding answers regardless of whether they're actually more correct. Never treat it as the sole, final signal.

🧑‍⚖️ Human Spot-Check (Slower, But It's Your Ground Truth)

Have a person manually review the top 2–3 candidate combinations' outputs on a subset of the golden set. This is what actually validates whether the LLM-judge's scores can be trusted at all — always do this before finalizing a decision based purely on automated scores.

📏 Deterministic Checks (Cheap, Exact, No Ambiguity)

For format compliance — valid JSON, required fields present, citation format followed — plain code can check this perfectly. A JSON parser either succeeds or it doesn't; there's no need to spend an expensive LLM call on something a simple script verifies instantly and exactly.


💻 Section 9: A Simple Evaluation Harness (With Code)

📌 What This Code Does (Read Before The Code!)

This runs every combination of model × template against the golden test set from Section 4, scores each output with a mix of cheap deterministic checks and an LLM judge, and produces the matrix from Section 7 automatically.

# Model x Template Evaluation Harness (Pseudocode)

def run_evaluation_matrix(golden_set, candidate_models, candidate_templates):
    results = {}

    for model in candidate_models:
        for template in candidate_templates:

            scores = []
            for case in golden_set:
                prompt = template.render(
                    context=case.retrieved_context,
                    question=case.question
                )
                answer = model.generate(prompt)

                # Cheap deterministic check first — no LLM call needed
                format_ok = validate_format(answer, expected_format=template.output_format)

                # Only spend an LLM-judge call if the format is even valid
                if format_ok:
                    judge_score = llm_judge.score(
                        question=case.question,
                        context=case.retrieved_context,
                        ideal_answer=case.ideal_answer,
                        candidate_answer=answer
                    )
                else:
                    judge_score = 0   # invalid format auto-fails

                scores.append(judge_score)

            results[(model.name, template.name)] = average(scores)

    return results   # → the matrix from Section 7
✅ Notice the Pattern: Cheap, deterministic checks run first and can short-circuit the more expensive LLM-judge call — the same "cheap filter before expensive precision" funnel shape we've seen throughout this series, now applied to evaluation itself.


❓ Frequently Asked Questions

How do you evaluate an LLM for a RAG system?

Score it against RAG-specific criteria rather than general leaderboards: groundedness to the retrieved context, honest refusal on unanswerable questions, instruction/format compliance, and latency and cost per query (Section 5) — all measured on a small, deliberately hard test set (Section 4).

Why do the LLM and the prompt template need to be tested together?

Because the same model can produce dramatically different results depending on how the prompt is structured (Section 2) — testing one while assuming the other is fixed can lead to the wrong conclusion about which model is actually best, as the Model Z example in Section 7 shows directly.

What is a "golden test set" in RAG evaluation?

A small, fixed set of realistic (question, retrieved context, ideal answer) triples — usually 30–50 cases — deliberately including multi-chunk, numeric, and unanswerable questions, used as a consistent benchmark for comparing model and template combinations (Section 4).

Can I use an LLM to judge another LLM's answers?

Yes, and it scales well across a full model × template matrix, but LLM-as-judge has known biases — like favoring longer or more confident answers — so its scores should always be cross-checked with a human spot-check on the top contenders (Section 8), never trusted alone.

What's the single highest-impact prompt template change to test first?

Adding an explicit instruction telling the model to answer only from the provided context and say "I don't know" otherwise (Section 6) — in our evaluation matrix example, this one change lifted every candidate model's groundedness score more than any other single tweak.


🛡️ Section 10: Pitfalls Beginners Commonly Fall Into

🎭 Trusting LLM-as-Judge Blindly

Always cross-check a sample of judge scores against real human judgment (Section 8) — an uncorrected judge bias can quietly crown the wrong winner.

🔀 Changing the Model and Template at the Same Time, Untracked

If you tweak both in the same test run without the full matrix (Section 7), you can never tell which change actually caused the improvement — don't shortcut the grid.

🧪 A Golden Set With No Hard Cases

If every test case is an easy, single-chunk lookup, every model and template will score near-perfectly and you'll learn nothing useful — the hard cases from Section 4 are exactly where real differences show up.

💸 Deciding on Quality Before Considering Cost

Bring cost and latency into the same decision from the start (Section 5) — a "winning" combination that's unaffordable at your real traffic volume isn't actually the winner, as our Model Z example in Section 7 showed directly.


🎓 Section 11: Cheat Sheet — Picking Your LLM + Prompt Template

Step 1: Build a Small, Honest Golden Test Set
  • 30–50 cases, including multi-chunk, numeric, and no-answer-in-context questions
Step 2: Shortlist 2–3 Models and 2–3 Templates
  • Vary templates on context placement, "don't guess" instructions, and few-shot examples
Step 3: Run the Full Matrix (Section 7)
  • Every model × every template, scored on the same golden set — never one variable at a time
Step 4: Score With Layered Checks
  • Deterministic format checks → LLM-as-judge → human spot-check the top contenders
Step 5: Weigh Quality Against Cost & Latency
  • Pick the combination that's best on the metrics that matter for your actual use case, not just the top raw score

🎉 Final Summary

🍳 The LLM is the chef; the prompt template is the recipe card — a great chef with a vague recipe card still produces inconsistent results
🧩 Model and template are tangled together — evaluate them as pairs in one matrix, never assume one is fixed while testing the other
👻 Hallucination means confidently stating something the context doesn't support — groundedness and honest refusal matter more here than general leaderboard rank
🏅 A small, honest golden test set with hard cases baked in reveals more truth than a huge, easy one
🧪 Layer your scoring — deterministic checks, then LLM-as-judge, then human spot-checks — and never trust the judge alone
✅ The Core Lesson:

You don't need a sprawling RAG evaluation framework to make this decision well — you need a small, honest test set, a matrix that tests model and template together, and layered scoring you actually trust. Get this one decision right, and the clean, precise context your earlier pipeline stages worked so hard to produce finally gets the thoughtful, grounded answer it deserves.


Happy Building! Test the Pair, Not Just the Part. 🔥

Comments