Skip to main content

How to Evaluate LLM Outputs: Validators, Metrics, Judges and Human Review

Calculating read time…

An LLM output can fail in four completely different ways, and no single scoring method catches all four — a well-run evaluation stack pairs each failure type with the cheapest evaluator that can reliably catch it: deterministic validators for structural rule-breaks, reference-based metrics for tasks with a known right answer, LLM-as-judge for nuanced quality judgments, and trained human reviewers for the ambiguous, high-risk edge cases nothing else can safely score. 🧭

The stakes of getting this wrong are concrete rather than academic. A team that only runs an LLM-as-judge pass will happily approve a response that hallucinated a field name, because "sounds helpful and confident" and "matches the required schema" are different questions entirely. A team that only runs deterministic checks will ship a technically-valid, well-formatted response that is condescending, off-policy, or simply wrong in a way no regex could ever catch. Production reliability comes from matching the failure mode to the right judge — not from picking one evaluation method and hoping it generalizes. ⚠️

🔀 Quick Comparison

Evaluator Best for Needs Speed / cost Blind spot
Deterministic JSON validity, required fields, tool-call shape, forbidden content patterns A written rule or schema Milliseconds, near-free Can't judge whether the answer is actually good
Reference-based Routing, classification, extraction, tool selection, refusal decisions A labeled ground-truth dataset Fast, cheap, deterministic Useless once there's no single correct answer
LLM-as-judge Helpfulness, tone, completeness, clarity, policy alignment on open-ended replies A fixed rubric + human calibration Seconds, per-call cost Can be biased, inconsistent, or gamed by verbosity
Human review Ambiguous, high-risk, policy-edge, UX-sensitive cases Trained reviewers, rubrics, calibration Slowest, most expensive Doesn't scale to every production request

1. The Four-Layer Evaluator Stack, at a Glance

Kid analogy: imagine a school science fair with four judges at four different tables. The first judge just checks that your poster board actually has a title, a hypothesis, and a conclusion glued on — they don't read a word of it, they just check the pieces are there. The second judge has an answer key for the "what's the boiling point of water" question and marks you right or wrong. The third judge is an older student with a detailed grading rubric who reads your whole project and scores how clearly you explained your thinking. The fourth judge is a science teacher who only gets called over when something's genuinely confusing, borderline, or a bit risky — like a project involving fire. Every project passes through some combination of these four judges, and no one judge could do all four jobs well. 🎪

Google's Vertex AI Gen AI evaluation service formalizes almost exactly this split in its own metric taxonomy: it separates computation-based metrics, which score against ground truth with deterministic algorithms, from rubric-based metrics, which incorporate an LLM judge and can use either dynamically generated ("adaptive") rubrics or fixed ("static") ones depending on how much per-prompt nuance the task needs, describing rubrics as the criteria for rating a response and metrics as the score measuring output against those rubrics, with rubric-based metrics recommended for writing quality, safety, and instruction-following tasks that are hard to score with deterministic algorithms alone. Azure AI Foundry's evaluation framework draws the same line, combining built-in deterministic measurements like exact match, F1, and BLEU with AI-assisted metrics for dimensions such as groundedness, relevance, coherence, and fluency as a pre-deployment quality gate before an agent reaches production. Two major cloud vendors converging on the same "deterministic first, model-graded second" split independently is a strong signal that this isn't a stylistic preference — it reflects what actually holds up under production load.

✅ Worked example: Anthropic's own guidance on building evaluations for Claude lays out three grading methods in order of preference — code-based grading (fast, reliable, and the best option whenever a task allows it), model-based grading (Claude grading Claude on things like tone or free-form accuracy), and human grading (the most capable method, reserved for cases the first two can't handle because it is comparatively slow and expensive). That ordering — cheap and reliable first, expensive and flexible last — is the same instinct behind the four-gate stack in this post.

🎯 Use this when: you're deciding how to spend a limited evaluation budget across a pipeline with more than one kind of possible failure — which is almost every production LLM system.

2. Gate 1 — Deterministic Validators

Kid analogy: before a roller coaster car is allowed to leave the station, an attendant walks down the row checking that every single lap bar is clicked shut. They're not judging whether you're going to enjoy the ride — they're checking one binary, non-negotiable fact per seat. That's a deterministic validator: a yes/no rule with zero room for interpretation. 🎢

A deterministic validator enforces objective, checkable rules against an LLM output: is this valid JSON, are the required fields present, are the value types correct, was the mandated tool actually called, does the response avoid exposing something it was told never to expose, and does it avoid making a claim the system has no basis for. None of these questions require judgment about quality — they're pass/fail checks that a short script, a JSON Schema validator, or a regex can answer in milliseconds, without ever calling another model.

A 2026 research benchmark on trustworthiness scoring for LLM structured outputs makes the production case for this gate concretely: teams that had tried using LLMs for document-processing automation found that sporadic errors and edge cases in structured outputs made pure automation unreliable at scale, and that scoring the trustworthiness of each field lets a team send only the 1–5% of genuinely uncertain outputs to a human reviewer while the remaining 95–99% flow through automatically. That ratio only holds because the deterministic layer is doing the first, cheapest pass — catching schema violations and obviously malformed fields before anything more expensive gets involved.

Without this gate, what breaks in production is usually silent at first: a downstream system expecting a field called customer_id receives customerId instead, a parser throws on a stray trailing comma, or an agent skips a required compliance disclosure tool call entirely — and nobody notices until a customer complaint or an audit surfaces it, because a fluent-sounding response gave no visual signal that anything was wrong.

Building a minimal deterministic gate, step by step:

  1. Define the contract first — the JSON Schema, required tool signature, or output format the response must satisfy, independent of any specific prompt.
  2. Write validators as pure functions with no model calls: schema validation, type coercion checks, enum membership checks, and pattern matches for disallowed content (PII patterns, banned phrases, unsupported absolute claims).
  3. Run the full validator set against every candidate output before anything else touches it — this should be the very first node in an eval pipeline or a pre-deploy CI job, since it's the cheapest possible filter.
  4. Log failures with the specific rule that broke (not just "invalid"), so failure-mode analysis can group by root cause instead of one big undifferentiated bucket.
  5. Route hard failures to auto-reject or auto-retry, and route borderline or unusual failures into the sampling pool for reference-based or human review, rather than letting them fail silently.

💡 Where this gate runs out of road: a response can pass every deterministic check — perfectly valid JSON, every required field present, the correct tool called — and still be substantively wrong: a support agent that calls issue_refund with a syntactically valid amount that is nonetheless the wrong amount. Structural correctness and correctness-of-content are different axes, and only the next two gates evaluate the second one.

🎯 Use this when: the failure you're worried about is structural — malformed output, missing fields, the wrong tool, or a hard policy line being crossed — not whether the content itself is good.

3. Gate 2 — Reference-Based Metrics

Kid analogy: think of a spelling test graded against an answer key. The teacher doesn't debate whether "definately" is a reasonable spelling — there's one correct spelling, and the answer either matches it or it doesn't. Reference-based metrics work the same way: they only make sense when there's a known-correct label to compare against. ✏️

Reference-based metrics compare a model's output against a known label using accuracy, precision, recall, or F1 (and, for generated text against a reference string, metrics like exact match or BLEU). They're the right tool for tasks with a genuinely correct answer: routing a support ticket to the right queue, classifying an email's intent, extracting a specific field from a document, choosing which of a fixed set of tools to call, or deciding whether to refuse a request. In every one of these cases, a human labeler (or an existing system of record) can produce a ground-truth label in advance, and the model's output either matches it or doesn't.

This is precisely the "computation-based metrics" category in Google's Gen AI evaluation service, described as evaluating responses with deterministic algorithms — usually against ground truth — producing a numerical score such as 0.0 to 1.0, and recommended specifically for situations where ground truth is available and can be matched with a deterministic method. Azure AI Foundry's evaluation framework lists the identical family of built-in metrics — exact match, F1 score, and BLEU score — for the same reason: they're objective, reproducible, and don't require an LLM call to compute.

Precision, recall, and F1 matter more than a single accuracy number for most of these tasks because the cost of a false positive and a false negative is rarely symmetric. A refusal classifier with 95% accuracy but 60% recall on genuinely harmful requests is quietly letting 40% of the dangerous cases straight through — a number that a single "percent correct" headline metric would hide completely. The reason teams move past a single aggregate accuracy figure is exactly this: it can look excellent while masking a lopsided, high-consequence error pattern in one class.

Without this gate, teams either fall back to reading transcripts by hand for every prompt/model change (which doesn't scale past a handful of examples) or skip regression testing on classification-style tasks entirely — meaning a prompt tweak that quietly drops routing accuracy from 96% to 89% ships unnoticed until a spike in misrouted tickets shows up in a support queue weeks later.

✅ Worked example: an intent classifier for a customer support router has a golden set of 2,000 labeled tickets. Every prompt change is scored for per-class precision and recall against that set before merge — the same "check against a known answer" instinct as the deterministic gate, except now the check requires a labeled example instead of a fixed rule.

💡 Where this gate runs out of road: that same support router, when asked to draft the actual reply to a customer rather than just pick a category, no longer has a single correct answer to compare against — two very differently worded replies can both be excellent. Reference-based scoring goes quiet exactly where open-ended generation begins, which is where Gate 3 takes over.

🎯 Use this when: the task has one correct answer and you can afford to build (and maintain) a labeled dataset — routing, classification, extraction, tool selection, and refusal decisions are the classic cases.

4. Gate 3 — LLM-as-Judge

Kid analogy: there's no single "correct" way to write a great book report — two students can write completely different reports and both deserve an A. So instead of an answer key, the teacher hands out a rubric: does it summarize the plot, does it give an opinion with reasons, is it well organized? An older student who has read the rubric carefully can grade a big stack of reports consistently using that same checklist — faster than the teacher, though not quite as reliable as the teacher herself. That's LLM-as-judge. 📖

LLM-as-judge uses a model to evaluate qualities that don't reduce to a single correct string: helpfulness, completeness, tone, clarity, and policy alignment on open-ended output. It only works reliably when the rubric is fixed and specific rather than a vague "rate this 1–5 for quality," when the judge is asked to cite evidence for its verdict rather than just assert a score, when the exact judge prompt and model version are pinned and versioned so scores stay comparable over time, and when the judge's scores are periodically checked against human ratings to catch drift or bias.

Google's Gen AI evaluation service implements this pattern as "adaptive rubrics," where rubrics are dynamically generated for each prompt and the response is evaluated with granular, explainable pass-or-fail verdicts specific to that prompt, alongside a "static rubrics" mode for when the exact same scoring criteria need to apply across every prompt in a set. Arize's Phoenix evaluation library documents the same architecture from the plumbing side: it plugs a model in as judge to grade hallucinations, factuality, helpfulness, toxicity, and custom rubrics, and its own guidance recommends using a smaller, cheaper model as the judge for large evaluation batches since the judge's job is classification rather than generation — a direct, practical acknowledgment of the cost governance concern that shows up later in this post.

The two biases every LLM-as-judge setup needs to actively check for are self-preference (a model tends to rate its own family's outputs a little more favorably than a truly neutral judge would) and length bias (longer, more confident-sounding answers tend to score higher independent of whether they're actually more correct). Anthropic's own guidance on building evals for Claude places model-based grading squarely between code-based and human grading in its recommended ordering — capable of handling free-form tasks a script can't, but requiring the same skepticism about consistency that any single automated grader deserves before it's trusted at scale.

Without calibration against human ratings, an LLM-as-judge setup can quietly reward the wrong thing for months: a judge with an undetected verbosity bias will keep nudging a team toward longer, padded responses because "more thorough-sounding" outputs win its comparisons, even while real user satisfaction with those longer responses is flat or declining.

✅ Worked example: the same support-router product from Gate 2 uses LLM-as-judge to score the drafted reply itself against a five-point rubric (addresses the actual question, correct tone, no unsupported promises, appropriately concise, grammatically clean) — reusing the golden set of tickets from Gate 2, but now scoring open-ended text instead of a category label.

💡 Where this gate runs out of road: imagine a reply that is fluent, on-tone, and rubric-compliant — but is responding to a customer who mentioned they're in a genuinely distressing personal situation, and a judge rubric built around "helpfulness and tone" simply wasn't designed to weigh that kind of human, high-stakes nuance. That's exactly the case Gate 4 exists for.

🎯 Use this when: the output is open-ended (no single correct string) but the quality dimensions you care about can be written down as a specific, evidence-requiring rubric — and you're willing to calibrate and re-check that rubric against humans regularly.

5. Gate 4 — Human Evaluation

Kid analogy: back at the science fair, most projects get scored by the rubric-following older student — but every so often, a project touches something genuinely tricky: an ethical question, a safety concern, a judgment call the rubric never anticipated. That's when the actual teacher gets pulled over, because some calls need a real, experienced human. 🧑‍🏫

Human evaluation is necessary for ambiguous cases (where reasonable people could disagree on the right answer), high-risk cases (medical, legal, financial, or safety-adjacent content), policy-edge cases (right at the boundary of what's allowed), and UX-sensitive cases where the "feel" of an interaction matters as much as its literal correctness. Doing this well requires trained reviewers who share a common understanding of the rubric, calibration sessions where reviewers score the same examples and reconcile disagreements before scoring independently, double-scoring a sample of cases to measure inter-rater agreement, and risk-based sampling that concentrates human attention on the highest-stakes slice of traffic rather than spreading it thin and even across everything.

Anthropic's own guidance is direct about where human grading sits in the hierarchy: it's described as the most capable grading method because it can be applied to almost any task, but also the slowest and most expensive — which is exactly why the guidance is to avoid designing evals that require human grading whenever a faster method will do, and to reserve it for what genuinely needs it. The industry has also been consolidating tooling around exactly this problem: Anthropic's 2025 move to bring the Humanloop team in-house was explicitly framed around strengthening human-feedback and evaluation tooling for enterprise deployments — a sign that structured human review is being treated as core evaluation infrastructure rather than an afterthought bolted onto automated scoring.

Without this gate, the failure mode isn't usually a dramatic single incident — it's a slow drift where the automated gates all report green while a genuinely uncomfortable pattern (a subtly condescending tone toward a specific user group, a policy-edge response that's technically compliant but clearly against the spirit of a guideline) goes completely undetected, because no deterministic rule, labeled dataset, or rubric anticipated it.

✅ Worked example: the support-router pipeline routes every reply that the LLM-judge scores below a threshold, plus a fixed 2% random sample of everything else, to a pool of two trained human reviewers who score independently before comparing notes — the same golden-set thread running through all four gates, now closed out by the slowest and most expensive one, applied only where it's actually needed.

💡 A trap worth naming: a single reviewer's opinion is not "human evaluation" — it's one person's opinion. Without calibration and double-scoring, disagreement between reviewers gets silently averaged into a single number that looks precise but is actually hiding real disagreement about what "good" means.

🎯 Use this when: the case is ambiguous, high-risk, right at a policy boundary, or so UX-sensitive that no rubric written in advance could be trusted to fairly score it alone.

6. Hands-On Lab: Build a Tiny Layered Harness

This is a small, disposable exercise — a five-prompt toy dataset and a few dozen lines of code, not a production system. The goal is to feel the four gates working together before scaling any of them up.

1
Write five toy prompts that ask a model to return JSON like {"intent": "...", "reply": "..."} for a made-up customer message ("my order hasn't arrived", "I want a refund", etc). Pick intents from a fixed list of five categories you define yourself.
2
Write Gate 1: a validator function that parses the output as JSON, confirms both keys exist, and confirms intent is one of your five allowed categories. Expect to see: at least one of your five toy prompts fail this check the first time you run it — that's normal and is the point of running it first.
3
Hand-label the correct intent for each of your five prompts yourself. Write Gate 2: a tiny script that compares the model's intent field to your label and reports accuracy across the five. Expect to see: a plain fraction like 4/5 — this is your entire "reference-based metric" for this toy task.
4
Write Gate 3: a second prompt that hands the model its own reply text and a three-item rubric (addresses the message, correct tone, no made-up promises), asking for a pass/fail verdict on each item plus one sentence of justification per item. Expect to see: the judge occasionally disagreeing with your own gut read on tone — write that disagreement down instead of dismissing it, since it's your first real calibration data point.

💡 Common first-timer mistake: asking the judge for a single 1–10 score instead of a per-item pass/fail with justification. A bare number gives you nothing to debug when the judge is wrong; a justified per-item verdict tells you exactly which rubric line it misapplied.

Bridging to production: this five-prompt toy is the exact same shape as the support-router example threaded through Gates 1–4 above — the only differences at production scale are the size of the labeled set (thousands of tickets instead of five), the number of reviewers doing Gate 4 (a trained pool instead of you), and the fact that all three automated gates run in CI on every change instead of by hand in a notebook.

7. Rolling This Out at Enterprise Scale

Moving from a notebook exercise to an enterprise eval pipeline raises questions that don't show up at small scale:
  • Ownership and governance: someone needs to own each golden dataset and each rubric as a living artifact — not a one-time export — with a clear approver for changes to pass/fail thresholds.
  • Test-set versioning and drift: golden datasets should be versioned like code, and reviewed on a schedule against real production traffic samples, because user behavior and phrasing shift over months even when the underlying task doesn't.
  • CI-gated evaluation: every prompt or model change runs the deterministic and reference-based gates automatically before merge, with the LLM-judge gate flagging regressions for review rather than silently blocking (since judge scores carry more noise than deterministic checks).
  • Access control and data governance: evaluation sets built from real user traffic contain the same sensitive data production traffic does, and need the same access controls, retention limits, and redaction as production logs — not a lighter "it's just for testing" standard.
  • Cost governance for LLM-as-judge at scale: judge calls are a real, recurring line item once evaluation volume grows — using a smaller model as judge for high-volume batches, sampling rather than scoring every single production request, and caching judge scores for unchanged prompt/response pairs all keep this bounded.
  • Separate dashboards for training-time vs. inference-time metrics: a metric like validation loss during fine-tuning answers a different question than live production groundedness or refusal rate, and conflating them on one dashboard makes both harder to act on.
  • Alerting for quality regressions: a live monitoring threshold (a jump in deterministic-gate failure rate, a drop in judge scores, a spike in human-escalation rate) should page a human the same way a latency or error-rate spike would.

💡 A currency note worth planning around: OpenAI has announced that its standalone Evals platform will become read-only for existing users on October 31, 2026 and shut down entirely on November 30, 2026, with functionality consolidating elsewhere in its API surface. Teams with eval pipelines built directly on that platform should treat this the same way they'd treat any vendor deprecation — plan a migration path rather than discovering it during an incident.

🎯 Use this when: your evaluation pipeline has moved from "a script one engineer runs before merging" to "a system multiple teams depend on to ship safely" — that transition is exactly when governance questions stop being optional.

8. Common Mistakes

  • Relying on one aggregate score instead of a metric suite. A single "quality: 87%" number can be flat for months while precision on one critical class quietly collapses — aggregation is exactly what hides an uneven failure pattern.
  • Using the same model as both generator and judge without bias checks. A model grading its own family's outputs has a documented tendency toward mild self-preference; without periodic calibration against human ratings, that bias compounds silently every time it's used to approve a change.
  • No held-out test set, or eval-on-training-data leakage. If the same examples used to iterate on a prompt are also used to claim it works, the reported score measures overfitting to that specific set, not generalization to new traffic.
  • Treating latency and cost as someone else's problem. A response that scores perfectly on every quality gate but takes 40 seconds or costs ten times the target per call has still failed the product requirement — cost and latency need to be first-class eval dimensions, not an afterthought checked separately by an infra team.
  • Treating an offline eval pass as sufficient without live monitoring. Offline evals run against a fixed dataset; production traffic drifts continuously, and a pipeline that only checks quality at merge time has no way to notice a slow regression that appears only after deployment.
  • Letting golden datasets go stale as user behavior shifts. A dataset built a year ago reflects how users phrased requests a year ago — as language, products, and edge cases evolve, an unreviewed golden set slowly stops representing the traffic it's meant to test.

❓ FAQ

Do I need all four evaluator types for every project?

No — match the evaluators to the failure types you actually have. A pure classification task might only need Gates 1 and 2; an open-ended assistant handling sensitive topics needs all four.

Can LLM-as-judge replace human evaluation entirely once it's well calibrated?

Calibration reduces how often a judge disagrees with humans on average, but it doesn't give the judge the ability to weigh genuinely ambiguous, high-risk, or policy-edge cases the way a trained human can — those categories are exactly why Gate 4 exists as a permanent layer, not a temporary stopgap.

How big should a golden dataset be before it's useful?

There's no universal number, but a common pattern is to start small and high-signal — enough examples to cover your known edge cases and each major task variant — and grow the set as new failure patterns are discovered in production, rather than trying to guess a "final" size up front.

What's the single biggest reason a well-scored eval suite still misses production problems?

Stale test data. An eval suite is only as good as how well its dataset represents current traffic, and traffic changes faster than most teams update their golden sets — which is why scheduled dataset review belongs in governance, not just in the initial setup.

Should cost and latency block a deployment the same way a quality regression would?

In most production systems, yes — a response that's technically excellent but too slow or too expensive to run at the required volume has still failed a real product requirement, so treating cost and latency thresholds as CI gates alongside quality gates is the more defensible default.

🔗 References & Further Reading

Official / primary documentation consulted for fact-checking:

  • Google Cloud, Vertex AI — Define your evaluation metrics (computation-based vs. rubric-based metrics): docs.cloud.google.com/vertex-ai/generative-ai/docs/models/determine-eval
  • Google Cloud, Vertex AI — Details for managed rubric-based metrics (adaptive and static rubrics): docs.cloud.google.com/vertex-ai/generative-ai/docs/models/rubric-metric-details
  • Microsoft, Azure AI Foundry evaluation framework and built-in/AI-assisted metrics documentation: learn.microsoft.com (Azure AI Foundry evaluation docs)
  • Anthropic — Create strong empirical evaluations and Building evals and test cases, docs.anthropic.com and platform.claude.com/cookbook
  • Arize — Using Anthropic Claude as judge with arize-phoenix-evals, arize.com/docs
  • OpenAI — Evaluation best practices and Evals platform deprecation timeline, developers.openai.com/api/docs/guides

Disclaimer: the explanations, examples, and analogies above are original synthesis written from an independent understanding of these concepts, verified against the primary sources above for factual accuracy — no text is reproduced verbatim from any source, and all product and company names are trademarks of their respective owners, mentioned here for identification purposes only.

📝 Summary

  • Different LLM failure types need different judges — no single evaluator covers structural errors, factual mismatches, nuanced quality, and genuine ambiguity all at once.
  • Gate 1, deterministic validators, catch structural rule-breaks fast and cheap, and should run first on everything.
  • Gate 2, reference-based metrics, handle tasks with a known correct answer — routing, classification, extraction, tool selection, refusal.
  • Gate 3, LLM-as-judge, scores open-ended quality against a fixed, evidence-requiring rubric, calibrated regularly against humans.
  • Gate 4, human evaluation, is reserved for the ambiguous, high-risk, and policy-edge cases nothing automated can safely resolve alone.
  • At enterprise scale, this stack needs ownership, versioned datasets, CI gates, data governance, cost discipline, and live alerting — not just a one-time build.
  • The most common mistakes all share one root cause: trusting a single score, a single model, or a single dataset snapshot to represent something that's actually multi-dimensional and constantly shifting.

That's the whole stack — four gates, each doing the one job it's actually good at. Build the cheap ones first, save the expensive one for what truly needs it, and revisit all four as your traffic and your product keep changing. 🚀

Comments