Runtime Grounding Evaluation: Using Live Traffic to Generate Continuous Quality Signals
Runtime grounding evaluation is the practice of scoring, on an ongoing basis, whether the answers an LLM system actually sends to real users are supported by the source material it was given — not just whether a model passed a test suite before launch, but whether the same system stays grounded once it is handling thousands of unpredictable questions a day. 🧭
Here is why the stakes are real: a retrieval pipeline that scored 95% faithful in a pre-launch benchmark can quietly drift the moment a knowledge base is re-indexed, a new document format breaks chunking, or a traffic spike pushes users toward questions the eval set never covered. Without a live signal, that drift is invisible until a customer posts a screenshot of a confidently wrong answer, or a regulator asks why a financial assistant fabricated a policy term. Teams that only test offline are, in effect, driving by looking in the rearview mirror. 🚨
Original diagram: live traffic flowing through a grounding check into a continuous quality signal.
📑 In This Post
- 1. What Runtime Grounding Evaluation Actually Means
- 2. Offline Grounding Tests vs. Runtime Grounding Signals
- 3. How a Grounding Score Gets Computed: Faithfulness and the RAG Triad
- 4. Turning Single Scores Into a Continuous Signal
- 5. Blocking in Real Time: Guardrail-Layer Grounding Checks
- 6. Closing the Loop: Drift, Stale Golden Sets, and Feedback
- 7. Rolling Out Runtime Grounding Evaluation at Enterprise Scale
- 8. Common Mistakes
- ❓ FAQ
- 🔗 References & Further Reading
- 📝 Summary
🔀 Quick Comparison: Offline Grounding Tests vs. Runtime Grounding Signals
| Dimension | Offline / Pre-Launch Eval | Runtime / Live-Traffic Signal |
|---|---|---|
| What it measures | A fixed set of golden questions before a model or prompt ships | Every sampled real answer as it goes out the door |
| Data source | Curated test set, written or approved by humans | Live traffic evidence: real query, real retrieved context, real answer |
| Frequency | Run in CI, on a schedule, or before a release | Continuous — scores accumulate into a rolling time series |
| Catches | Known failure modes the test set anticipated | Unknown drift: new topics, index changes, edge cases nobody wrote a test for |
| Typical tooling | Ragas, promptfoo, offline MLflow evaluate() runs | Arize online evals, Databricks Lakehouse Monitoring, Bedrock/Azure runtime guardrails |
| Blind spot | Can look perfect and still miss production reality | Needs sampling and cost control, or judge bills spiral |
Neither replaces the other. Offline evals are the seatbelt you check before the car leaves the factory; runtime signals are the sensor that keeps checking the seatbelt is still buckled on every trip after that. This post is mostly about the second one, because it is the piece most teams under-build.
1️⃣ What Runtime Grounding Evaluation Actually Means
Kid analogy first
Imagine a librarian who answers questions all day long. A teacher can quiz the librarian once, in the morning, with ten sample questions, and give a report card. That report card tells you how the librarian did on those ten questions. But what you really want to know is: is the librarian still checking the actual books before answering the hundredth question of the afternoon, or has she started making things up because she's tired and the questions keep coming? Runtime grounding evaluation is having someone quietly re-check a sample of her afternoon answers against the real books, all day, every day — not just grading the morning quiz.
A production LLM system faces the same problem. A model can pass every offline test and still start inventing details once it meets a document format the eval set never included, a knowledge base that got re-indexed overnight, or a user phrasing a question in a way nobody anticipated.
Stage by stage, this is what "runtime" adds on top of an evaluation metric that could otherwise be run offline:
- The input is real, not curated. The query and retrieved context come from whatever a live user actually typed and whatever the retriever actually returned — including malformed queries, partial retrievals, and edge cases no test author thought to write down.
- The check runs at or near the moment of generation. Whether it's a guardrail blocking the response before it reaches the user, or a sampled judge scoring it seconds after, the evaluation is tied to a specific, timestamped production event.
- The output accumulates, it doesn't just pass or fail once. Each scored response becomes one point in a running series, so a team can see a trend — a slow decline over a week — not just a single verdict.
🎯 Use this when: you need to know whether a system that already shipped is still behaving the way it did in your pre-launch tests.
2️⃣ Offline Grounding Tests vs. Runtime Grounding Signals
Kid analogy first
A fire drill is an offline test: everyone knows it's coming, the building is empty of real danger, and it tells you whether people remember the exit route. A smoke detector is a runtime signal: it's watching the actual air, all the time, for the actual thing you care about, with no advance notice. You need both, and neither one substitutes for the other.
Offline grounding evaluation is what most teams build first, because it's tractable: assemble a golden set of questions with known-good source documents, run the pipeline against them, and score faithfulness with a metric library. This is valuable and necessary — but it only tells you about the questions someone thought to write down, evaluated against a knowledge base frozen at one point in time.
Runtime grounding evaluation inherits the same underlying scoring logic — usually still an LLM comparing a generated answer to retrieved context — but points it at whatever traffic is actually arriving right now. That distinction matters because the two failure modes it catches are different in kind, not just in degree: an offline test tells you the system can be grounded under conditions you anticipated; a runtime signal tells you it is grounded under conditions you didn't.
🎯 Use this when: you're deciding where in your evaluation budget to invest next — offline coverage for known risks, runtime signal for the risks you can't yet name.
3️⃣ How a Grounding Score Gets Computed: Faithfulness and the RAG Triad
Kid analogy first
Picture a book report where the teacher doesn't just ask "is this a good sentence?" — she asks, "can you show me the page in the book that says this?" for every single sentence in the report. If a sentence can't be traced back to a page, it doesn't count as truthful, even if it sounds smart. That's the entire idea behind a grounding or faithfulness score.
Mechanically, most current tools break a generated answer into individual claims and check each one against the retrieved context, rather than scoring the answer as one indivisible blob. This decomposition matters: an answer can be 80% supported and 20% fabricated, and a claim-level check surfaces exactly which 20%, instead of averaging it away into a single ambiguous number.
The three questions a grounding-aware suite actually asks
- Context relevance — did the retriever hand the generator anything useful in the first place? If the context is irrelevant, everything downstream is built on a bad foundation.
- Groundedness / faithfulness — of the claims in the answer, how many are supported by that context? This is the number most people mean when they say "hallucination rate," just measured as its inverse.
- Answer relevance — even if every claim is grounded, does the answer actually address what the user asked? A grounded non-answer is still a bad answer.
🎯 Use this when: you're choosing which library to prototype with, or explaining to a stakeholder why "the model hallucinated" needs to be broken into retrieval failure vs. generation failure before anyone can fix it.
4️⃣ Turning Single Scores Into a Continuous Signal
Kid analogy first
One thermometer reading tells you today's temperature. A hundred thermometer readings, one per day, plotted on a wall calendar, tell you it's getting colder. A single grounding score on one production answer is the first kind of measurement; runtime grounding evaluation is built to produce the second kind — a trend you can act on, not a snapshot you forget by lunchtime.
Getting from "one score" to "a signal" requires three things working together: a place to capture what happened (a trace), a way to decide which traces get scored without scoring literally everything (a sampler), and a store that turns scattered scores into a time series a human or an alert rule can read.
Here's the stage-by-stage flow that makes a signal like this trustworthy rather than just noisy:
- Capture the trace. Every production call is logged as a structured record: the user query, whatever context was retrieved, and the final answer, ideally with enough metadata to reconstruct exactly what the model saw.
- Sample deliberately. Judging every single production call with another LLM call is usually not affordable at scale, so most teams sample — for example, a fixed percentage of traffic, plus 100% of traffic that trips a cheaper heuristic filter first.
- Score with a consistent judge configuration. The same judge model, prompt template, and scoring rubric need to stay fixed across a time window, or a shift in scores could just mean someone swapped judge models, not that grounding got worse.
- Write the score into a time-indexed store. Aggregate by hour or by day, segment by feature, model version, or customer tier, so a regression in one slice doesn't get diluted into a healthy-looking global average.
- Alert on the trend, not the single data point. A single low-scoring answer is expected noise; a sustained downward slope across a rolling window is the actual signal worth paging someone about.
🎯 Use this when: your team has already shipped a RAG or agent system and is asking "how would we even know if this got worse?" — the answer is this pipeline, not another offline benchmark.
5️⃣ Blocking in Real Time: Guardrail-Layer Grounding Checks
Kid analogy first
A referee who blows the whistle the instant a foul happens is different from a coach who reviews game tape afterward and says "we should train that better." Both matter, but only the referee stops the play before it counts. Some grounding checks are built to act like the referee: they inspect a response before it ever reaches the user, and they can refuse to let it through.
GROUNDING or RELEVANCE — it can be set to BLOCK a response outright, replacing it with fallback messaging, or to NONE, which still records the detection in the trace without stopping the response. Picture a support bot whose source document says: returns without a receipt get store credit, and returns made with a receipt within 30 days get a refund. A user asks how long they have to return an item if they've lost the receipt. An answer of "you have 30 days" would be ungrounded — that 30-day window belongs to a different clause in the source and was never attached to receipt-less returns at all. An answer of "you'll receive store credit" is grounded and correct, but doesn't actually address the time-window question that was asked, so it fails on relevance instead. That's the distinction Bedrock's check exists to catch: two separate ways an answer can go wrong, scored separately instead of blended into one pass/fail. Microsoft's Azure AI Content Safety ships a comparable capability: a groundedness detection endpoint, currently in public preview, that a team can call to check whether a specific answer is actually backed by the documents it was supposed to draw from — sitting in the same product family as Azure's separate prompt-injection and protected-material checks.The mechanical difference between this and the sampled, after-the-fact judge from the previous section is timing and consequence:
- It runs inline, in the request path, which means it adds latency to every single call it's applied to, not just a sampled subset.
- It can take an action, not just log a score. A guardrail can block, redact, or substitute a fallback message; a monitoring judge can only tell you afterward that something was wrong.
- It has to be cheap and fast enough to run on 100% of traffic (or close to it) if it's protecting every response, which is why guardrail-layer checks tend to be lighter-weight, more narrowly scoped classifiers than the deep, claim-by-claim judges used for offline or sampled monitoring.
- It pairs naturally with a canary rollout. Teams shipping a new prompt or model version often route only a small percentage of live traffic to it at first specifically so a guardrail-triggered spike in blocked responses shows up against a limited slice of users, not the entire customer base, before the rollout widens.
🎯 Use this when: the cost of a single ungrounded answer reaching a user is high enough to justify added latency on every request — regulated industries, medical or legal guidance, or anywhere a wrong answer creates liability.
6️⃣ Closing the Loop: Drift, Stale Golden Sets, and Feedback
Kid analogy first
A map that was accurate five years ago can still get you lost today if the city built three new roads and closed one bridge. The map isn't wrong because it was drawn badly — it's wrong because the city changed and the map didn't. A golden evaluation dataset is a map of what "good" looked like on the day it was written; if nobody updates it, it quietly stops describing the world your users actually live in.
This is where runtime grounding evaluation earns its keep beyond monitoring: the flagged, low-scoring production cases it surfaces are exactly the raw material a stale golden dataset needs. A case that a runtime judge scores as ungrounded, and that a human reviewer confirms really was ungrounded, is a new, real, high-value test case — arguably more valuable than a synthetic one, because it's proof the failure actually happens with real users.
Original diagram: the seven-step runtime grounding feedback loop, from live traffic capture through to the golden dataset.
The loop closes in roughly this order:
- A runtime judge or guardrail flags a live answer as ungrounded.
- A human reviewer confirms it's a genuine failure, not a judge false positive.
- The confirmed case — query, retrieved context, and the bad answer — gets added to the golden evaluation set with an annotated correct answer.
- The next CI-gated offline eval run now includes this case, so a regression on it blocks a deploy before it ever reaches production again.
- The dataset's owner periodically reviews whether the mix of cases still reflects current traffic patterns, retiring cases that no longer represent real usage.
🎯 Use this when: your offline eval scores have been flat for months while your production incident count hasn't — a common sign the golden set stopped representing reality a while ago.
7️⃣ Rolling Out Runtime Grounding Evaluation at Enterprise Scale
Kid analogy first
One kid checking their own homework is easy to organize. Coordinating a hundred classrooms across a school district so every teacher grades consistently, every grade gets recorded somewhere the principal can see, and nobody accidentally sees another student's private test answers — that needs real rules, not just good intentions. Enterprise rollout of runtime grounding evaluation is the district-wide version of that problem.
Ownership and governance
Someone specific needs to own the eval pipeline itself — its judge prompts, its thresholds, its sampling logic — the same way a specific team owns a production database, rather than leaving it as an informal side project of whichever engineer built the first version.
Test-set versioning and dataset drift
Golden datasets should be versioned like code: every addition, removal, or answer correction goes through review, and every evaluation run records which dataset version it used, so a score change can always be traced to either a model change or a dataset change — never left ambiguous between the two.
CI-gated evaluation for model and prompt changes
The flow that makes this concrete usually looks like:
- An engineer opens a pull request that changes a prompt template or swaps a model version.
- CI automatically runs the versioned golden set through the proposed change, scoring faithfulness (and any other gated metrics) with the standard judge configuration.
- The build compares the result against an agreed floor, not against the previous run's score, so a slow multi-release decline can't sneak through one small drop at a time.
- If any gated metric falls below the floor, the merge is blocked, and the reviewer sees the specific failing cases plus the judge's rationale for each — not just a red X.
# illustrative CI step — not copied from any vendor pipeline
- name: run-grounding-eval
run: |
python run_eval.py \
--dataset golden_set_v14.jsonl \
--judge-model eval-judge-v2 \
--metric faithfulness \
--fail-below 0.90
# pipeline fails the build if faithfulness on the
# versioned golden set drops below the agreed floor
Access control and data governance for evaluation data
Runtime evaluation data is, by construction, a copy of real user traffic — which means it inherits whatever privacy and compliance obligations apply to that traffic. Access to raw traces should be scoped the same way access to production logs is scoped, and any case promoted into a shared golden dataset should be reviewed for what personal or sensitive information it might be carrying forward.
Cost governance for LLM-as-judge calls at scale
A judge call is itself an LLM call, with its own token cost, and at production volume that cost is not negligible. Sampling strategy (section 4), a cheaper first-pass classifier before an expensive judge, and periodic review of which judge model is actually necessary for which metric are all legitimate cost levers — but they should be decided deliberately and revisited, not left as whatever default a prototype happened to ship with.
Observability: training-time metrics vs. inference-time metrics
A dashboard built for training-time metrics — loss curves, offline benchmark scores — answers a different question than one built for inference-time metrics — live groundedness, latency, cost per call. Both matter, but conflating them onto one view tends to bury the inference-time signal, which is usually the one that changes on a shorter time horizon and needs faster reaction.
Alerting for quality regressions
An alert threshold set on a rolling window, not a single low score, and routed to whoever actually owns the pipeline (not a generic on-call queue that doesn't know what faithfulness means), is what turns a continuous signal into something the organization actually acts on rather than a dashboard nobody watches.
🎯 Use this when: more than one team is building on the same underlying LLM platform and needs a shared, auditable answer to "is this still grounded?" rather than each team reinventing its own ad hoc check.
8️⃣ Common Mistakes
- Relying on a single aggregate score. A blended "quality score" that averages faithfulness, relevance, and tone hides which one actually regressed. Keep the RAG-triad-style metrics separate so a drop in groundedness isn't masked by a stable relevance score.
- Using the same model as both generator and judge, with no bias check. A judge sharing blind spots or stylistic preferences with the generator it's grading can systematically over-score its own kind of mistake. Periodically cross-check with a different judge model or a human-labeled sample.
- No held-out test set / eval-on-training-data leakage. If the same documents used to build or fine-tune a retrieval or generation component also appear in the golden eval set, the score reflects memorization, not generalization to genuinely new questions.
- Ignoring latency and cost as first-class eval dimensions. A guardrail that catches every ungrounded answer but doubles response latency, or a judge pipeline that quietly becomes the largest line item in the LLM bill, is a quality win with an unmeasured operational cost — both belong on the same dashboard as faithfulness.
- Treating an offline eval pass as sufficient without live monitoring. As section 2 covers, a green offline run answers "can this be grounded under conditions we thought of" — not "is it grounded right now, under today's traffic." Shipping without a runtime signal is shipping blind to the second question.
- Letting golden datasets go stale as user behavior or source content shifts. A dataset frozen at launch measures how well a system matches launch-day expectations, which drift further from current reality with every passing month the dataset goes un-reviewed.
- Never red-teaming the grounding check itself. A guardrail or judge that scores well on ordinary traffic can still have a blind spot that an adversarially worded prompt walks straight through — phrasing designed specifically to smuggle an unsupported claim past the check. Treating the check as a fixed, trusted wall rather than something worth attacking on purpose means the first real adversarial input a team sees may be the one that got through.
❓ FAQ
Is runtime grounding evaluation the same thing as a guardrail?
Not quite. A guardrail (like Bedrock's contextual grounding check or Azure's groundedness detection) is one implementation pattern — an inline check that can block a response before it reaches a user. Runtime grounding evaluation is the broader practice, which also includes sampled, after-the-fact judges (like Arize's online evaluations) that produce a monitoring signal without necessarily blocking anything. A mature setup usually uses both.
How often should a production system be sampled for grounding checks?
There's no single correct percentage — it depends on traffic volume, judge cost, and how much risk a delayed detection carries. What matters more than the exact number is stratifying the sample so rare, high-stakes intents aren't drowned out by high-volume, low-stakes ones, and keeping the sampling method consistent enough that a change in scores reflects a real quality shift rather than a change in what got sampled.
Can LLM-as-judge scoring be trusted for something as important as hallucination detection?
It's a useful, scalable proxy, not an infallible ground truth. Teams building judge-based groundedness metrics, including the group behind TruLens's RAG triad, have published benchmarking work specifically because judge implementations vary in accuracy and needed deliberate tuning. Validate a judge against a human-labeled sample before trusting its scores at scale, and re-validate periodically.
Do I need a vector database or RAG pipeline for grounding evaluation to matter?
Grounding evaluation matters most wherever a model is expected to base its answer on some external source — a retrieved document, a tool's output, a provided file — rather than open-ended generation. RAG is the most common shape today, but the same faithfulness idea applies to any system where an answer is supposed to trace back to specific evidence.
What's the single most common reason teams skip runtime grounding evaluation?
Usually cost and perceived complexity — running an extra LLM call to judge every sampled response feels like pure overhead compared to a one-time offline eval. In practice, a modest sampling rate with a cheap first-pass filter keeps judge costs manageable, and the cost of an undetected production regression — in trust, support load, or compliance exposure — is typically far higher than the cost of the judge calls that would have caught it.
🔗 References & Further Reading
- AWS: Use contextual grounding check to filter hallucinations in responses
- Microsoft Learn: Groundedness detection (preview) — Azure AI Content Safety
- Microsoft Learn: What is Azure AI Content Safety?
- Databricks: What is Mosaic AI Agent Evaluation?
- Databricks: Monitor apps deployed using Agent Framework (Lakehouse Monitoring for GenAI)
- Arize AI: Online LLM Evaluations
- Arize AI Docs: Evaluation overview (Phoenix)
- Ragas paper: "Ragas: Automated Evaluation of Retrieval Augmented Generation" (arXiv:2309.15217)
- TruLens Docs: The RAG Triad
- Snowflake Engineering Blog: Benchmarking LLM-as-a-Judge for the RAG Triad Metrics
All product names (Amazon Bedrock, AWS, Azure AI Content Safety, Microsoft, Databricks, Mosaic AI, Arize, Phoenix, Ragas, TruLens, Snowflake, and any other tools mentioned) are trademarks of their respective owners and are referenced here for identification and educational purposes only.
📝 Summary
- Runtime grounding evaluation scores real, live answers on an ongoing basis — it's the smoke detector to offline evaluation's fire drill.
- Offline tests and runtime signals catch different failure modes; you need both, and runtime is the one most teams under-build.
- Faithfulness scoring works by decomposing an answer into claims and checking each against retrieved context — Ragas and TruLens's RAG triad are the two frameworks that made this mainstream.
- A trustworthy signal needs trace capture, deliberate (stratified) sampling, a consistent judge configuration, time-indexed storage, and trend-based alerting — not just one score.
- Guardrail-layer checks like Bedrock's contextual grounding check and Azure's groundedness detection act inline, in real time, and can block a response outright — a different tool from sampled monitoring.
- Flagged production failures should feed back into the golden dataset and CI-gated evals, closing the loop against both behavioral and content drift.
- Enterprise rollout needs explicit ownership, versioned test sets, CI gating, access control, cost governance for judge calls, separated training-time vs. inference-time dashboards, and trend-based alerting.
- The most common mistakes all share a theme: collapsing distinct signals into one number, and trusting a point-in-time check to stand in for an ongoing one.
If there's one habit worth taking from this post, it's this: treat every answer your system sends as a small piece of evidence about how it's actually doing, not just as a transaction that's over once it's sent. That's the whole idea behind making grounding evaluation continuous instead of occasional. Happy shipping — and happy monitoring! 🚀
Comments
Post a Comment