Skip to main content

The 5-Layer Logging Stack Every Production LLM Team Needs

Calculating read time…

The single biggest reason an LLM evaluation program collapses six months after launch isn't a bad metric — it's a missing field in a log line. If a system doesn't record which prompt version, which retrieved documents, and which grader produced a given score, then "the eval passed" is a sentence nobody can actually stand behind. Logging primitives are the small, boring, non-negotiable pieces of information a team writes down at request time so that every later question about quality — did this regress, why did it fail, can we reproduce it, is the judge lying to us — has an answer instead of a shrug. 🧱

The stakes are not abstract. When Klarna's OpenAI-powered assistant scaled to roughly 2.3 million conversations in its first month, handling about two-thirds of the company's customer service chats, the company was managing a system operating at a volume where a single unlogged failure mode could touch hundreds of thousands of real customers before a human ever noticed. Public reporting later described Klarna dialing the assistant back toward a hybrid human-plus-AI model — a reminder that impressive rollout numbers and durable quality are two different claims, and only one of them can be checked against a log. A team that can't reconstruct "what exactly did the model see, and who scored it, and against what version" is flying blind at exactly the volume where blindness is most expensive. 🚨

Diagram showing five stacked logging layers: Identity, Contract, Grounding, Judgment, and Fleet

🔀 Quick Comparison: Three Logging Altitudes

"Logging" isn't one activity — it happens at three altitudes, and mixing them up is where most instrumentation plans go wrong.

Altitude Answers the question Lives in Retention
Request-level "What happened on this one call?" Trace/span store (e.g. an OpenTelemetry backend) Days to weeks
Run-level "Did this batch of changes make things better or worse?" Evaluation/experiment store, tied to a dataset snapshot Months, versioned like code
Fleet-level "Is the whole system healthy right now, across every user?" Metrics/alerting dashboards Rolling windows, aggregated

The five layers below aren't a fourth altitude — they're the actual fields that get written at all three altitudes. Layer 1 and 2 mostly live at request-level, Layer 4 bridges request-level and run-level, and Layer 5 is what turns request-level detail into fleet-level signal.

1️⃣ The Identity Layer: proving which request you're even looking at

Kid analogy: imagine a school where every kid's homework paper is unsigned and undated. If the teacher says "this answer is wrong," nobody can tell whose paper it was, which assignment it belonged to, or whether it was even graded against this week's rules or last week's. The Identity Layer is the name, date, and assignment number written at the top of the page — without it, nothing else on the page can be trusted.

In plain terms, this layer is five fields on every single request:

  • a request ID — one unique tag for this exact call
  • trace/span IDs — the thread that ties this call to everything else that happened during the same request
  • the prompt template + version that was actually rendered
  • the model requested vs. the model that actually answered (these are not always the same)
  • the inference settings — temperature, max tokens — that shaped the reply

Here's what that looks like as a real log line, not a theory:

{
  "request_id": "req-84f21a",
  "trace_id": "trc-9c0e77",
  "prompt_version": "refund-flow-v14",
  "model_requested": "support-model",
  "model_responded": "support-model-2026-06-11",
  "temperature": 0.2
}

Real example. This isn't a made-up wishlist — the industry is actively standardizing exactly these fields. OpenTelemetry's GenAI working group has been building a shared set of gen_ai.* attributes since 2024 (provider name, requested model, responding model, token counts) so a call to OpenAI and a call to Anthropic produce telemetry that looks the same shape. It's still marked experimental in 2026, but adopting that shared vocabulary now beats inventing your own and migrating later.

✅ Worked example: A support-chat feature logs request_id, prompt_version=refund-flow-v14, and model=support-model-2026-06 on every call. When a customer complains about a wrong refund amount three weeks later, the on-call engineer pulls up that exact request ID and sees precisely which prompt version and model generated the reply — no guesswork about "which version were we even running that week."

💡 Harder case: a provider silently routes a request to an updated model snapshot behind the same model name. If you only log the model name you asked for, and not the model version that actually replied, an entire class of regressions becomes invisible — the logs will insist nothing changed, right up until the fleet-level metrics say otherwise.

🎯 Use this when: you need to reproduce, audit, or explain any single past response — this is the layer that makes "show me exactly what happened" answerable at all.

2️⃣ The Contract Layer: defining what "correct" was supposed to mean

Kid analogy: a rubric handed out before an essay is due is completely different from a rubric invented after the essay is already graded. If the teacher decides afterward what counts as a good essay, every grade is really just an opinion dressed up as a score. The Contract Layer is the rubric written down before the model ever answers — what shape the answer must take, how risky getting it wrong is, and which rules it's not allowed to break.

In plain terms, three things go here:

  • a risk tier — a password reset and a casual product tip do not deserve the same scrutiny
  • the output shape a correct answer must have (required fields, a schema to validate against)
  • which policy version applied at the time (a disclaimer rule, a tone rule, a refusal rule)
{
  "risk_tier": "high",
  "required_fields": ["order_id", "amount", "eta"],
  "schema_valid": true,
  "policy_version": "disclaimer-policy-v2"
}

Real example. A financial-services chatbot running at Klarna's scale carries real regulatory risk — a wrong refund amount is not the same mistake as an awkward sentence. That's why not every request can be treated the same way: some need to be flagged high-risk, and every one of them needs a record of what "correct" was supposed to look like, before anyone grades the answer.

✅ Worked example: the same support-chat feature from Layer 1 tags every refund-related request as risk_tier=high and requires a schema_valid=true/false flag on the output. A grading pipeline can then treat a malformed refund response as an automatic failure, without needing a human or a judge model to weigh in at all.

💡 Harder case: the policy itself changes mid-quarter — a new disclaimer requirement rolls out. If the applicable policy version isn't logged alongside the response, a request answered correctly under the old policy will look like a failure under the new one during retroactive review, and nobody can tell the difference between "the model broke a rule" and "we're grading it against a rule that didn't exist yet."

🎯 Use this when: you're building any scoring rubric, automated grader, or pass/fail gate — the contract has to exist and be logged before a score means anything.

3️⃣ The Grounding Layer: the receipts behind the answer

Kid analogy: think of a kid doing a research project who's asked to show their work — which books they used, which page had the fact, whether they double-checked it. A confident-sounding answer with no receipts could be right, or could be a guess dressed up as a fact. The Grounding Layer is those receipts: what the system actually looked at and did before it spoke.

In plain terms, log two things every time the system looks something up or uses a tool:

  • retrieval: which document IDs came back, their rank/score, and which index version was live
  • tool calls: which tool, what arguments, success or failure, latency, and whether it retried or timed out
{
  "tool": "order_lookup",
  "args": {"order_id": "4471"},
  "tool_status": "success",
  "latency_ms": 210,
  "retried": false
}

Real example. Most "wrong answer" bugs in RAG and agent systems aren't the model reasoning badly — they're a retrieval miss or a failed tool call wearing a model failure's disguise. That's why teams building on tools like LangSmith turn real production traces (the actual retrieval and tool steps a request took) directly into evaluation datasets, instead of writing synthetic test cases that skip the exact steps that break.

✅ Worked example: the refund assistant retrieves the customer's order record via a tool call before answering. The trace logs tool=order_lookup, the order ID it fetched, and a tool_status=success flag. When the refund amount quoted turns out to be wrong, the trace shows immediately whether the model misread a correct order record (a model problem) or was handed a stale one (a data problem).

💡 Harder case: the order-lookup tool times out, and a fallback path silently answers from the model's general knowledge instead. If the fallback trigger and the timeout aren't logged, the resulting answer looks identical in the logs to a normal, properly-grounded response — until a customer disputes a refund number that was never real.

🎯 Use this when: debugging any RAG, agent, or tool-using system — this is almost always where the real failure lives, even when the symptom shows up in the final text.

4️⃣ The Judgment Layer: making sure the scorer can be trusted too

Kid analogy: imagine two different teachers grading the same stack of essays with two different answer keys, and nobody wrote down which teacher used which key. A student's grade going from a B to a D between two rounds might mean their writing got worse — or it might just mean a stricter teacher graded the second round. The Judgment Layer is the label on the answer key itself: which grader, using which rules, graded this.

In plain terms, log the grader the same way you'd log anything else that produces an important number:

  • which judge model + judge prompt version was used (if it's an LLM-as-judge)
  • which rubric version it graded against
  • anything that was checked by a deterministic rule instead of a judge's opinion
{
  "judge_model": "grader-2026-03",
  "rubric_version": "refund-rubric-v3",
  "result": "pass",
  "graded_by": "llm_judge"
}

Real example. OpenTelemetry's GenAI conventions are now extending to a standard event just for evaluator results — treating a judge's score as its own logged record, not a number that shows up unexplained in a spreadsheet. That matters because judge drift is real: swap the judge model without logging it, and every past score becomes an apples-to-oranges comparison.

✅ Worked example: the refund assistant's automated grader logs judge_model=grader-2026-03 and rubric_version=refund-rubric-v3 next to every score. When quarterly accuracy appears to jump five points, the team's first check is whether the judge or rubric changed that quarter — before anyone celebrates a model improvement that never actually happened.

💡 Harder case: the same model family is used as both the thing being tested and the judge doing the testing. Without a logged, independent record of judge identity and a periodic bias check against human ratings, a model can end up subtly rewarding its own stylistic habits, and the score will look flawless right up until real users disagree with it.

🎯 Use this when: comparing scores across time, across model versions, or across teams — never trust a score trend without confirming the grader stayed constant underneath it.

5️⃣ The Fleet Layer: connecting one request to the whole system

Kid analogy: one kid's grade tells you about one kid. A whole class's grades, watched over time, tell you whether the teaching itself is working — and whether it's time to change the lesson plan for everyone. The Fleet Layer is the class-wide report card: live quality signals, alert thresholds, and links to whatever rollout decision they triggered.

In plain terms, this is a live dashboard fed by every request below it:

  • live quality/safety scores sampled from real traffic
  • alert thresholds that turn a dip into a page, not a missed footnote
  • canary metrics comparing a new version against the old one on a slice of traffic
  • a link from every incident back to the exact run or model version that caused it
{
  "metric": "schema_valid_rate",
  "canary_value": 0.91,
  "baseline_value": 0.98,
  "alert": "fired",
  "linked_run_id": "run-2026-09-10-canary"
}

Real example. Public reporting on Klarna's assistant is a real-world version of this: a rollout that looked spectacular in month-one adoption numbers was later described as being rebalanced toward a hybrid human-and-AI model. That's exactly the pattern live, fleet-wide signal exists to catch early — before a headline does it instead.

✅ Worked example: a canary rollout sends 5% of refund traffic to a new prompt version, and a dashboard tracks its schema-validity rate against the existing version's baseline in real time. The moment the canary's validity rate drops more than an agreed threshold, an alert fires and links straight to the offending run — no waiting for the next weekly eval report.

💡 Harder case: a quality metric degrades slowly across weeks rather than dropping sharply. A static alert threshold tuned for sudden failures will miss a slow drift entirely — this is exactly the class of problem that needs a trend-aware alert, not just a floor value.

🎯 Use this when: deciding whether to ship, roll back, or hold a change — fleet-level signal is the only layer that speaks for the whole user base at once.

🧪 Hands-On Lab: Build a Five-Layer Log Line Yourself

This is a small, disposable exercise — no production system required, just a text editor. The goal is to feel, in your own hands, why a log line with all five layers is a completely different object than a log line with just the model's raw output.

1
Open a plain text file and write down a fake customer question, e.g. "Can I get a refund for order #4471?" Give it a made-up request_id: req-0001 — this is your Identity Layer starting point.
2
Under it, add risk_tier: high and required_fields: [amount, order_id, eta]. You've just written a Contract — before you write the model's answer, you've written what a correct answer has to contain.
3
Make up a fake tool result, e.g. order_lookup(4471) → status: shipped, amount: $42.00, and write it down as its own line, separate from the model's eventual answer. That separation is the whole point of the Grounding Layer — the evidence is visible on its own, not baked invisibly into the final sentence.
4
Now write a made-up model answer, then grade it yourself against step 2's contract, and log the grade as graded_by: you, rubric: refund-rubric-v1, result: pass. Expect to see: a grade that references a specific rubric name, not just "pass" floating with nothing attached to it.
5
Imagine ten more fake requests like this one and picture them all rolling up into one number: "9/10 passed." That single rolled-up number is your first, tiny taste of the Fleet Layer — and notice you can only trust it because every request behind it carries the same four layers underneath.

Common first-timer mistake: writing the grade ("pass") without writing down which rubric version produced it. Go back and add the rubric name to every grade before moving on — this is the exact habit the Judgment Layer is trying to build.

Bridge to production: everything you just wrote by hand — a request ID, a contract, a piece of evidence, a versioned grade — is precisely what tools like LangSmith, Arize Phoenix, or a custom OpenTelemetry pipeline automate at scale, on every real request, without you typing a single line by hand.

🏢 Rolling This Out at Enterprise Scale

A five-layer log line for one request is easy. Making it a durable, governed practice across dozens of teams and millions of requests is a different problem, and it tends to break in the same handful of places.

  1. One team owns the schema. If every squad invents its own field names for these five layers, cross-team comparisons break exactly when leadership asks for them.
  2. Version your golden dataset like code. A dataset frozen six months ago reflects six-month-old user behavior — put a date and a review cycle on it.
  3. Gate merges on eval results. A prompt or model change shouldn't ship until it clears the current eval suite — treat an eval regression exactly like a failing unit test.
  4. Lock down access to eval data. Grounding and Identity data can contain real customer PII. Same masking and retention rules as any other production data — no "it's just for evals" exemption.
  5. Sample judge calls, don't grade everything. LLM-as-judge calls cost real tokens. Above a few hundred thousand requests a day, judging a sample is usually the only affordable option.
  6. Keep training-time and inference-time dashboards separate. "How the model did in fine-tuning" and "how it's doing in front of users this week" are different questions — merging the dashboards hides the drift you're trying to catch.
  7. Alert on the trend, not just the floor. A metric that's "fine" but down 15% week-over-week deserves its own alert, separate from a static threshold.

🎯 Use this when: more than one team is producing or consuming eval data — governance problems that don't matter at ten requests a day become blocking problems at ten million.

⚠️ Common Mistakes

Each of these shows up repeatedly, and each one traces back to skipping one of the five layers above.

  • Relying on a single aggregate score. A single "quality: 87%" number hides which layer actually failed. A metric suite that separately reports schema validity, groundedness, and judge confidence lets a team fix the right thing instead of guessing.
  • Using the same model as both generator and judge, with no bias check. Without a logged, independently-tracked judge identity (Layer 4) and a periodic comparison against human ratings, a model can quietly reward outputs that resemble its own style rather than outputs that are actually correct.
  • No held-out test set, or eval-on-training-data leakage. If the data used to grade a model overlaps with data the model (or a fine-tuned version of it) has already seen, the resulting score measures memorization, not the generalization that will actually be exercised in production.
  • Treating latency and cost as someone else's problem. A response that's technically correct but arrives after the user has given up, or costs ten times the acceptable per-request budget, is a production failure — it belongs in the same eval suite as correctness, not a separate infrastructure dashboard nobody on the eval team looks at.
  • Treating an offline eval pass as the finish line. A model can clear every golden-dataset test and still fail on the long tail of real traffic the golden dataset never anticipated — that's exactly what the Fleet Layer's live telemetry exists to catch, and skipping it turns "we tested it" into a false sense of security.
  • Letting golden datasets go stale. User behavior, product features, and adversarial patterns shift constantly. A dataset frozen at launch grades a system against a world that no longer exists, and a passing score on it stops meaning very much.

❓ FAQ

Do I need all five layers from day one?

No — a small prototype can start with just the Identity and Contract layers. But add Grounding as soon as retrieval or tools enter the system, and add Judgment the moment an automated grader (rather than a human) starts producing scores anyone will act on.

Isn't logging full prompts and completions a privacy risk?

Yes, if done carelessly — full request/response text can contain customer PII. Treat the Grounding and Identity layers with the same masking, access control, and retention limits as any other customer data store, not as an exemption because it's "for evals."

What's the difference between the Contract Layer and the Judgment Layer?

The Contract Layer is written before the model answers — it defines what "correct" means. The Judgment Layer is written after — it records who actually checked the answer against that contract, and with what tool. Confusing the two is how a grading rubric quietly turns into a moving target.

Should I build my own logging schema or adopt an existing standard?

Adopt an existing, evolving standard like the OpenTelemetry GenAI semantic conventions where you can, even though it's still experimental — a shared vocabulary means your telemetry works with off-the-shelf dashboards and tools instead of requiring custom parsing for every backend you ever connect.

How do I know if my logging is "enough"?

A useful test: pick a real incident from the past month and try to answer, using only your logs, which prompt version ran, what evidence it saw, who graded it, and how the fleet-wide metric reacted. Any question you can't answer marks the layer you're missing.

🔗 References & Further Reading

Official/primary documentation used for fact-checking:

  • OpenTelemetry — GenAI Semantic Conventions repository and specification (opentelemetry.io / github.com/open-telemetry)
  • OpenAI — API and platform documentation (platform.openai.com/docs)
  • Anthropic — API and platform documentation (docs.claude.com)
  • LangChain/LangSmith — official evaluation and observability documentation (docs.smith.langchain.com)

This post synthesizes and explains general industry practice in its own words; it does not reproduce text from any source verbatim. Product and company names (OpenTelemetry, OpenAI, Anthropic, LangSmith, Klarna, and others) are trademarks of their respective owners and are referenced here for identification and educational purposes only, with no endorsement implied.

📝 Summary

  • The Identity Layer proves which exact request, prompt version, and model you're looking at.
  • The Contract Layer writes down what "correct" meant before anyone graded the answer.
  • The Grounding Layer keeps the receipts — retrieved evidence and tool calls — separate and visible.
  • The Judgment Layer makes sure the scorer itself is versioned and trustworthy, not just the score.
  • The Fleet Layer rolls thousands of individually-logged requests into one live, actionable signal.
  • The hands-on lab shows all five layers by hand, in miniature, before any tooling is involved.
  • At enterprise scale, ownership, versioning, CI-gating, governance, cost control, and drift-aware alerting turn this from a good habit into a durable system.
  • Every common mistake in this space is really just one of the five layers, skipped.

None of these five layers are exciting. That's exactly why they're the ones worth getting right first — a clever metric built on top of ungoverned logging is a house built on a foundation nobody inspected. Log the boring fields, and the interesting evaluation questions actually become answerable. 🧭

Comments