This is the shared vocabulary every LLM reliability and evaluation engineer needs before they can read a dashboard, write a test plan, or debug a production incident without guessing what a term means. None of these words are decoration — each one maps to a specific decision someone made about how a system is measured, gated, or trusted. 📘
📑 In This Post
- Core evaluation foundations
- Output quality and grading
- Classical metrics and measurement
- Reliability, performance, and latency
- Experiment design and statistical confidence
- Evaluation dataset quality
- RAG evaluation vocabulary
- AI agent evaluation vocabulary
- Safety, security, and governance
- Cost and operational efficiency
- Summary
📌1. Core Evaluation Foundations
- Task success rate — the share of test cases (or real interactions) where the system actually achieved the user's goal, not just produced a plausible-looking response.
- Evaluation dataset / eval set — the fixed collection of inputs, and usually expected outputs, that a model or system is scored against; its quality sets the ceiling on how much you can trust any score.
- Golden dataset / golden set — a small, carefully curated, high-trust subset of the eval set, usually reviewed by experts, used as the stable long-term benchmark for comparing model versions over time.
- Ground truth / reference answer — the accepted correct output for a test case, against which a model's actual output is compared; for open-ended tasks this may be a rubric rather than one exact string.
- Test case — a single input (plus context) paired with its expected behavior, ground truth, or scoring criteria — the atomic unit an eval set is built from.
- Rubric — a written set of criteria describing what makes an answer good, bad, or partially correct, used to keep human or LLM-judge scoring consistent across many raters and runs.
- Metric — a specific, defined measurement computed from outputs (accuracy, latency, groundedness score) that turns raw behavior into a comparable number.
- Quality threshold — the minimum acceptable value for a metric, below which an output or a release is considered a failure rather than a pass.
- Quality gate / release gate — the checkpoint in a pipeline where a model or prompt change must clear defined thresholds before it's allowed to progress toward production traffic.
- Baseline — the current production system's performance, used as the reference point a new candidate must match or beat before it's allowed to replace it.
- Candidate system — the new model, prompt, or configuration being evaluated for possible promotion, compared directly against the baseline.
- Regression — a drop in performance on something that used to work, introduced by a change that was intended to improve something else.
- Error analysis — the manual, structured process of reading through failures to find patterns, root causes, and categories of mistake, rather than stopping at "the score went down."
- Data slice — a defined subset of the eval set or production traffic (a language, a task type, a user segment) that's scored separately so an aggregate number can't hide a localized problem.
- Production monitoring — the ongoing collection of metrics and signals from a live system, used to catch problems that only appear after release, not just at launch.
- Observability — the broader capability to ask new questions about a running system's behavior after the fact, using logs, metrics, and traces, without having predicted every question in advance.
- Trace — the full recorded path of a single request through a system — every step, call, and decision it took — used to reconstruct exactly what happened for one specific interaction.
- Span — one timed, named segment within a trace (a single retrieval call, a single model call), the building block traces are made of.
📌2. Output Quality and Grading
- Correctness — whether the factual content or logic of an answer is actually right, independent of how well it's written or presented.
- Relevance — whether the answer actually addresses what was asked, rather than being accurate about something adjacent or off-topic.
- Completeness — whether the answer covers everything the question required, rather than a technically correct but partial response.
- Helpfulness — a holistic judgment of whether the answer actually moved the user closer to solving their problem, combining correctness, relevance, and tone.
- Groundedness — whether the claims in an answer are actually supported by the source material or context the model was given, rather than invented.
- Faithfulness — closely related to groundedness: whether the answer accurately represents what the source material says, without distorting or overstating it.
- Hallucination — content the model states with confidence that is false, unsupported, or fabricated — the single most reputationally expensive failure mode in production LLM systems.
- Citation correctness — whether a cited source actually says what the model claims it says, checked citation by citation, not just that a citation exists.
- Citation coverage — the share of factual claims in an answer that are backed by a citation at all, regardless of whether each citation is individually correct.
- Format adherence — whether the output follows the requested structure (length, tone, sections, style) that was specified in the prompt or system instructions.
- Structured output — output constrained to a specific machine-readable shape, such as JSON or a table, rather than free-form prose.
- Schema validity — whether a structured output actually parses and conforms to its defined schema — the field names, types, and required keys it was supposed to have.
- Deterministic validator — a non-LLM check (a regex, a parser, a rule) that verifies a specific, objective property of an output with no ambiguity or judgment involved.
- Rule-based evaluation — scoring built from explicit, hand-written rules rather than a model's judgment — fast, cheap, and reliable for narrow, well-defined checks.
- LLM-as-a-judge — using a separate model call to score or compare outputs against a rubric, scaling evaluation far beyond what human review can cover, at the cost of inheriting the judge's own biases.
- Judge prompt — the instructions given to the judge model describing exactly how to score or compare outputs, which functions as the rubric translated into an executable form.
- Judge calibration — the process of checking and adjusting a judge model's scores against human labels, so its verdicts can be trusted to track real quality rather than the judge's own quirks.
- Human evaluation — scoring or comparison performed by trained people rather than an automated metric or judge, used where nuance and judgment matter most.
- Pairwise comparison — presenting a rater (human or LLM) with two outputs side by side and asking which is better, rather than scoring each output in isolation.
- Preference evaluation — evaluation built around which of two or more outputs people prefer, rather than an absolute quality score — often more reliable than absolute scoring for subjective tasks.
- Inter-annotator agreement — the degree to which different human raters give the same label to the same example, used to check whether a rubric is clear and consistently applied.
📌3. Classical Metrics and Measurement
- Accuracy — the share of predictions that were correct out of all predictions made; simple, but misleading on its own when classes or outcomes are imbalanced.
- Precision — out of everything the system flagged as positive, the share that was actually positive — high precision means few false alarms.
- Recall — out of everything that was actually positive, the share the system successfully caught — high recall means few misses.
- F1 score — the harmonic mean of precision and recall, giving one combined number when you care about both and neither alone tells the full story.
- Exact match — a strict scoring rule where the output must match the reference answer character-for-character (or after light normalization) to count as correct.
- Partial match — a scoring rule that gives partial credit when an output is close to correct but not identical, more forgiving than exact match for open-ended tasks.
- False positive — the system flagged something as true, unsafe, or positive when it actually wasn't — the source of unnecessary alarms or wrongful blocks.
- False negative — the system missed something that actually was true, unsafe, or positive — often the more dangerous error type in safety-critical checks.
- True positive — the system correctly flagged something that actually was positive.
- True negative — the system correctly did not flag something that actually was negative.
- Confusion matrix — a table laying out true/false positives and negatives together, showing exactly what kinds of mistakes a classifier makes, not just how often it's wrong.
- Classification threshold — the cutoff score above which a model's continuous output is treated as a positive prediction; moving it trades precision against recall.
- Calibration — whether a model's confidence scores actually match its real accuracy — a well-calibrated model that says "80% confident" is right about 80% of the time.
- Confidence score — a number a model or classifier attaches to its own prediction, meant to indicate how likely that prediction is to be correct.
- Pass@k — the probability that at least one of k independently sampled outputs for a task is correct, commonly used to evaluate code generation and other tasks where you can generate multiple attempts.
📌4. Reliability, Performance, and Latency
- Latency — the time between a request being sent and a response being received; the metric users feel most directly.
- Throughput — the number of requests (or tokens) a system can process per unit of time at a given level of resources.
- Availability — the share of time a service is actually up and successfully serving requests, usually expressed as a percentage over a period.
- Error rate — the share of requests that fail outright, as opposed to succeeding with poor quality — a distinct problem from low-quality-but-successful responses.
- Timeout rate — the share of requests that failed specifically because they took longer than the system was willing to wait.
- Retry rate — the share of requests that had to be attempted more than once before succeeding, a useful early signal of instability even before outright failures appear.
- Fallback rate — the share of requests served by a backup path (a smaller model, a cached answer, a simpler rule) instead of the primary system, usually because the primary was unavailable or too slow.
- Queue time — the time a request spends waiting for a resource to become available before processing even starts, separate from actual processing time.
- Concurrency — the number of requests being processed by a system at the same moment, which drives resource contention and, past a point, latency.
- Rate limit — a cap on how many requests a client or the system as a whole is allowed to make in a given time window, used to protect capacity and control cost.
- Tail latency — the latency experienced by the slowest fraction of requests, which matters disproportionately because it's what your most frustrated users actually feel.
- p50 / median — the latency value below which half of all requests complete; the "typical" experience, but blind to how bad the worst requests are.
- p90 — the latency value below which 90% of requests complete; the slowest 10% take longer than this.
- p95 — the latency value below which 95% of requests complete; the slowest 5% take longer than this.
- p99 — the latency value below which 99% of requests complete; the slowest 1% take longer than this, and at real scale that 1% can still be thousands of unhappy users.
- Time to First Token (TTFT) — how long a user waits after sending a request before the model starts streaming back any output at all — the metric that drives perceived responsiveness.
- Time to Last Token — the total time from request to the very last token of the response, i.e. full completion time.
- Tokens per second — the rate at which a model generates output once it has started, which determines how fast a long response finishes streaming.
- Service Level Indicator (SLI) — the actual measured value of a metric that matters to users, such as real latency or real error rate — the "what we measured."
- Service Level Objective (SLO) — the internal target you set for an SLI (for example, p95 latency under a set number of seconds), used to drive engineering priorities.
- Service Level Agreement (SLA) — a formal, often contractual commitment to a service level, typically with consequences if it's not met — stricter and more external-facing than an SLO.
- Error budget — the amount of unreliability you're allowed to spend against an SLO before it's breached, used to balance shipping speed against stability deliberately rather than by accident.
✅ Percentile latency, made concrete: if p95 latency is 8 seconds, that means 95% of requests complete in 8 seconds or less, while the slowest 5% take longer than 8 seconds. p50 tells you what a typical user feels; p95 and p99 tell you how bad it gets for the users having the worst experience — and at production scale, that slowest slice can still be a very large number of real, frustrated people.
📌5. Experiment Design and Statistical Confidence
- Offline evaluation — running a candidate against a fixed, versioned test set before any real user is exposed to it; fast and repeatable, but limited to what the test set anticipated.
- Online evaluation — evaluating a candidate against real, live traffic once it's actually serving some portion of real users.
- Continuous evaluation — evaluation that runs on an ongoing basis after full rollout, not just at launch, to catch drift and decay that appear later.
- A/B testing — randomly splitting traffic between two versions and comparing their outcomes statistically, to attribute any difference to the change itself rather than chance.
- Canary rollout — gradually increasing the share of traffic sent to a new version while the old version keeps serving the rest, so a regression affects a small, controlled slice of users first.
- Shadow testing — sending a copy of real traffic to a candidate system without showing its responses to users, so its real-world behavior can be checked with zero user-facing risk.
- Holdout set — a portion of data deliberately kept out of training and tuning, reserved purely for evaluating how the system generalizes to unseen cases.
- Control group — in an experiment, the group that continues to receive the existing (baseline) system, used as the point of comparison.
- Treatment group — in an experiment, the group that receives the new candidate system being tested.
- Sample size — the number of observations a test or comparison is based on; too small a sample size means real effects can be invisible and noise can look like a pattern.
- Variance — how spread out a metric's values are; high variance means a single measurement is a less trustworthy estimate of the true value.
- Confidence interval — a range around a measured value that likely contains the true value, communicating how much uncertainty a single point estimate is hiding.
- Statistical significance — a measured difference is unlikely to be due to random chance alone, given the sample size and variance observed.
- Effect size — how large a difference actually is in practical terms, independent of whether it's statistically significant — a tiny but statistically significant difference may not matter to anyone.
- Statistical power — the probability that a test will correctly detect a real effect if one actually exists, given the sample size and expected effect size.
- Bootstrap confidence interval — a confidence interval built by repeatedly resampling the observed data itself, useful when a metric's true distribution is unknown or doesn't fit a standard formula.
- Practical significance — whether a statistically real difference is actually large enough to matter for users or the business, distinct from statistical significance alone.
- Rollback — reverting a release back to the previous known-good version, typically triggered when a monitored metric crosses a defined failure threshold.
📌6. Evaluation Dataset Quality
- Representativeness — how closely an eval set's mix of examples matches the real distribution of what the system actually encounters in production.
- Coverage — how completely an eval set spans the range of task types, languages, and scenarios the system needs to handle, rather than clustering around a few easy cases.
- Stratified sampling — deliberately sampling a test set so specific subgroups (languages, task types, user segments) are proportionally represented, rather than left to chance.
- Edge case — an unusual, boundary, or rarely occurring input that still needs to be handled correctly, and is disproportionately likely to expose a bug.
- Adversarial test case — an input deliberately crafted to break, mislead, or exploit a system, used to test robustness rather than typical behavior.
- Red teaming — a structured, ongoing effort by people specifically trying to find safety and robustness failures in a system before real attackers do.
- Synthetic data — artificially generated examples (often by another model) used to expand or diversify an eval or training set beyond what was collected naturally.
- Train split — the portion of data used to actually teach or fine-tune a model.
- Validation split — the portion of data used during development to tune choices and catch problems before final evaluation, kept separate from training.
- Test split — the portion of data reserved for final, unbiased evaluation, untouched during training or tuning decisions.
- Dataset versioning — tracking exactly which version of a dataset was used for a given training run or evaluation, so results can be reproduced or audited later.
- Data drift — a gradual change in the statistical properties of incoming data over time, compared to what a model was built or tested against.
- Distribution shift — a broader, sometimes sudden mismatch between the data a system now sees and the data it was trained or evaluated on, which can silently degrade performance.
- Data leakage — information from outside the intended training or evaluation scope (often from the answer itself, or from the test set) accidentally influencing a result, making performance look better than it really is.
- Benchmark contamination — a specific form of leakage where benchmark test data has ended up inside a model's training data, inflating its reported score on that benchmark.
- Near-duplicate detection — identifying examples that are not identical but close enough in meaning or wording to count as leakage, typically using embedding similarity rather than exact-match checks.
- Annotation guideline — the written instructions given to human labelers describing exactly how to label or score examples consistently.
- Label quality — how accurate, consistent, and well-justified the human-assigned labels in a dataset actually are — a training and evaluation set is only as trustworthy as its labels.
📌7. RAG Evaluation Vocabulary
- Retrieval — the step where a system searches an external knowledge source for information relevant to a query, before generating an answer.
- Retrieval precision — out of the documents or chunks retrieved, the share that were actually relevant to the query.
- Retrieval recall — out of all the truly relevant documents that existed, the share the retrieval step actually found.
- Context precision — how much of the retrieved context that gets passed to the model is actually useful, rather than irrelevant padding.
- Context recall — whether the retrieved context contains everything needed to fully and correctly answer the question.
- Answer relevance — whether the final generated answer actually addresses the user's question, independent of whether the retrieval step did its job well.
- Answer faithfulness — whether the final answer accurately reflects what the retrieved context actually says, without adding unsupported claims.
- Context utilization — how much of the retrieved, relevant context the model actually used in constructing its answer, as opposed to ignoring it.
- Chunking strategy — the method used to split source documents into smaller pieces for indexing and retrieval, which strongly affects both retrieval quality and answer faithfulness.
- Embeddings — numeric vector representations of text that capture meaning, allowing similarity between pieces of text to be computed mathematically.
- Reranking — a second pass that reorders an initial set of retrieved candidates by relevance, usually with a more precise (and more expensive) model than the first retrieval pass.
- Top-k retrieval — returning the k most relevant results for a query, where k is a tunable number balancing completeness against noise and cost.
- Retrieval failure — a case where the retrieval step fails to surface any genuinely relevant document, which downstream generation cannot recover from no matter how good the model is.
- Citation grounding — verifying that every cited source in an answer actually supports the specific claim it's attached to, not just that a citation is present.
- Abstention — the system correctly declining to answer when it doesn't have sufficient grounded information, rather than guessing.
- No-answer quality — how well a system handles the case where the honest answer is "I don't know" or "this isn't covered" — a distinct skill from answering well when it does have the information.
- Knowledge freshness — how up to date the information in the retrieval index actually is relative to the real world, which directly caps how current a RAG system's answers can be.
📌8. AI Agent Evaluation Vocabulary
- AI agent — a system that uses an LLM to plan and take a sequence of actions (often calling tools or other systems) toward a goal, rather than just producing a single text response.
- End-to-end task success — whether the agent actually achieved the overall goal by the end of its run, the single most important agent metric, distinct from how good any individual step looked.
- Tool selection accuracy — whether the agent chose the correct tool for a given step, out of the tools available to it.
- Function calling — the mechanism by which a model outputs a structured request to invoke an external function or API, rather than only generating text.
- Tool-call validity — whether a function call the model generated is well-formed and executable — correct name, correct structure — regardless of whether it was the right call to make.
- Argument accuracy — whether the specific values the agent passed into a tool call were correct, not just whether the call's structure was valid.
- Tool execution success rate — the share of tool calls that actually completed successfully once executed, as opposed to failing at runtime.
- Tool error handling — how well an agent responds when a tool call fails — whether it retries sensibly, informs the user, or tries an alternative, rather than pretending the call succeeded.
- Trajectory — the full sequence of steps, decisions, and tool calls an agent took to complete (or fail) a task, evaluated as a whole rather than one action at a time.
- Step-level evaluation — scoring each individual action in a trajectory on its own merits, used to pinpoint exactly where a multi-step task went wrong.
- Plan quality — whether the sequence of steps an agent chose was logical and efficient for the goal, independent of whether it happened to reach the right outcome anyway.
- Multi-turn evaluation — scoring a full conversation or session as a whole, since individually reasonable replies can still add up to a session that never actually resolved the user's need.
- State management — how correctly a system tracks and updates the relevant facts about an ongoing task or conversation as it progresses.
- Memory correctness — whether information a system is supposed to remember across turns or sessions stays accurate and isn't corrupted, dropped, or confused with something else.
- Context management — how well a system selects, prioritizes, and trims what information stays available to the model as a conversation or task grows longer than its context window.
- Retry — automatically attempting a failed action again, typically after a brief delay, before treating it as a hard failure.
- Backoff — increasing the delay between successive retry attempts, so repeated failures don't hammer an already-struggling dependency.
- Fallback — switching to an alternative method, tool, or simpler response when the primary approach fails or isn't available.
- Circuit breaker — a mechanism that stops sending requests to a dependency after it starts failing repeatedly, to prevent cascading failure while it's unhealthy.
- Escalation — handing a task off to a human (or a higher-capability system) when an agent can't complete it safely or successfully on its own.
- Human-in-the-loop — a design where a person reviews, approves, or can intervene in an agent's actions, rather than letting it act fully autonomously.
- Autonomy rate — the share of tasks an agent completes successfully without needing human intervention or escalation.
- Recovery rate — the share of times an agent successfully gets back on track after hitting an error or an unexpected state, rather than failing outright.
- Safe tool use — an agent's ability to use powerful or risky tools (payments, file deletion, external messaging) only within intended, safe boundaries.
- Idempotency — the property that performing the same action multiple times has the same effect as performing it once, which is what makes safe retries of tool calls possible.
📌9. Safety, Security, and Governance
- Policy compliance — whether a system's outputs and actions stay within the organization's defined rules for acceptable content and behavior.
- Safety evaluation — testing specifically aimed at finding harmful, dangerous, or policy-violating outputs, run separately from general quality evaluation.
- Prompt injection — an attempt to override a system's original instructions by embedding new, malicious instructions inside user input or content the model processes.
- Jailbreak — a technique designed to manipulate a model into bypassing its safety guidelines and producing content or actions it would normally refuse.
- Indirect prompt injection — a prompt injection attack delivered through content the model reads as data (a webpage, a document, an email) rather than typed directly by the user.
- Data exfiltration — a system being manipulated into leaking sensitive information it had access to, to an unauthorized party.
- PII leakage — a system exposing personally identifiable information in its output when it should have withheld or redacted it.
- Sensitive-data handling — the set of practices governing how a system detects, protects, and limits exposure of sensitive information throughout its pipeline.
- Authorization boundary — the defined limit of what actions and data a system or agent is permitted to access, which every safety-relevant check assumes is being enforced.
- Least privilege — the principle of giving a system or agent only the minimum access it needs to do its job, and nothing more, to limit the damage of any single failure or exploit.
- Unsafe tool use — an agent invoking a tool in a way that causes real-world harm or risk, whether from a bad decision, a manipulation, or a bug.
- Correct refusal — the system appropriately declining a genuinely unsafe or out-of-policy request, which is the desired behavior, not a failure.
- Over-refusal — the system incorrectly declining a legitimate, safe request, mistaking it for something it should refuse — a real quality failure, not a safe default.
- Under-refusal — the system failing to decline a request it should have refused, letting an unsafe or policy-violating output through.
- Audit trail — a persistent, reviewable record of what a system did, when, and why, used to reconstruct events after the fact for debugging, compliance, or incident review.
- Governance — the overall structure of ownership, policy, and approval processes that decides how models and systems are allowed to be built, tested, and released.
- Incident response — the defined process for detecting, containing, and recovering from a production failure or safety event, including who is responsible for each step.
📌10. Cost and Operational Efficiency
- Input tokens — the units of text sent into a model as part of a request (the prompt, context, and history), which is one basis models are typically billed on.
- Output tokens — the units of text a model generates in its response, usually billed at a different rate than input tokens.
- Token usage — the total volume of input and output tokens consumed, the core driver of both cost and latency for a given workload.
- Cost per request — the average cost of serving a single request, combining token usage, model choice, and any tool or retrieval costs involved.
- Cost per successful task — cost measured against actual successful outcomes rather than raw requests, which is what really matters when some requests fail or need retries.
- Token efficiency — how much useful output or task progress a system achieves per token consumed, a measure of how "wasteful" a prompt or workflow is.
- Cache hit rate — the share of requests answered from a cached previous result instead of a fresh model call, directly reducing both cost and latency.
- Prompt caching — a technique where the model provider or system reuses previously processed portions of a prompt (like a long system prompt) to avoid reprocessing them on every request.
- Model routing — automatically sending different requests to different models based on complexity or requirements, so easy queries use a cheaper model and hard ones use a stronger one.
- Cost-quality trade-off — the explicit tension between using a more capable (and more expensive) model or approach versus a cheaper one that's good enough for the task.
- Capacity planning — forecasting the compute, GPU, and infrastructure resources a system will need ahead of demand, so scaling doesn't become a reactive scramble.
- GPU utilization — how much of the available GPU compute capacity is actually being used productively, versus sitting idle while still being paid for.
- Cost anomaly — an unexpected spike or pattern in spend that doesn't match expected usage, often the first signal of a bug, an abuse pattern, or a misconfigured retry loop.
📌📝 Summary
- Foundations define the machinery of evaluation itself: eval sets, gates, baselines, and the discipline of error analysis.
- Quality and grading terms describe what "good" means and how it gets judged, from groundedness to LLM-as-a-judge.
- Classical metrics (precision, recall, F1, calibration) are the measurement toolkit borrowed from decades of prior ML practice.
- Reliability and latency terms, especially percentiles, describe the experience of your worst-off users, not your average one.
- Experiment design and statistics are what separate a real signal from noise you shouldn't have acted on.
- Dataset quality terms govern whether anything measured against that data can be trusted in the first place.
- RAG and agent vocabularies extend evaluation beyond single-turn text into retrieval correctness and multi-step action.
- Safety, governance, and cost terms round out what "production-grade" actually requires beyond raw quality.
Keep this list close — half of debugging a bad eval conversation is making sure everyone in the room means the same thing by the same word. 🚀
Comments
Post a Comment