Skip to main content

LLM Evaluation Glossary: 180+ Terms Explained

Calculating read time…

This is the shared vocabulary every LLM reliability and evaluation engineer needs before they can read a dashboard, write a test plan, or debug a production incident without guessing what a term means. None of these words are decoration — each one maps to a specific decision someone made about how a system is measured, gated, or trusted. 📘

📌1. Core Evaluation Foundations

  • Task success rate — the share of test cases (or real interactions) where the system actually achieved the user's goal, not just produced a plausible-looking response.
  • Evaluation dataset / eval set — the fixed collection of inputs, and usually expected outputs, that a model or system is scored against; its quality sets the ceiling on how much you can trust any score.
  • Golden dataset / golden set — a small, carefully curated, high-trust subset of the eval set, usually reviewed by experts, used as the stable long-term benchmark for comparing model versions over time.
  • Ground truth / reference answer — the accepted correct output for a test case, against which a model's actual output is compared; for open-ended tasks this may be a rubric rather than one exact string.
  • Test case — a single input (plus context) paired with its expected behavior, ground truth, or scoring criteria — the atomic unit an eval set is built from.
  • Rubric — a written set of criteria describing what makes an answer good, bad, or partially correct, used to keep human or LLM-judge scoring consistent across many raters and runs.
  • Metric — a specific, defined measurement computed from outputs (accuracy, latency, groundedness score) that turns raw behavior into a comparable number.
  • Quality threshold — the minimum acceptable value for a metric, below which an output or a release is considered a failure rather than a pass.
  • Quality gate / release gate — the checkpoint in a pipeline where a model or prompt change must clear defined thresholds before it's allowed to progress toward production traffic.
  • Baseline — the current production system's performance, used as the reference point a new candidate must match or beat before it's allowed to replace it.
  • Candidate system — the new model, prompt, or configuration being evaluated for possible promotion, compared directly against the baseline.
  • Regression — a drop in performance on something that used to work, introduced by a change that was intended to improve something else.
  • Error analysis — the manual, structured process of reading through failures to find patterns, root causes, and categories of mistake, rather than stopping at "the score went down."
  • Data slice — a defined subset of the eval set or production traffic (a language, a task type, a user segment) that's scored separately so an aggregate number can't hide a localized problem.
  • Production monitoring — the ongoing collection of metrics and signals from a live system, used to catch problems that only appear after release, not just at launch.
  • Observability — the broader capability to ask new questions about a running system's behavior after the fact, using logs, metrics, and traces, without having predicted every question in advance.
  • Trace — the full recorded path of a single request through a system — every step, call, and decision it took — used to reconstruct exactly what happened for one specific interaction.
  • Span — one timed, named segment within a trace (a single retrieval call, a single model call), the building block traces are made of.

📌2. Output Quality and Grading

  • Correctness — whether the factual content or logic of an answer is actually right, independent of how well it's written or presented.
  • Relevance — whether the answer actually addresses what was asked, rather than being accurate about something adjacent or off-topic.
  • Completeness — whether the answer covers everything the question required, rather than a technically correct but partial response.
  • Helpfulness — a holistic judgment of whether the answer actually moved the user closer to solving their problem, combining correctness, relevance, and tone.
  • Groundedness — whether the claims in an answer are actually supported by the source material or context the model was given, rather than invented.
  • Faithfulness — closely related to groundedness: whether the answer accurately represents what the source material says, without distorting or overstating it.
  • Hallucination — content the model states with confidence that is false, unsupported, or fabricated — the single most reputationally expensive failure mode in production LLM systems.
  • Citation correctness — whether a cited source actually says what the model claims it says, checked citation by citation, not just that a citation exists.
  • Citation coverage — the share of factual claims in an answer that are backed by a citation at all, regardless of whether each citation is individually correct.
  • Format adherence — whether the output follows the requested structure (length, tone, sections, style) that was specified in the prompt or system instructions.
  • Structured output — output constrained to a specific machine-readable shape, such as JSON or a table, rather than free-form prose.
  • Schema validity — whether a structured output actually parses and conforms to its defined schema — the field names, types, and required keys it was supposed to have.
  • Deterministic validator — a non-LLM check (a regex, a parser, a rule) that verifies a specific, objective property of an output with no ambiguity or judgment involved.
  • Rule-based evaluation — scoring built from explicit, hand-written rules rather than a model's judgment — fast, cheap, and reliable for narrow, well-defined checks.
  • LLM-as-a-judge — using a separate model call to score or compare outputs against a rubric, scaling evaluation far beyond what human review can cover, at the cost of inheriting the judge's own biases.
  • Judge prompt — the instructions given to the judge model describing exactly how to score or compare outputs, which functions as the rubric translated into an executable form.
  • Judge calibration — the process of checking and adjusting a judge model's scores against human labels, so its verdicts can be trusted to track real quality rather than the judge's own quirks.
  • Human evaluation — scoring or comparison performed by trained people rather than an automated metric or judge, used where nuance and judgment matter most.
  • Pairwise comparison — presenting a rater (human or LLM) with two outputs side by side and asking which is better, rather than scoring each output in isolation.
  • Preference evaluation — evaluation built around which of two or more outputs people prefer, rather than an absolute quality score — often more reliable than absolute scoring for subjective tasks.
  • Inter-annotator agreement — the degree to which different human raters give the same label to the same example, used to check whether a rubric is clear and consistently applied.

📌3. Classical Metrics and Measurement

  • Accuracy — the share of predictions that were correct out of all predictions made; simple, but misleading on its own when classes or outcomes are imbalanced.
  • Precision — out of everything the system flagged as positive, the share that was actually positive — high precision means few false alarms.
  • Recall — out of everything that was actually positive, the share the system successfully caught — high recall means few misses.
  • F1 score — the harmonic mean of precision and recall, giving one combined number when you care about both and neither alone tells the full story.
  • Exact match — a strict scoring rule where the output must match the reference answer character-for-character (or after light normalization) to count as correct.
  • Partial match — a scoring rule that gives partial credit when an output is close to correct but not identical, more forgiving than exact match for open-ended tasks.
  • False positive — the system flagged something as true, unsafe, or positive when it actually wasn't — the source of unnecessary alarms or wrongful blocks.
  • False negative — the system missed something that actually was true, unsafe, or positive — often the more dangerous error type in safety-critical checks.
  • True positive — the system correctly flagged something that actually was positive.
  • True negative — the system correctly did not flag something that actually was negative.
  • Confusion matrix — a table laying out true/false positives and negatives together, showing exactly what kinds of mistakes a classifier makes, not just how often it's wrong.
  • Classification threshold — the cutoff score above which a model's continuous output is treated as a positive prediction; moving it trades precision against recall.
  • Calibration — whether a model's confidence scores actually match its real accuracy — a well-calibrated model that says "80% confident" is right about 80% of the time.
  • Confidence score — a number a model or classifier attaches to its own prediction, meant to indicate how likely that prediction is to be correct.
  • Pass@k — the probability that at least one of k independently sampled outputs for a task is correct, commonly used to evaluate code generation and other tasks where you can generate multiple attempts.

📌4. Reliability, Performance, and Latency

  • Latency — the time between a request being sent and a response being received; the metric users feel most directly.
  • Throughput — the number of requests (or tokens) a system can process per unit of time at a given level of resources.
  • Availability — the share of time a service is actually up and successfully serving requests, usually expressed as a percentage over a period.
  • Error rate — the share of requests that fail outright, as opposed to succeeding with poor quality — a distinct problem from low-quality-but-successful responses.
  • Timeout rate — the share of requests that failed specifically because they took longer than the system was willing to wait.
  • Retry rate — the share of requests that had to be attempted more than once before succeeding, a useful early signal of instability even before outright failures appear.
  • Fallback rate — the share of requests served by a backup path (a smaller model, a cached answer, a simpler rule) instead of the primary system, usually because the primary was unavailable or too slow.
  • Queue time — the time a request spends waiting for a resource to become available before processing even starts, separate from actual processing time.
  • Concurrency — the number of requests being processed by a system at the same moment, which drives resource contention and, past a point, latency.
  • Rate limit — a cap on how many requests a client or the system as a whole is allowed to make in a given time window, used to protect capacity and control cost.
  • Tail latency — the latency experienced by the slowest fraction of requests, which matters disproportionately because it's what your most frustrated users actually feel.
  • p50 / median — the latency value below which half of all requests complete; the "typical" experience, but blind to how bad the worst requests are.
  • p90 — the latency value below which 90% of requests complete; the slowest 10% take longer than this.
  • p95 — the latency value below which 95% of requests complete; the slowest 5% take longer than this.
  • p99 — the latency value below which 99% of requests complete; the slowest 1% take longer than this, and at real scale that 1% can still be thousands of unhappy users.
  • Time to First Token (TTFT) — how long a user waits after sending a request before the model starts streaming back any output at all — the metric that drives perceived responsiveness.
  • Time to Last Token — the total time from request to the very last token of the response, i.e. full completion time.
  • Tokens per second — the rate at which a model generates output once it has started, which determines how fast a long response finishes streaming.
  • Service Level Indicator (SLI) — the actual measured value of a metric that matters to users, such as real latency or real error rate — the "what we measured."
  • Service Level Objective (SLO) — the internal target you set for an SLI (for example, p95 latency under a set number of seconds), used to drive engineering priorities.
  • Service Level Agreement (SLA) — a formal, often contractual commitment to a service level, typically with consequences if it's not met — stricter and more external-facing than an SLO.
  • Error budget — the amount of unreliability you're allowed to spend against an SLO before it's breached, used to balance shipping speed against stability deliberately rather than by accident.

✅ Percentile latency, made concrete: if p95 latency is 8 seconds, that means 95% of requests complete in 8 seconds or less, while the slowest 5% take longer than 8 seconds. p50 tells you what a typical user feels; p95 and p99 tell you how bad it gets for the users having the worst experience — and at production scale, that slowest slice can still be a very large number of real, frustrated people.

📌5. Experiment Design and Statistical Confidence

  • Offline evaluation — running a candidate against a fixed, versioned test set before any real user is exposed to it; fast and repeatable, but limited to what the test set anticipated.
  • Online evaluation — evaluating a candidate against real, live traffic once it's actually serving some portion of real users.
  • Continuous evaluation — evaluation that runs on an ongoing basis after full rollout, not just at launch, to catch drift and decay that appear later.
  • A/B testing — randomly splitting traffic between two versions and comparing their outcomes statistically, to attribute any difference to the change itself rather than chance.
  • Canary rollout — gradually increasing the share of traffic sent to a new version while the old version keeps serving the rest, so a regression affects a small, controlled slice of users first.
  • Shadow testing — sending a copy of real traffic to a candidate system without showing its responses to users, so its real-world behavior can be checked with zero user-facing risk.
  • Holdout set — a portion of data deliberately kept out of training and tuning, reserved purely for evaluating how the system generalizes to unseen cases.
  • Control group — in an experiment, the group that continues to receive the existing (baseline) system, used as the point of comparison.
  • Treatment group — in an experiment, the group that receives the new candidate system being tested.
  • Sample size — the number of observations a test or comparison is based on; too small a sample size means real effects can be invisible and noise can look like a pattern.
  • Variance — how spread out a metric's values are; high variance means a single measurement is a less trustworthy estimate of the true value.
  • Confidence interval — a range around a measured value that likely contains the true value, communicating how much uncertainty a single point estimate is hiding.
  • Statistical significance — a measured difference is unlikely to be due to random chance alone, given the sample size and variance observed.
  • Effect size — how large a difference actually is in practical terms, independent of whether it's statistically significant — a tiny but statistically significant difference may not matter to anyone.
  • Statistical power — the probability that a test will correctly detect a real effect if one actually exists, given the sample size and expected effect size.
  • Bootstrap confidence interval — a confidence interval built by repeatedly resampling the observed data itself, useful when a metric's true distribution is unknown or doesn't fit a standard formula.
  • Practical significance — whether a statistically real difference is actually large enough to matter for users or the business, distinct from statistical significance alone.
  • Rollback — reverting a release back to the previous known-good version, typically triggered when a monitored metric crosses a defined failure threshold.

📌6. Evaluation Dataset Quality

  • Representativeness — how closely an eval set's mix of examples matches the real distribution of what the system actually encounters in production.
  • Coverage — how completely an eval set spans the range of task types, languages, and scenarios the system needs to handle, rather than clustering around a few easy cases.
  • Stratified sampling — deliberately sampling a test set so specific subgroups (languages, task types, user segments) are proportionally represented, rather than left to chance.
  • Edge case — an unusual, boundary, or rarely occurring input that still needs to be handled correctly, and is disproportionately likely to expose a bug.
  • Adversarial test case — an input deliberately crafted to break, mislead, or exploit a system, used to test robustness rather than typical behavior.
  • Red teaming — a structured, ongoing effort by people specifically trying to find safety and robustness failures in a system before real attackers do.
  • Synthetic data — artificially generated examples (often by another model) used to expand or diversify an eval or training set beyond what was collected naturally.
  • Train split — the portion of data used to actually teach or fine-tune a model.
  • Validation split — the portion of data used during development to tune choices and catch problems before final evaluation, kept separate from training.
  • Test split — the portion of data reserved for final, unbiased evaluation, untouched during training or tuning decisions.
  • Dataset versioning — tracking exactly which version of a dataset was used for a given training run or evaluation, so results can be reproduced or audited later.
  • Data drift — a gradual change in the statistical properties of incoming data over time, compared to what a model was built or tested against.
  • Distribution shift — a broader, sometimes sudden mismatch between the data a system now sees and the data it was trained or evaluated on, which can silently degrade performance.
  • Data leakage — information from outside the intended training or evaluation scope (often from the answer itself, or from the test set) accidentally influencing a result, making performance look better than it really is.
  • Benchmark contamination — a specific form of leakage where benchmark test data has ended up inside a model's training data, inflating its reported score on that benchmark.
  • Near-duplicate detection — identifying examples that are not identical but close enough in meaning or wording to count as leakage, typically using embedding similarity rather than exact-match checks.
  • Annotation guideline — the written instructions given to human labelers describing exactly how to label or score examples consistently.
  • Label quality — how accurate, consistent, and well-justified the human-assigned labels in a dataset actually are — a training and evaluation set is only as trustworthy as its labels.

📌7. RAG Evaluation Vocabulary

  • Retrieval — the step where a system searches an external knowledge source for information relevant to a query, before generating an answer.
  • Retrieval precision — out of the documents or chunks retrieved, the share that were actually relevant to the query.
  • Retrieval recall — out of all the truly relevant documents that existed, the share the retrieval step actually found.
  • Context precision — how much of the retrieved context that gets passed to the model is actually useful, rather than irrelevant padding.
  • Context recall — whether the retrieved context contains everything needed to fully and correctly answer the question.
  • Answer relevance — whether the final generated answer actually addresses the user's question, independent of whether the retrieval step did its job well.
  • Answer faithfulness — whether the final answer accurately reflects what the retrieved context actually says, without adding unsupported claims.
  • Context utilization — how much of the retrieved, relevant context the model actually used in constructing its answer, as opposed to ignoring it.
  • Chunking strategy — the method used to split source documents into smaller pieces for indexing and retrieval, which strongly affects both retrieval quality and answer faithfulness.
  • Embeddings — numeric vector representations of text that capture meaning, allowing similarity between pieces of text to be computed mathematically.
  • Reranking — a second pass that reorders an initial set of retrieved candidates by relevance, usually with a more precise (and more expensive) model than the first retrieval pass.
  • Top-k retrieval — returning the k most relevant results for a query, where k is a tunable number balancing completeness against noise and cost.
  • Retrieval failure — a case where the retrieval step fails to surface any genuinely relevant document, which downstream generation cannot recover from no matter how good the model is.
  • Citation grounding — verifying that every cited source in an answer actually supports the specific claim it's attached to, not just that a citation is present.
  • Abstention — the system correctly declining to answer when it doesn't have sufficient grounded information, rather than guessing.
  • No-answer quality — how well a system handles the case where the honest answer is "I don't know" or "this isn't covered" — a distinct skill from answering well when it does have the information.
  • Knowledge freshness — how up to date the information in the retrieval index actually is relative to the real world, which directly caps how current a RAG system's answers can be.

📌8. AI Agent Evaluation Vocabulary

  • AI agent — a system that uses an LLM to plan and take a sequence of actions (often calling tools or other systems) toward a goal, rather than just producing a single text response.
  • End-to-end task success — whether the agent actually achieved the overall goal by the end of its run, the single most important agent metric, distinct from how good any individual step looked.
  • Tool selection accuracy — whether the agent chose the correct tool for a given step, out of the tools available to it.
  • Function calling — the mechanism by which a model outputs a structured request to invoke an external function or API, rather than only generating text.
  • Tool-call validity — whether a function call the model generated is well-formed and executable — correct name, correct structure — regardless of whether it was the right call to make.
  • Argument accuracy — whether the specific values the agent passed into a tool call were correct, not just whether the call's structure was valid.
  • Tool execution success rate — the share of tool calls that actually completed successfully once executed, as opposed to failing at runtime.
  • Tool error handling — how well an agent responds when a tool call fails — whether it retries sensibly, informs the user, or tries an alternative, rather than pretending the call succeeded.
  • Trajectory — the full sequence of steps, decisions, and tool calls an agent took to complete (or fail) a task, evaluated as a whole rather than one action at a time.
  • Step-level evaluation — scoring each individual action in a trajectory on its own merits, used to pinpoint exactly where a multi-step task went wrong.
  • Plan quality — whether the sequence of steps an agent chose was logical and efficient for the goal, independent of whether it happened to reach the right outcome anyway.
  • Multi-turn evaluation — scoring a full conversation or session as a whole, since individually reasonable replies can still add up to a session that never actually resolved the user's need.
  • State management — how correctly a system tracks and updates the relevant facts about an ongoing task or conversation as it progresses.
  • Memory correctness — whether information a system is supposed to remember across turns or sessions stays accurate and isn't corrupted, dropped, or confused with something else.
  • Context management — how well a system selects, prioritizes, and trims what information stays available to the model as a conversation or task grows longer than its context window.
  • Retry — automatically attempting a failed action again, typically after a brief delay, before treating it as a hard failure.
  • Backoff — increasing the delay between successive retry attempts, so repeated failures don't hammer an already-struggling dependency.
  • Fallback — switching to an alternative method, tool, or simpler response when the primary approach fails or isn't available.
  • Circuit breaker — a mechanism that stops sending requests to a dependency after it starts failing repeatedly, to prevent cascading failure while it's unhealthy.
  • Escalation — handing a task off to a human (or a higher-capability system) when an agent can't complete it safely or successfully on its own.
  • Human-in-the-loop — a design where a person reviews, approves, or can intervene in an agent's actions, rather than letting it act fully autonomously.
  • Autonomy rate — the share of tasks an agent completes successfully without needing human intervention or escalation.
  • Recovery rate — the share of times an agent successfully gets back on track after hitting an error or an unexpected state, rather than failing outright.
  • Safe tool use — an agent's ability to use powerful or risky tools (payments, file deletion, external messaging) only within intended, safe boundaries.
  • Idempotency — the property that performing the same action multiple times has the same effect as performing it once, which is what makes safe retries of tool calls possible.

📌9. Safety, Security, and Governance

  • Policy compliance — whether a system's outputs and actions stay within the organization's defined rules for acceptable content and behavior.
  • Safety evaluation — testing specifically aimed at finding harmful, dangerous, or policy-violating outputs, run separately from general quality evaluation.
  • Prompt injection — an attempt to override a system's original instructions by embedding new, malicious instructions inside user input or content the model processes.
  • Jailbreak — a technique designed to manipulate a model into bypassing its safety guidelines and producing content or actions it would normally refuse.
  • Indirect prompt injection — a prompt injection attack delivered through content the model reads as data (a webpage, a document, an email) rather than typed directly by the user.
  • Data exfiltration — a system being manipulated into leaking sensitive information it had access to, to an unauthorized party.
  • PII leakage — a system exposing personally identifiable information in its output when it should have withheld or redacted it.
  • Sensitive-data handling — the set of practices governing how a system detects, protects, and limits exposure of sensitive information throughout its pipeline.
  • Authorization boundary — the defined limit of what actions and data a system or agent is permitted to access, which every safety-relevant check assumes is being enforced.
  • Least privilege — the principle of giving a system or agent only the minimum access it needs to do its job, and nothing more, to limit the damage of any single failure or exploit.
  • Unsafe tool use — an agent invoking a tool in a way that causes real-world harm or risk, whether from a bad decision, a manipulation, or a bug.
  • Correct refusal — the system appropriately declining a genuinely unsafe or out-of-policy request, which is the desired behavior, not a failure.
  • Over-refusal — the system incorrectly declining a legitimate, safe request, mistaking it for something it should refuse — a real quality failure, not a safe default.
  • Under-refusal — the system failing to decline a request it should have refused, letting an unsafe or policy-violating output through.
  • Audit trail — a persistent, reviewable record of what a system did, when, and why, used to reconstruct events after the fact for debugging, compliance, or incident review.
  • Governance — the overall structure of ownership, policy, and approval processes that decides how models and systems are allowed to be built, tested, and released.
  • Incident response — the defined process for detecting, containing, and recovering from a production failure or safety event, including who is responsible for each step.

📌10. Cost and Operational Efficiency

  • Input tokens — the units of text sent into a model as part of a request (the prompt, context, and history), which is one basis models are typically billed on.
  • Output tokens — the units of text a model generates in its response, usually billed at a different rate than input tokens.
  • Token usage — the total volume of input and output tokens consumed, the core driver of both cost and latency for a given workload.
  • Cost per request — the average cost of serving a single request, combining token usage, model choice, and any tool or retrieval costs involved.
  • Cost per successful task — cost measured against actual successful outcomes rather than raw requests, which is what really matters when some requests fail or need retries.
  • Token efficiency — how much useful output or task progress a system achieves per token consumed, a measure of how "wasteful" a prompt or workflow is.
  • Cache hit rate — the share of requests answered from a cached previous result instead of a fresh model call, directly reducing both cost and latency.
  • Prompt caching — a technique where the model provider or system reuses previously processed portions of a prompt (like a long system prompt) to avoid reprocessing them on every request.
  • Model routing — automatically sending different requests to different models based on complexity or requirements, so easy queries use a cheaper model and hard ones use a stronger one.
  • Cost-quality trade-off — the explicit tension between using a more capable (and more expensive) model or approach versus a cheaper one that's good enough for the task.
  • Capacity planning — forecasting the compute, GPU, and infrastructure resources a system will need ahead of demand, so scaling doesn't become a reactive scramble.
  • GPU utilization — how much of the available GPU compute capacity is actually being used productively, versus sitting idle while still being paid for.
  • Cost anomaly — an unexpected spike or pattern in spend that doesn't match expected usage, often the first signal of a bug, an abuse pattern, or a misconfigured retry loop.

📌📝 Summary

  • Foundations define the machinery of evaluation itself: eval sets, gates, baselines, and the discipline of error analysis.
  • Quality and grading terms describe what "good" means and how it gets judged, from groundedness to LLM-as-a-judge.
  • Classical metrics (precision, recall, F1, calibration) are the measurement toolkit borrowed from decades of prior ML practice.
  • Reliability and latency terms, especially percentiles, describe the experience of your worst-off users, not your average one.
  • Experiment design and statistics are what separate a real signal from noise you shouldn't have acted on.
  • Dataset quality terms govern whether anything measured against that data can be trusted in the first place.
  • RAG and agent vocabularies extend evaluation beyond single-turn text into retrieval correctness and multi-step action.
  • Safety, governance, and cost terms round out what "production-grade" actually requires beyond raw quality.

Keep this list close — half of debugging a bad eval conversation is making sure everyone in the room means the same thing by the same word. 🚀

Comments