Skip to main content

Decoding Strategies in LLMs: Greedy, Beam Search, Top-K & Top-P Explained

Calculating read time…

Decoding strategy is the algorithm that turns an LLM's probability distribution over the next token into the single token that actually gets written — it is the last, decisive step in every single word a model like ChatGPT, Claude, or Gemini produces, and it happens billions of times a day across production systems most people have never heard of. 🧠

This matters because the same underlying model can feel completely different depending on this one setting. A support chatbot that hallucinates policy details, a code assistant that keeps repeating itself, a translation engine that drops half a sentence, a data-extraction pipeline that returns a different JSON shape on every run — these are rarely "the model is bad." More often, they are decoding-strategy problems: the wrong algorithm, or the wrong knob, chosen for the job. Enterprises that treat decoding configuration as a first-class engineering decision (not a leftover default) ship noticeably more reliable AI products. ⚙️

Diagram showing the LLM generation pipeline from context to tokenizer to model to logits to softmax to decoding strategy to next token

🔀 Quick Comparison

Strategy Deterministic? How it picks Best for Watch out for
Greedy Yes Always the single highest-probability token Structured extraction, classification, tool-call arguments Repetition loops, generic phrasing
Beam Search Yes Tracks B best sequences, keeps the top overall Machine translation, summarization with a "correct" target Expensive, biased toward short/safe output
Top-K Sampling No Random pick from the K most likely tokens Open-source model defaults, chat with bounded variety Fixed K can be too wide or too narrow depending on context
Top-P (Nucleus) No Random pick from the smallest set covering probability P Chatbots, creative writing, most default API configs Still needs a temperature/P pairing that's actually tested

1️⃣ How LLMs Actually Choose the Next Word

Here's the part beginners usually skip past, and it's the part that makes everything else click: an LLM does not "know" the next word — it scores every possible next word and hands that scorecard to a decoding strategy.

Walk through it with a sentence any RAG engineer would recognize: "Our retriever returned the most...":

  1. Tokenization — the text is chopped into sub-word pieces using an algorithm called Byte Pair Encoding (BPE). "Our retriever returned the most" doesn't stay as five whole words; a BPE-style tokenizer would typically split it into fragments such as Our, Ġretriever, Ġreturned, Ġthe, Ġmost (the "Ġ" is just how many BPE tokenizers mark a leading space).
  2. Vocabulary lookup — every token maps to an integer ID from the model's fixed vocabulary. The model never sees words; it sees numbers.
  3. Forward pass — those integers flow through the transformer's layers and come out the other side as logits: one raw, unbounded score for every single token in the vocabulary (tens of thousands of them).
  4. Softmax — logits alone are meaningless (they can be negative, huge, tiny). The softmax function squashes them into a proper probability distribution that sums to 1, so the model can meaningfully compare how likely "relevant" is against "similar" as the next word.
Illustrative next-token candidates after "Our retriever returned the most":
  relevant     18.2%
  similar       9.4%
  recent        6.1%
  accurate      5.3%
  useful        4.0%
  (long tail across remaining vocabulary)

(These figures are illustrative, chosen to demonstrate the shape of a realistic probability distribution — not output pulled from a specific model run.)

💬 In Plain English: Think of the model as a librarian who, when asked "what comes next?", doesn't hand you one book — it hands you a ranked stack of every possible book with a probability sticky-note on each. Decoding strategy is the rule you give the librarian for which book to actually pull off the stack. That's it. Everything below is just different versions of that rule.

🎯 Use this when: onboarding a new engineer onto a GenAI team — this is the mental model that makes every downstream API parameter (temperature, top_p, top_k) stop feeling like magic numbers.

2️⃣ Greedy Decoding

Greedy decoding always takes the single highest-probability token at every step, appends it, and repeats until it hits an end-of-sequence token or a length limit. No randomness, no lookahead, no second-guessing.

✅ Real-world example — structured data extraction: A very common production pattern (widely documented across OpenAI's own developer community and enterprise integration guides) is to set an LLM's temperature parameter to 0 whenever the task is classification, tool-call argument generation, or pulling structured fields out of unstructured text. A temperature of 0 collapses the sampling distribution so the model effectively behaves like greedy decoding — the same input should keep producing the same output shape, which is exactly what a downstream system parsing that output into a database record needs.
💡 Key warning: "Deterministic" is not an absolute guarantee. Provider documentation for major LLM APIs explicitly notes that even at temperature 0, floating-point non-determinism, batching effects, and load-balancing across GPUs mean you may still see small variation between calls. If your pipeline needs byte-for-byte reproducibility (e.g. audit logs, regression tests), you still need output validation — not just a temperature setting — as a safety net.

The failure mode to know: greedy decoding can walk into a repetition loop, where "the the the..." or a repeated phrase becomes locally the highest-probability choice at every step and the model gets stuck. This is one of the most common "why does my chatbot keep repeating itself" tickets in production LLM support channels.

🎯 Use this when: the output must be predictable and machine-parseable — JSON extraction, SQL generation, classification labels, tool-call arguments — not when you want varied, natural-sounding prose.

3️⃣ Beam Search

Greedy decoding is short-sighted: the best word right now isn't always part of the best full sentence. Beam search fixes this by keeping several candidate sequences alive at once (the "beam," with width B), expanding every one of them at each step, scoring the results, and keeping only the top B overall — repeating until every surviving beam reaches a stop condition.

✅ Real-world example — Google's production translation system: Google's own research paper on its Neural Machine Translation system (the architecture that powered Google Translate) describes using beam search with a length-normalization procedure and a coverage penalty during inference, specifically to stop the decoder from producing sentences that skip over words in the source text. This is beam search doing exactly what it's good at: picking a globally coherent sequence rather than a locally greedy one, in a system translating billions of sentences a day.

Here's the arithmetic beginners find clarifying. Say the context is "The nightly ingestion job finished" and, with beam width B = 2, the top two first-token candidates are successfully (0.45) and processing (0.30). Expanding both by one more token and multiplying probabilities along each path:

"successfully and"  →  0.45 × 0.50 = 0.225   ← highest, this beam wins
"successfully with" →  0.45 × 0.30 = 0.135
"processing the"     →  0.30 × 0.55 = 0.165
"processing 80"       →  0.30 × 0.25 = 0.075

Note that a greedy decoder would have locked in successfully and never even checked whether processing the might have beaten successfully with down the line. Beam search keeps that door open by carrying multiple hypotheses in parallel — which is exactly the property that matters for a well-formed sentence, not just a locally likely next word.

💡 Key warning: Beam search is B times more expensive to run than greedy decoding, since it scores B parallel sequences at every step. It also tends to favor short, "safe" outputs unless you add a length-normalization penalty — academic literature on production NMT systems generally reports beam widths in roughly the 4–12 range as the practical sweet spot, since very large beams have been shown to actually hurt output quality despite costing more compute.

🎯 Use this when: there's a genuinely "more correct" target sequence to find — translation, summarization with a reference answer, speech-to-text — not for open-ended chat, where beam search tends to produce bland, hedge-everything text.

4️⃣ Top-K Sampling

Greedy and beam search are both deterministic — same input, same output, every time. Top-K and Top-P flip that: they introduce controlled randomness, which is what makes chat responses feel varied instead of robotic.

Top-K sampling keeps only the K highest-probability tokens, throws away the rest, renormalizes the remaining probabilities with softmax so they sum to 1 again, and then randomly samples one token from that shortened list.

✅ Real-world example — open-model defaults: Top-K is one of the standard exposed generation parameters across major open-weight model families and cloud AI platforms — Google's Gemini API, for instance, exposes a topK field, and open-source inference stacks commonly ship with a default in the 40–50 range. It's popular precisely because it's simple to reason about: "never consider more than K candidates" is an easy constraint to explain to a non-technical stakeholder.
💡 Key warning: A fixed K doesn't adapt to how "peaked" or "flat" the distribution is at a given step. If the model is 95% certain the next token is a closing quotation mark, a K of 40 still forces it to consider 39 near-irrelevant alternatives — wasting sampling budget on noise. This exact shortcoming is what Top-P was designed to fix.

🎯 Use this when: you want simple, bounded randomness and you're working with a model/platform where Top-K is the primary exposed knob (e.g. Gemini, many self-hosted open-weight deployments) — pair it with a moderate temperature rather than tuning both aggressively.

5️⃣ Top-P (Nucleus) Sampling

Top-P sampling keeps adding the next most probable token to a candidate set until the cumulative probability crosses a threshold P — that variable-sized set is called the nucleus. The set is renormalized, and a token is randomly sampled from it. Unlike Top-K, the number of candidates considered shrinks or grows automatically depending on how confident the model is at that step.

Take a different scenario: an internal support-ticket router deciding which team a ticket goes to, with candidates billing (0.30), engineering (0.25), sales (0.20), legal (0.15), and a threshold P = 0.7. Adding billing alone gets to 0.30 (still under 0.7), adding engineering gets to 0.55 (still under), adding sales pushes it to 0.75 (over the threshold) — so those three form the nucleus. Renormalized, they become roughly billing 40%, engineering 33%, sales 27%, and one is sampled at random — legal never even enters the running for this step.

✅ Real-world example — the industry-standard chatbot default: Top-P (nucleus sampling) is the default sampling strategy exposed by essentially every major commercial LLM API — OpenAI's Chat Completions API and Anthropic's Messages API both expose a top_p parameter alongside temperature, and both providers' own guidance advises adjusting one or the other, not both aggressively at once, because they interact in ways that are hard to predict together. This is the parameter combination behind most of the "natural-sounding but not incoherent" behavior you see in ChatGPT- and Claude-style consumer chat products.
💡 Key warning: Top-P is not immune to the "false precision" trap. If one token already dominates the distribution at 95% probability, a P of 0.9 can still collapse the nucleus down to that single token — so a "high" P value doesn't automatically guarantee diverse output. Test with real prompts from your product, not just the parameter's theoretical range.

🎯 Use this when: building conversational products, creative-writing tools, or anything where "natural-sounding and non-repetitive" matters more than strict reproducibility — this is the safest general-purpose default across the major commercial APIs.

6️⃣ Temperature: The Dial Behind All Four

The source article this post is built on focuses on the four selection algorithms above, but no production discussion of decoding is complete without temperature, because it changes the shape of the probability distribution those algorithms then select from. Temperature divides each logit by a value T before the softmax step: low T sharpens the distribution (the model becomes more confident/predictable), high T flattens it (more tokens become plausible, output gets more surprising).

✅ Real-world example — reasoning models lock it down: As of the most recent provider documentation, OpenAI's reasoning-model line (its "o-series" and successors) fixes temperature, top_p, and sample count to their defaults and rejects non-default values with an error — the rationale given is that the model's internal chain-of-thought scaffolding needs consistent sampling behavior to work reliably. Separately, Anthropic's own Claude Platform documentation notes that temperature, top_p, and top_k are not accepted at all on its newest model generations, with guidance to control behavior through prompting instead. This is a genuinely current architectural shift worth knowing: the industry trend for the most capable "reasoning" models is moving away from exposing raw sampling knobs, not toward exposing more of them.
💡 Key warning: Don't assume a temperature/top_p contract from one model generation carries forward to the next. Teams have shipped production incidents by hardcoding a temperature: 0.7 payload that a newer reasoning-tier model silently rejects or ignores. Always check the specific model's current parameter contract before deploying — this is exactly the kind of fast-moving detail that goes stale in a training corpus.

7️⃣ How Production APIs Actually Combine These

Platform Primary knobs exposed Notable behavior
OpenAI Chat Completions temperature (0–2), top_p Reasoning-tier models fix both to defaults; classic models default to temperature=1, top_p=1
Anthropic Messages API temperature (0–1), top_p, top_k Newest model generations reject non-default sampling params entirely; guidance shifts to prompting
Google Gemini API temperature, topP, topK, candidateCount One of the few major APIs still exposing Top-K directly alongside Top-P
Code-completion tools (Copilot-style) temperature, usually low or fixed Practitioner reports describe these tools defaulting near temperature 0 for inline completions, prioritizing predictability over creativity

💬 In Plain English: No serious production system uses "just Top-K" or "just Top-P" in isolation — they're layered. A typical real request pipeline filters by Top-K or Top-P and applies a temperature scale and often adds repetition/frequency penalties on top. The four strategies in this post are the building blocks; production configs are usually a specific, tested combination of two or three of them.

8️⃣ Enterprise Rollout at Scale

Decoding strategy stops being a "just pick a number" decision the moment more than one team is calling the same model. What scales well in practice:

  1. Named decoding profiles, not raw parameters. Instead of every team choosing its own temperature/top_p pair, define a small set of approved profiles in a shared config — e.g. extraction (temperature 0, greedy-equivalent), chat-default (temperature 0.7, top_p 0.9), creative (temperature 1.0, top_p 0.95) — and require every service to reference a profile by name.
  2. Governance and ownership. One team (often Platform/ML Infra) owns the profile definitions; product teams request changes through review rather than editing values inline, the same way infrastructure teams govern shared Terraform modules.
  3. CI enforcement for determinism-sensitive paths. For extraction/classification services, add automated tests that call the same prompt multiple times and assert output-schema stability — this catches a silently-changed default (e.g. a provider changing what "temperature 0" means under the hood) before it reaches production.
  4. Monitoring metrics, not just latency and cost. Track repetition rate, output-length variance, and schema-validation failure rate per decoding profile over time. A spike in repetition rate on a "chat-default" profile is often the first signal that a model version upgrade quietly changed sampling behavior.

9️⃣ Common Mistakes

  1. Tuning temperature and top_p aggressively at the same time. Because they both reshape the same probability distribution, changing both makes it very hard to isolate which one caused a regression in output quality. Provider guidance consistently recommends holding one at its default while tuning the other.
  2. Assuming "temperature 0" means bit-for-bit reproducibility. As covered above, provider documentation itself flags that near-zero temperature reduces but doesn't guarantee determinism. Teams that build compliance or audit workflows on this assumption without an output-validation layer are exposed to subtle drift.
  3. Using beam search for open-ended chat. Beam search optimizes for the single "best" sequence, which in conversational contexts tends to produce safe, generic, repetitive-feeling text — it's built for tasks with a defined correct target (like translation), not for open creative dialogue.
  4. Hardcoding sampling parameters across model versions. As shown with the newest reasoning-tier models from multiple providers rejecting or ignoring temperature/top_p/top_k entirely, a payload that worked on last quarter's model can silently break — or silently get ignored — on this quarter's model.
  5. Treating Top-K's fixed cutoff as always "safe." A static K that works well for a flat, ambiguous distribution can be needlessly restrictive for a peaked, high-confidence one (and vice versa) — this is precisely the gap Top-P was designed to close, and it's worth testing both on your actual traffic rather than defaulting by habit.

❓ FAQ

Q: Is greedy decoding the same as setting temperature to 0?

A: Functionally, yes for most practical purposes — a temperature of 0 collapses the probability distribution so the top token is chosen with near-certainty, which behaves like greedy decoding. Provider documentation notes it isn't a mathematically absolute guarantee of determinism due to floating-point and infrastructure effects, but for everyday engineering purposes, treat them as equivalent.

Q: Should I use Top-K or Top-P?

A: Top-P is the more widely defaulted-to choice across major commercial chat APIs because it adapts to how confident the model is at each step, rather than using a fixed candidate count. Top-K is simpler to reason about and remains common on platforms like Gemini and many self-hosted deployments. Many production stacks actually use both together with a temperature setting layered on top.

Q: Why did my Claude or GPT API call suddenly reject my temperature parameter?

A: The newest reasoning-oriented models from multiple major providers have moved to fixing sampling parameters at their defaults (or rejecting non-default values outright) to keep their internal reasoning process consistent. If you upgraded to a newer model and see a 400-style error mentioning temperature, top_p, or top_k, check that specific model's current parameter contract in the provider's docs rather than assuming your old payload still applies.

Q: Is beam search still used in modern LLM chat products?

A: Rarely for open-ended chat — it's expensive and tends to produce bland output. It remains genuinely valuable for tasks with a defined "correct" target sequence, like machine translation, where Google's own published research on its production translation system describes using beam search with length normalization and a coverage penalty.

Q: What's the single most common decoding mistake in production RAG or agent systems?

A: Using a "chatty" decoding profile (moderate-to-high temperature, top_p around 0.9) for a step that actually needs deterministic output — like generating a tool call, a SQL query, or a JSON payload that a downstream parser depends on. Extraction and structured-output steps generally want a near-zero-temperature, greedy-equivalent profile; only the natural-language-facing steps need sampling-based creativity.

🔗 References & Further Reading

  • Wu et al., "Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation" — arXiv:1609.08144
  • OpenAI Platform Documentation — Reasoning models guide (temperature/top_p/n fixed at defaults) — platform.openai.com
  • Anthropic Claude Platform Documentation — Messages API parameter reference — platform.claude.com
  • Hugging Face transformers library documentation — text generation strategies — huggingface.co

All product names (OpenAI, ChatGPT, Anthropic, Claude, Google, Gemini, Google Translate) are trademarks of their respective owners and are referenced here only for factual, educational comparison. All technical claims above are synthesized and explained in original wording based on publicly available documentation and research — no text has been reproduced verbatim from any source.

📝 Summary

  • LLMs generate text by scoring every possible next token and handing that scored list to a decoding strategy — the model predicts probabilities; decoding strategy decides.
  • Greedy decoding always takes the top token — fast, deterministic, ideal for extraction and tool calls, but prone to repetition loops.
  • Beam search tracks multiple candidate sequences to avoid greedy's short-sightedness — powers real translation systems like Google's, at a real compute cost.
  • Top-K sampling randomly samples from a fixed-size pool of top candidates — simple, common on Gemini and open-weight stacks, but doesn't adapt to model confidence.
  • Top-P (nucleus) sampling randomly samples from a dynamically-sized pool based on cumulative probability — the de facto default across most commercial chat APIs.
  • Temperature reshapes the distribution before any of the above run — and the newest reasoning-tier models are increasingly locking it down entirely.
  • At enterprise scale, decoding strategy should be a governed, named configuration — not a per-team, per-call guess.

That's the full picture from "what is a logit" all the way to "why did my API call just 400." If you're building a RAG pipeline, pair this with how you configure retrieval-side determinism too — happy building! 🚀

Comments