Skip to main content

Chain-of-Draft Prompting: The Prompt Engineer's Guide to Faster, Cheaper LLM Reasoning

Calculating read time…

Chain-of-Draft (CoD) is a prompting technique that asks a language model to reason toward an answer using terse, five-word-or-fewer intermediate notes instead of full explanatory sentences — and it can match or beat Chain-of-Thought's accuracy while writing as little as roughly a tenth of the reasoning text. For anyone whose job is designing prompts rather than training models, that's a genuinely new lever. 🧠

This matters because reasoning-style prompting has a cost most teams underprice: every "let's think step by step" instruction can turn a cheap API call into a slow, token-heavy one, and in latency-sensitive products — live chat, voice assistants, real-time meeting tools — that verbosity is the difference between a response that feels instant and one that feels broken. A prompt engineer who only knows Chain-of-Thought is missing a technique built specifically to fix this trade-off. ⚡

Diagram comparing Chain-of-Thought's verbose reasoning steps against Chain-of-Draft's terse five-word reasoning steps for the same question

🔀 Quick Comparison

Approach What the model writes Token cost Best for
Standard / Direct Just the final answer, no reasoning Lowest Trivial lookups, simple classification
Chain-of-Thought (CoT) Full, explanatory reasoning sentences Highest Complex multi-step reasoning, when tokens/latency aren't the bottleneck
Chain-of-Draft (CoD) Terse notes, ≤5 words per reasoning step Low — a fraction of CoT's Latency-sensitive or cost-sensitive reasoning tasks

1️⃣ What Chain-of-Draft Actually Is

🔑 Three words to know before we go further:

  • Token — roughly a word or word-fragment. LLM providers charge by the token and also take time to generate each one, so more tokens = more cost and more waiting.
  • Latency — the delay between asking a model something and getting the full reply back. Long reasoning = long latency, which you feel as "it's thinking..." on screen.
  • Few-shot examples — a couple of worked examples you paste into the prompt itself to show the model exactly the style of answer you want, before it sees the real question.

If you've ever used ChatGPT, Claude, or Gemini and watched it "think out loud" before answering a tricky question, you've seen Chain-of-Thought reasoning in action. It's genuinely useful — but every single word of that thinking is a token the model had to generate, and every token takes a small slice of time. Chain-of-Draft is simply a way to keep that same careful thinking while making the model write far less of it down.

Chain-of-Thought prompting works by asking a model to "think step by step," and the model responds with full, human-readable sentences of reasoning before it commits to an answer. That works well, but it's verbose — a model will often write several paragraphs to solve a problem a person could sketch out in a few lines of scratch paper.

Chain-of-Draft, introduced by a research team at Zoom Communications, takes a different cue from human behavior: when people actually work through a problem, they jot down shorthand — partial calculations, single words, quick symbols — not full sentences. CoD prompts a model to do the same: keep reasoning, but hold every intermediate note to roughly five words or under. Critically, nothing in the decoding engine or the surrounding code actually enforces that limit — it's taught purely through a handful of few-shot examples in the prompt that demonstrate the terse style, and the model picks up and generalizes the pattern on its own.

💬 In Plain English: If Chain-of-Thought is asking someone to show their full work on an exam, Chain-of-Draft is asking them to solve it on a sticky note. Same underlying thinking, radically less writing.

🎯 Use this when: you're prompt-engineering a reasoning task where every output token has a real cost — in money, latency, or both.

2️⃣ Why It Was Built: A Real Latency Problem

✅ Real-world example — Zoom's own reasoning latency problem: Chain-of-Draft comes out of Zoom Communications' research team — the same company that has shipped a real-time "AI Companion" meeting assistant since 2023. Multiple independent technology outlets covering the paper's release describe this as the practical motivation behind the work: a reasoning step that takes several seconds to write out in full CoT-style prose is a poor fit for a product that has to summarize or respond inside a live video call. CoD was built to keep the reasoning quality of CoT while cutting the output length — and therefore the response time — dramatically.

It also directly answers a separate, well-documented failure mode of reasoning models: overthinking. Research into "reasoning" models has shown they can generate surprisingly long chains of thought even for trivial problems — burning tokens and time on something a calculator solves instantly. A prompting technique that structurally discourages verbosity, rather than relying on the model to self-regulate, is a direct countermeasure to that.

💡 Key warning: CoD is not a universal replacement for CoT. Some tasks genuinely need extended deliberation, self-correction, or pulling in outside knowledge mid-reasoning — squeezing those into five-word steps can strip out information the model actually needed to reach a correct answer. Treat CoD as the right tool for well-defined, latency-sensitive reasoning tasks, not a blanket upgrade.

📈 Why This Matters Even More Right Now

Zoom's meeting assistant is one example, but the same latency pressure shows up across most of what's actually getting built in AI products today:

  1. Voice AI agents. Real-time voice assistants — the kind that answer a phone call or talk back to you live — have almost no room for a multi-paragraph internal monologue. Every extra second of "thinking" is a second of awkward silence on the call. A terse reasoning style is a much more natural fit for voice than verbose CoT.
  2. Agentic AI workflows. A coding agent or browser agent doesn't call an LLM once — it calls it repeatedly, once per step, sometimes dozens of times to finish one task. If each of those internal calls reasons in full CoT-style prose, the verbosity doesn't just add up, it multiplies across every step in the chain. Trimming each step's reasoning has an outsized effect on the agent's total time-to-finish.
  3. High-volume support copilots. A helpdesk assistant triaging thousands of tickets a day doesn't need a paragraph of reasoning per ticket — it needs a fast, auditable trail of how it reached HIGH/MEDIUM/LOW, at a cost that still makes sense multiplied by that ticket volume.

3️⃣ CoD vs. CoT vs. Standard Prompting — Worked Example

Take a simple word problem: "A warehouse had 20 pallets. It shipped 12 and then received 3 times as many new pallets as it shipped. How many pallets does it have now?" Here's how the three approaches actually differ in what the model writes on its way to the answer:

STANDARD:
  A: 44

CHAIN-OF-THOUGHT:
  A: First, the warehouse starts with 20 pallets.
     Next, it ships 12 pallets, leaving 20 - 12 = 8 pallets.
     Then, it receives 3 times the shipped amount, which is
     3 x 12 = 36 new pallets.
     Finally, adding these together: 8 + 36 = 44 pallets.
     The answer is 44.

CHAIN-OF-DRAFT:
  A: 20 - 12 = 8
     3 x 12 = 36
     8 + 36 = 44
     #### 44

All three can land on the correct answer. The difference is entirely in how much the model had to write to get there — and in production, that written-out reasoning is exactly what you're paying for and waiting on, token by token.

🎯 Use this when: explaining the technique to a non-technical stakeholder — this side-by-side is usually the moment it clicks.

4️⃣ How to Actually Write a Chain-of-Draft Prompt

  1. Write a clear system instruction stating the reasoning-length rule directly — something like: "Work through the problem, but hold each intermediate note to roughly five words. Give the final answer after a separator line."
  2. Provide 2–4 few-shot examples in the exact CoD style you want — this is the part that actually teaches the constraint, since it's a soft guideline, not an enforced limit. Skipping this step is the single most common reason CoD prompts underperform.
  3. Pick a clear answer separator (like ####) so downstream code can reliably split the terse reasoning from the final answer, the same way you'd structure any other parseable LLM output.
  4. Test on your actual task distribution, not just easy examples — arithmetic and simple commonsense reasoning respond very well to CoD; tasks requiring long-context synthesis need more validation before you commit to it in production.
✅ Practical example — a support-ticket triage prompt: A prompt engineer building an internal ticket router could write: "Work out the urgency using short notes of five words or fewer each, then give HIGH, MEDIUM, or LOW after '####'." Paired with two or three short worked examples in that format, this keeps reasoning traceable for audit purposes without paying full CoT-length costs on every single ticket — a meaningful saving at high ticket volume.

5️⃣ Benchmark Results Prompt Engineers Should Know

✅ Real-world example — the paper's own headline result: The Chain-of-Draft paper's own abstract reports that the technique matches or beats Chain-of-Thought's accuracy while needing only about one-thirteenth of CoT's token count — roughly a 7.6% share — across the reasoning benchmarks the authors tested. Independent technology press covering the release corroborates this with task-specific detail: on the GSM8K arithmetic-reasoning benchmark and on BIG-bench commonsense tasks like date and sports understanding, the paper reports CoD reaching accuracy close to or matching CoT while cutting output tokens dramatically — in one reported case, average output tokens for a sports-understanding task dropped by well over 90% compared to CoT, with latency falling in step.

💬 In Plain English: The headline number to remember isn't "CoD is smarter than CoT" — it's roughly "CoD gets you CoT-level answers for a fraction of the writing." For most production use cases, that's the number that actually shows up on your API bill and your response-time dashboard.

6️⃣ Enterprise Rollout at Scale

  1. Maintain a shared prompt-pattern library. Store your tested CoD few-shot templates (by task type — arithmetic, classification, triage) in a shared repo, the same way a platform team would version a shared code library, so individual product teams aren't each hand-rolling their own five-word-step examples from scratch.
  2. A/B test against your existing CoT prompts before switching. Accuracy parity is task-dependent — validate on your own eval set, not just the paper's benchmarks, before replacing a production CoT prompt.
  3. Track token cost and latency per prompting strategy as a first-class metric, alongside accuracy — this is what makes the token savings visible to the business, not just to the engineering team.
  4. Keep an escalation path back to full CoT (or CoT plus self-consistency) for harder cases — some enterprise deployments route by estimated task difficulty, using CoD as the fast default and CoT as the fallback for a subset of harder queries.

7️⃣ Common Mistakes

  1. Stating the five-word limit without few-shot examples. Because the constraint isn't enforced by any decoding mechanism, telling the model the rule in words alone is a much weaker signal than showing it two or three worked examples in that exact style.
  2. Applying CoD to tasks that genuinely need long deliberation. Multi-hop research questions, tasks needing external tool calls mid-reasoning, or problems where the model benefits from "thinking out loud" to catch its own errors are poor fits — the terse format can silently drop the exact context the model needed.
  3. Skipping a validation pass on your own eval set. Published benchmark parity doesn't automatically transfer to a different domain, language, or task shape — verify on your own data before shipping.
  4. Not giving the model a clear answer separator. Without one, downstream parsing has to guess where the "draft" ends and the answer begins, which reintroduces exactly the kind of fragile string-parsing that structured prompting is supposed to avoid.
  5. Treating CoD as a token-savings trick only. The latency reduction is often the bigger production win — for real-time or high-throughput systems, faster responses can matter more than the API cost line item.

❓ FAQ

Q: Is the five-word limit strictly enforced?

A: No. It's a soft guideline taught entirely through the few-shot examples in the prompt — there's no decoding-level mechanism forcing the model to stop at five words. This is exactly why the few-shot examples matter so much when writing a CoD prompt.

Q: Does Chain-of-Draft replace Chain-of-Thought entirely?

A: No — treat it as an additional tool, not a wholesale replacement. CoD is a strong default for well-defined, latency- or cost-sensitive reasoning tasks. Tasks that need extended deliberation, self-correction, or mid-reasoning knowledge retrieval are generally better served by full CoT, or CoT paired with self-consistency.

Q: Who developed Chain-of-Draft, and why does that matter?

A: It was introduced by researchers at Zoom Communications. It matters for context because it comes from a team building real-time products (like a live meeting assistant), which is a strong signal for why latency — not just accuracy — was a central design goal, not an afterthought.

Q: Does CoD work on any LLM, or only specific models?

A: The published research evaluates it across multiple major commercial models, and the underlying mechanism — few-shot demonstration of a terse reasoning style — doesn't depend on any model-specific feature. As with any prompting technique, actual performance varies by model and task, so validating on the specific model you're deploying is still the right move.

Q: What's the single biggest reason a CoD prompt underperforms in practice?

A: Missing or weak few-shot examples. Since the length constraint is learned purely by demonstration, a prompt that states the rule but shows the model verbose examples (or no examples at all) tends to drift back toward CoT-length reasoning.

🔗 References & Further Reading

All product and company names (Zoom, OpenAI, Anthropic) are trademarks of their respective owners and are referenced here only for factual, educational purposes. All technical claims above are synthesized and explained in original wording based on the cited primary research and publicly available reporting — no text has been reproduced verbatim from any source.

📝 Summary

  • Chain-of-Draft (CoD) prompts a model to reason in terse, ≤5-word steps instead of full CoT-style sentences, taught entirely through few-shot examples.
  • It was built by Zoom Communications' research team, motivated by real latency constraints in live products, and can match or beat CoT accuracy using a fraction of the tokens.
  • Writing an effective CoD prompt hinges on strong few-shot examples and a clear answer separator — the length rule alone isn't enough.
  • It's not a universal CoT replacement — tasks needing deep deliberation or mid-reasoning tool use still favor full CoT.
  • At enterprise scale, treat CoD prompts as governed, versioned templates with their own cost/latency/accuracy metrics — not a one-off prompt tweak.


Comments