Prompt caching is a feature offered by LLM providers that lets you store the unchanged part of a prompt — a long system instruction, a set of tool definitions, a reference document — so that repeated requests reusing that same content skip full reprocessing and get billed at a steep discount instead. It's the single highest-leverage lever most teams have for cutting both API cost and response latency, and it usually requires no change to what the model actually says back. 💸
This matters because unmanaged token spend is one of the most common ways an LLM feature quietly becomes unprofitable. A coding assistant that resends a 20,000-token codebase summary on every single autocomplete request, or a document Q&A tool that re-sends the same 100-page PDF for every follow-up question, pays full price for the same tokens over and over — and pays that price in both dollars and the seconds a user spends staring at a loading spinner. Anthropic's own published testing found that caching a 100,000-token document cut response time from 11.5 seconds down to 2.4 seconds on a repeated request — a 79% latency drop on exactly the kind of workload most production LLM apps run constantly. ⚡
The same request, twice: full price and full wait the first time, a fraction of both after that.
📑 In This Post
- What Is Prompt Caching, Really?
- Why Costs and Latency Balloon Without It
- How Prefix Caching Actually Works Under the Hood
- Provider-by-Provider Mechanics
- Designing Prompts for Cache-Friendliness
- Cache Lifetime, TTL, and Invalidation Tradeoffs
- Semantic Caching vs. Prefix Caching
- Measuring Cache Hit Rate and ROI
- Rolling This Out at Enterprise Scale
- Common Mistakes (and Why They Happen)
- Hands-On Lab: Measure a Cache Hit Yourself
- FAQ
- References & Further Reading
- Summary
🔀 Quick Comparison: How the Major Providers Cache
"Turn on caching" means something different depending on which provider you're calling. Here's the practical shape of each approach before we go deep.
| Provider | How Caching Is Enabled | Typical Discount | Default Lifetime |
|---|---|---|---|
| Anthropic (Claude) | Explicit cache_control breakpoints, or one automatic top-level breakpoint |
~90% off cache reads; cache writes cost more up front | 5 minutes (1 hour optional) |
| OpenAI / Azure OpenAI | Automatic by default; current-generation models also support explicit developer-placed breakpoints | Up to ~90% off cache reads on current models (older models: a smaller, model-dependent discount) | ~30 minutes on current models; roughly 5–10 min (or up to 24h extended) on older ones |
| Google Gemini | Automatic "implicit" caching, or manually created "explicit" caches for guaranteed savings | Varies; explicit caching guarantees a lower rate on cached tokens | Implicit: short-lived; explicit: configurable TTL, 1 hour default |
1. What Is Prompt Caching, Really?
🧒 Kid analogy: Imagine packing the exact same lunch — same sandwich, same snack, same drink — every single school day. Instead of walking to the store and buying fresh ingredients each morning, you'd batch-prep once and just grab the same prepped lunch from the fridge each day, only swapping out whatever's actually different, like today's homework folder. Prompt caching is that fridge: the parts of your prompt that never change get "pre-packed" once, and every later request just grabs them instead of rebuilding them from scratch.
Every request to an LLM has to convert its input text into a numeric representation the model can process — this happens through the model's attention mechanism, which produces intermediate values commonly called key-value (KV) pairs for every token in the prompt. Normally, that computation happens fresh for the entire prompt on every single call, even if 95% of the prompt is byte-for-byte identical to the last call. Prompt caching lets a provider store those already-computed KV values for a stable prefix of tokens, so a later request sharing that exact prefix can skip recomputing it and load the stored result instead — cutting both the compute cost and the time spent waiting.
The key word is prefix: caching works on the leading, unchanged portion of a prompt, in the exact order it appears — a change anywhere inside that stable portion invalidates the cache for everything after the change point. This is why prompt structure (Section 5) ends up mattering so much: put the parts that change every request at the very end, and the parts that stay the same come first and stay cacheable.
✅ Worked example: When Anthropic launched prompt caching, Notion was an early adopter for its Claude-powered Notion AI assistant. Notion's own leadership described the appeal in straightforward terms: the ability to make their AI assistant respond faster and more cheaply without giving up answer quality — exactly the two variables this whole post is about, cost and latency, moving in the same direction at once instead of trading off against each other.
📊 The actual math, redoable with a calculator: Take a support bot built on Claude Sonnet 5, using published per-token pricing of $2 per million tokens for standard input, $2.50 per million for a 5-minute cache write, and $0.20 per million for a cache read. Suppose its system prompt plus knowledge base runs 12,000 tokens (stable, cacheable) and each user turn adds about 300 new tokens, handling 5,000 requests in a day with traffic frequent enough that the cache rarely lapses.
| Scenario | Daily token cost | Monthly cost |
|---|---|---|
| No caching (all 12,300 tokens/request at standard rate) | $123.00 | $3,690.00 |
| With caching (1 write + 4,999 reads + fresh turn tokens) | $15.03 | $450.83 |
That's roughly $108 saved per day, or about $3,239 a month — an 88% reduction — on a single moderately-sized prompt, using nothing but Anthropic's published list pricing. The larger and more frequently reused the stable prefix, the bigger this gap gets, which is exactly why Anthropic's own 100,000-token benchmark (Section 1's opening figure) showed an even steeper 79% drop in response time on top of the cost savings.
💡 Where it gets harder: Caching only pays off when the same prefix actually comes back again before it expires. If your traffic pattern rarely repeats the same long prefix within the cache's lifetime, you can end up paying a cache-write premium on every request and never collecting the discount — which is exactly why Section 6's TTL tradeoffs matter as much as the mechanism itself.
🎯 Use this when: any part of your prompt — instructions, tools, reference documents, conversation history — repeats across more than one request.
2. Why Costs and Latency Balloon Without It
🧒 Kid analogy: Picture re-reading an entire textbook chapter from page one every single time you want to look up just one new fact at the end of it, instead of keeping a bookmark and jumping straight to where you left off. That's what an LLM does by default on every request — it "re-reads" the whole prompt from the beginning, even the parts it already processed a moment ago.
Three ordinary application patterns quietly multiply token spend and latency the same way: multi-turn conversations, where each new turn resends the entire conversation history so far, meaning turn ten reprocesses everything from turns one through nine all over again; retrieval-augmented generation, where the same retrieved passages or whole reference documents get attached to every follow-up question about them; and agentic tool use, where a growing list of tool definitions and prior tool call results gets included in every step of a multi-step agent loop. In each case, the "new" information in a given request is often a small fraction of the total tokens — the rest is identical content paid for and waited on again.
The latency cost isn't just about the extra tokens being billed — it's that processing more input tokens takes real compute time before the model can even start generating its answer, so every uncached duplicate token adds to the delay a user experiences before the first word of a response appears. This compounds badly in agent loops specifically, where a single user request might trigger many internal model calls in sequence, each one resending the same growing context.
✅ Worked example: A document-analysis tool includes a 3,000-token document in every request. Five follow-up questions about that same document means the model processes 15,000 tokens of completely identical document content across those five calls — paid for, and waited on, five separate times for content that never changed once.
💡 Key warning: This cost is easy to miss in early testing because a handful of manual test calls never reveals the multiplier — it only becomes visible (and expensive) once real traffic volume and real conversation length show up in production.
🎯 Use this when: you're reviewing a token-cost bill that seems disproportionate to how much your product's inputs actually change turn to turn.
3. How Prefix Caching Actually Works Under the Hood
🧒 Kid analogy: If you and a friend are both baking the exact same cake, and your friend already measured out and mixed all the dry ingredients, you don't need to redo that step — you can just start from their bowl and add the wet ingredients yourself. The cache is that pre-mixed bowl: someone (an earlier request) already did the first part of the work, and later requests build on top of it instead of starting over.
Mechanically, a caching-enabled request works in three steps. First, the provider checks whether a prompt's leading tokens — up to a marked or automatically detected point — match a prefix it processed recently and still has stored. Second, if a match is found, the provider loads that stored computation instead of reprocessing those tokens, and only the genuinely new tokens after the matching point get processed fresh. Third, if no match is found, the full prompt is processed as normal, and the prefix is written into the cache so a future request has something to match against — this first "cold" call typically costs a small premium over standard pricing specifically because writing to the cache is itself extra work.
Anthropic's documentation is explicit that this reference is exact and positional: the cache covers the full prefix — tools, system prompt, and messages, in that order — ending right at whichever content block carries a cache breakpoint, and a hash of that exact content is what gets matched on subsequent calls. This is why caching is sometimes described as "byte-exact": there's no fuzzy matching involved at this layer, which is the key distinction from the semantic caching technique covered in Section 7.
✅ Worked example: A support-bot prompt has a 4,000-token system prompt (stable), 15 tool definitions (stable), and a growing conversation history (changes every turn). Marking a cache breakpoint right after the tool definitions means the system prompt and tools get cached once and reused across every conversation using that bot — only the actual conversation content needs fresh processing each turn.
💡 Where it gets harder: A single-character difference anywhere inside the cached prefix — a timestamp accidentally embedded in a "static" system prompt, a reordered tool list, a dynamically injected user name early in the prompt — silently breaks the match for everything after it, and the request quietly falls back to full-price, full-latency processing with no error message telling you why.
🎯 Use this when: you want to understand why a caching setup that "should" be working isn't showing up in your cost or latency numbers.
4. Provider-by-Provider Mechanics
🧒 Kid analogy: Two different libraries might both let you "renew" a book instead of returning and re-checking it out, but one requires you to fill out a renewal slip yourself while the other just automatically renews it if you show up again before it's due. Same underlying idea, different amount of effort required from you — which is exactly the split between explicit and automatic caching across providers.
Anthropic's API supports both automatic and explicit caching. Explicit caching means placing a cache_control marker directly on specific content blocks — up to four breakpoints per request — giving fine control over exactly which sections get cached, useful when different parts of a prompt change at different frequencies (tools rarely change, retrieved context changes daily, for instance). Automatic caching adds one top-level marker and lets the system push the breakpoint forward to whatever the newest cacheable content happens to be, advancing it automatically as a multi-turn conversation grows — simpler to set up, at the cost of some of that fine-grained control. Anthropic's system also checks roughly 20 content-block positions before an explicit breakpoint for a matching cached prefix, so a single well-placed breakpoint near the end of a stable section is often enough without needing several.
OpenAI's approach has converged toward Anthropic's over time: earlier models cached automatically only, with a smaller, model-dependent discount and no separate charge for writing to the cache, while OpenAI's current-generation models support both automatic (implicit) placement and developer-placed explicit breakpoints, with cache reads discounted to roughly a tenth of the standard input rate and cache writes carrying a modest premium — a pricing shape that now looks a great deal like Anthropic's. Google's Gemini API offers both: implicit caching happens automatically on supported models with no cost-saving guarantee, while explicit caching lets a developer deliberately create a named cache object with a configurable time-to-live, guaranteeing the discount applies whenever that specific cache is referenced. Azure OpenAI generally mirrors whichever OpenAI model generation it's built on, so its caching behavior — automatic-only versus automatic-plus-explicit, and the size of the discount — depends on which underlying model version a deployment is running.
✅ Worked example: A team porting the same RAG chatbot from an older GPT model to Claude needs to actively add cache_control breakpoints to get predictable caching, since the older model relied entirely on automatic placement — but the same port to a current-generation GPT model can use that model's own explicit breakpoint controls instead, since both providers now expose similar fine-grained control, just through different parameter names.
💡 Key warning: "Automatic" doesn't mean "guaranteed." Several providers explicitly note that automatic or implicit caching comes with no guarantee of a cache hit or savings on any specific request — for workloads where the savings need to be predictable and auditable, explicit caching is the safer design choice even where it takes more setup work.
🎯 Use this when: you're choosing a model provider for a cost-sensitive, high-repetition workload, or migrating one between providers.
5. Designing Prompts for Cache-Friendliness
🧒 Kid analogy: If you're packing a suitcase you'll partly repack every day of a trip, you put the things you won't touch — like the shoes you're not wearing that day — at the bottom, and the things you'll dig out every morning near the top. Pack it the other way around and you're unpacking everything just to reach what actually changes. Prompt structure works the same way: stable stuff goes first, changing stuff goes last.
Because caching works on a prefix match, prompt order is not a cosmetic choice — it directly determines how much of a prompt is even eligible to be cached. The practical rule nearly every provider's documentation converges on is the same: place the most stable content first (system instructions, tool definitions, long reference documents, few-shot examples) and place the content that changes on every single call last (the newest user message, live session state, a timestamp). Content that comes after a cache breakpoint can change freely without invalidating anything before that point — but content placed before it needs to stay genuinely identical.
This has a subtle but important implication for multi-turn conversations specifically: because each new turn appends to the growing message history rather than rewriting it, a stable prefix (system prompt plus tools) followed by an append-only conversation naturally stays cache-friendly turn after turn, since nothing earlier in the sequence gets edited. Conversation designs that instead summarize, reorder, or rewrite earlier turns — a common technique for managing context length — break this property, since editing anything before the breakpoint invalidates the cached prefix from that point forward.
Order isn't stylistic here — it's the difference between a prompt that caches well and one that doesn't.
✅ Worked example: A coding assistant that includes a 10,000-token summary of the current codebase should place that summary immediately after the system prompt and before any per-request context like "the file the user currently has open" — the codebase summary stays cacheable across every autocomplete request in a session, while only the small, genuinely-changing "current file" section needs fresh processing each time.
💡 Where it gets harder: It's tempting to inject something small and "harmless" early in a prompt for convenience — a request ID, a rendered timestamp, a per-user greeting — without realizing it sits before the cache breakpoint and invalidates the entire prefix on every single call. Audit anything placed before a breakpoint for hidden per-request variability, not just obviously dynamic content.
🎯 Use this when: you're designing a new prompt template from scratch, or debugging why an existing one gets a lower cache hit rate than expected.
6. Cache Lifetime, TTL, and Invalidation Tradeoffs
🧒 Kid analogy: Leftovers in the fridge are only good for so long before you'd rather throw them out and cook fresh. A cache works the same way — it doesn't sit there forever; if nobody "reheats" it (reuses it) within a certain window, it gets cleared out, and the next request has to start from scratch again.
Every provider's cache has a time-to-live (TTL) — a window after which an unused cache entry is evicted. Anthropic's default is five minutes, and reusing that cached content resets the clock again at no extra charge, with an optional one-hour duration for workloads where requests are naturally spaced further apart. OpenAI's current-generation models default to roughly a 30-minute lifetime, while older models typically stay warm for somewhere between five and ten minutes of quiet before expiring, unless a longer, extended-retention option is configured. Gemini's explicit caches default to a one-hour TTL, configurable per cache.
The tradeoff is straightforward but easy to get backwards: a short TTL means less risk of paying to store content nobody ends up reusing, but a higher chance of missing a cache hit if requests are naturally spaced further apart than the window. A longer TTL captures more reuse for spaced-out traffic but means paying to keep content cached during idle periods when it might never be read again. The right choice depends entirely on your actual request cadence — a customer support bot handling a live back-and-forth conversation rarely needs more than the default short window, while a document-analysis workflow where a user reads for several minutes between questions may benefit meaningfully from the longer option.
✅ Worked example: A live chat support tool where users typically respond within 30 seconds of the previous message fits comfortably inside a 5-minute default TTL. A research assistant where users often read a long generated report for several minutes before asking a follow-up question is a much better candidate for the extended 1-hour caching option.
💡 Key warning: A longer TTL is not automatically "more efficient" — it usually comes with a higher write premium or an explicit ongoing storage cost, and paying that premium for content that turns out not to be reused within the window is a net loss, not a safety margin.
🎯 Use this when: your default cache configuration produces a lower hit rate than your actual traffic pattern should allow.
7. Semantic Caching vs. Prefix Caching
🧒 Kid analogy: A librarian who only helps you if you ask for a book by its exact, precise title is doing prefix-style matching. A librarian who recognizes that "that book about the boy wizard" and "the first Harry Potter book" mean the same thing is doing something closer to semantic matching — understanding the meaning behind different wordings, not just checking for an identical string.
Everything covered so far in this post is prefix caching — an exact, byte-level match on the leading portion of a prompt, handled entirely by the model provider. Semantic caching is a different, application-level technique: instead of matching exact prompt text, it compares the meaning of a new incoming question against previously answered questions using embedding similarity, and if a sufficiently similar question was already answered, it returns that stored answer directly, skipping a fresh model call entirely rather than just skipping the reprocessing of a shared prefix.
The two techniques solve different problems and are frequently used together rather than as alternatives: prefix caching reduces the cost and latency of the input side of a request that still needs a fresh, correct answer; semantic caching can skip generating a new answer altogether for a genuinely repeated question, at the real risk of returning a stale or subtly wrong answer if the similarity threshold is too loose or the underlying facts have changed since the cached answer was generated. Prefix caching has no correctness risk of this kind, since the model still generates a fresh answer every time — only the input processing is reused.
✅ Worked example: A customer-facing FAQ bot might use semantic caching to instantly return a stored answer to "what are your business hours" no matter how the question is phrased, while also using prefix caching underneath for the system prompt and knowledge-base excerpts, so even the small fraction of genuinely novel questions still process quickly and cheaply.
💡 Where it gets harder: Semantic caching needs its own eval discipline — the same golden-dataset and grading practices from earlier posts in this series apply directly, specifically to check that the similarity threshold isn't quietly serving a wrong or outdated answer to a question that only sounds similar to a previously cached one.
🎯 Use this when: a meaningful share of your traffic asks genuinely repeated questions in different words, not just repeated context.
8. Measuring Cache Hit Rate and ROI
🧒 Kid analogy: Installing a smoke detector and never checking the battery isn't really safety — it's just hoping. Turning on caching and never checking whether it's actually hitting is the same kind of blind trust; you need to look at the numbers to know the thing you set up is doing anything at all.
Every major provider surfaces cache usage directly in the API response: a count of tokens served from cache versus tokens processed fresh, typically alongside a separate count for tokens newly written to the cache on that call. Tracking the ratio of cached tokens to total input tokens over time — the cache hit rate — is the single most direct signal of whether a caching setup is actually working, and it should be treated as a first-class metric alongside cost and latency dashboards, not a one-time verification during initial setup.
A dropping hit rate over time is itself a diagnostic worth alerting on: it can mean a recent prompt template change accidentally introduced dynamic content before a cache breakpoint, a shift in traffic patterns spread requests further apart than the TTL window, or a personalization feature started injecting per-user content earlier in the prompt than intended. Because the write premium is only worth paying if the discount is collected often enough afterward, a team should periodically calculate the actual net savings — write costs incurred minus read discounts earned — rather than assuming any nonzero hit rate is automatically a win.
It's worth knowing what a realistic hit rate actually looks like before assuming your own number is too low or too high. OpenAI's own documentation walks through two concrete example deployments with real measured outcomes: a single-turn LLM-as-judge setup — a fixed grading rubric and a handful of few-shot examples ahead of the breakpoint, with the interaction being judged placed after it — reported a token cache-hit rate of roughly 70%. A multi-turn agent, which appended new tool calls and results onto a shared, growing prefix rather than rewriting earlier turns, reported a hit rate above 90%. The gap between those two numbers tells you something structural, not just anecdotal: an append-only, ever-growing conversation prefix caches dramatically better than a workload where most of the prompt is freshly generated context on every single call.
✅ Worked example: A team notices their cache hit rate dropped from 85% to 40% after a routine deploy. Checking recent prompt template changes reveals a new feature that injects the user's current timezone directly into the system prompt, ahead of the cache breakpoint — moving it to after the breakpoint restores the original hit rate immediately.
💡 Key warning: A high hit rate on a low-traffic prefix and a modest hit rate on your highest-volume prefix are not equally important — weight cache monitoring toward the prefixes carrying the most request volume, since that's where a percentage-point change in hit rate translates into the most real dollars and seconds. There's also a hard number worth knowing: with Anthropic's standard 5-minute-cache pricing (a 1.25× write cost and a 0.1× read cost relative to base input), the math works out so caching only beats paying standard input price outright once your hit rate clears roughly 22%. Below that, in a low-traffic, rarely-repeated prefix, the write premium isn't earning its keep, and caching can quietly cost more than not bothering at all.
🎯 Use this when: you've enabled caching and want to confirm — with numbers, not assumption — that it's actually delivering savings.
9. Rolling This Out at Enterprise Scale
🧒 Kid analogy: Sharing one bike between two siblings is simple to manage informally. A whole apartment building sharing a bike rack needs rules about whose bike is whose, who's responsible if one goes missing, and a system so two people don't grab the same one at once. Caching across a company's many prompts and teams needs that same shift from informal habit to shared, tracked infrastructure.
Prompt caching introduces its own specific set of enterprise concerns on top of the ones covered earlier in this series.
Data residency and isolation for cached content. A cache entry is, functionally, a temporarily stored copy of whatever was in the cached prefix — which can include real customer data if that data sits before the breakpoint. Providers generally isolate caches so they aren't shared across accounts or subscriptions, but a team handling regulated data still needs to confirm that caching doesn't create a retention window that conflicts with a data-handling policy or contractual deletion requirement.
Cost governance and budget forecasting. Because caching changes the actual unit economics of a prompt — a cache write costing more than standard input, a cache read costing far less — cost forecasting needs to model expected hit rate explicitly rather than pricing every request at either the full or the fully-discounted rate. A rollout with an optimistic assumed hit rate that doesn't match real traffic can produce a forecast that's wrong in either direction.
Template governance for cache-breakpoint placement. Section 5 and Section 8 both point at the same operational risk: an innocuous-looking prompt template edit can silently move dynamic content ahead of a cache breakpoint. Treating cache breakpoint placement as a reviewed, deliberate part of a prompt template — not an implementation detail buried in code — keeps this from drifting unnoticed across many engineers editing the same shared templates.
Observability integrated with the eval and CI infrastructure from earlier posts. Cache hit rate belongs on the same dashboards as prompt quality and cost metrics from this series' evaluation post, and ideally the same CI gate that catches a quality regression should also flag a prompt template change that measurably drops the expected cache hit rate for a high-volume prompt.
Provider-portability planning. Because caching mechanics genuinely differ across providers (Section 4), a company running multi-provider infrastructure — for redundancy, cost arbitrage, or avoiding vendor lock-in — needs its prompt templates and internal tooling to handle explicit breakpoint placement for providers that require it, rather than assuming every provider behaves like the most automatic one.
💡 Key warning: Treating caching purely as a backend performance optimization, invisible to the people who write and edit prompts, is how the Section 5 and Section 8 failure modes keep recurring — the people most likely to accidentally break a cache breakpoint are the ones who don't know one exists in the prompt they're editing.
🎯 Use this when: caching is being rolled out across more than one team's prompts, or a prompt handling regulated or sensitive data is being made cacheable for the first time.
10. Common Mistakes (and Why They Happen)
These patterns come up repeatedly for teams adopting prompt caching for the first time, each with an understandable cause.
Putting dynamic content before the cache breakpoint. This happens because a small, "obviously harmless" piece of dynamic content — a timestamp, a request ID, a personalized greeting — doesn't feel risky to place wherever is convenient in the prompt. Section 5's discipline of auditing everything before a breakpoint for hidden variability is the direct fix.
Enabling caching once and never checking the hit rate again. This happens because a successful initial test feels like proof the setup works permanently, when in fact template edits, traffic pattern shifts, and new features can all silently degrade it later. Section 8's ongoing monitoring discipline treats hit rate as a metric to watch continuously, not a box to check once.
Assuming a longer TTL is always better. This happens because a longer cache lifetime intuitively sounds like "more caching, more savings," without accounting for the write premium or storage cost paid regardless of whether that longer window actually gets used. Section 6's traffic-pattern-based reasoning is the corrective.
Confusing semantic caching's correctness risk with prefix caching's total safety. This happens because both are called "caching" and both save money, so it's easy to assume they carry the same risk profile. They don't — Section 7's distinction between an exact-match technique with no answer-correctness risk and a similarity-based technique that can serve a stale or wrong answer needs to inform how cautiously each is deployed.
Restructuring a prompt for caching without re-running the eval suite. This happens because reordering content or adding a breakpoint feels like a purely mechanical change that shouldn't affect output quality. In practice, moving content around a prompt can change how a model weighs different instructions, so any caching-motivated restructuring should go through the same CI-gated regression testing described in this series' evaluation post before shipping.
Ignoring caching mechanics differences when porting a prompt across providers. This happens because a prompt that "works" on one provider looks portable at a glance, when the caching behavior underneath — automatic versus explicit, byte-exact matching rules, TTL defaults — can differ enough to change the actual cost and latency profile of the same prompt on a different provider.
11. Hands-On Lab: Measure a Cache Hit Yourself
This lab needs only an API key for a provider that reports cache usage in its response, a long piece of text to serve as your stable context, and about fifteen minutes.
Find or write a reference text of at least a few thousand words — a long article, documentation page, or public-domain book excerpt works fine — and place it as the system prompt or the first message, well above the minimum token threshold your provider requires for caching to activate. Expect to see: a prompt where the reference text makes up the large majority of total tokens.
Send a first request asking a short question about the text, enabling caching per your provider's method (an explicit cache marker for Claude or a current-generation GPT model, or simply relying on automatic caching for an older GPT model or implicit Gemini caching), and record the response time and the usage field's token breakdown. Expect to see: this first call reported as a cache miss, or showing cache-write tokens rather than cache-read tokens.
Within the cache's TTL window (a couple of minutes is safe for any provider's default), send a second request with the exact same reference text and breakpoint placement, but a different follow-up question. Record the response time and usage breakdown again. Expect to see: a visibly faster response and a usage field showing a meaningful share of tokens served from cache rather than processed fresh.
Now deliberately break the cache: change one word somewhere in the middle of your reference text (before the breakpoint) and resend the same follow-up question from Step 3. Expect to see: the request reported as a cache miss again, and the response time returning to roughly what it was in Step 2 — direct, hands-on proof of how positional and exact the prefix match really is.
Common first-timer mistake: waiting too long between Step 2 and Step 3 and accidentally letting the cache expire, then concluding caching "isn't working" when it was simply never given a chance to hit. Check the TTL for whichever provider you're testing before you conclude anything from a missed cache hit.
Those four requests demonstrate everything Sections 3 through 6 describe in miniature: a cache write on the first call, a cache hit on the second, an intentional invalidation on the third, and a TTL boundary you can observe directly by timing your own requests. Production caching setups differ mainly in scale and automation — the underlying mechanics you just watched happen by hand are exactly what's running under a high-volume application, just triggered by real user traffic instead of a manual script.
❓ FAQ
Does prompt caching change the model's actual answers?
No — prefix caching only reuses already-computed processing of the input tokens; the model still generates a fresh response every time based on the full prompt content. It affects cost and latency, not what the model says. Semantic caching is the exception, since it can return a previously generated answer directly rather than generating a new one.
Is prompt caching worth setting up for a low-traffic application?
It depends on repetition, not raw volume. Even a low-traffic app can benefit meaningfully if each user has a long, stable context (a big document, a long system prompt) that gets reused across several questions in the same session — the savings come from reuse within a short window, not from overall scale.
Why did my cache hit rate suddenly drop after a deploy?
The most common cause is a prompt template edit that introduced dynamic content — a timestamp, a personalized field, a reordered section — somewhere before the cache breakpoint. Diffing the most recent prompt template changes against what's actually placed before the breakpoint usually finds it quickly.
Can I combine prompt caching with prompt optimization or evaluation from earlier in this series?
Yes, and it's worth doing deliberately rather than as an afterthought. Any prompt restructuring done to improve cache-friendliness should go back through the same eval suite used to validate the prompt originally, since reordering content can occasionally shift how a model weighs different instructions even when the words themselves haven't changed.
Is cached content stored permanently, and is it secure?
No — caches are ephemeral by design, expiring automatically after their TTL window regardless of use, and providers generally isolate cached content so it isn't shared across different accounts or subscriptions. Teams handling regulated data should still confirm the specific retention and isolation guarantees in their provider's documentation rather than assuming they match a competitor's.
🔗 References & Further Reading
Official/primary documentation consulted for accuracy:
- Anthropic, "Prompt caching with Claude" (customer announcement) — anthropic.com/news/prompt-caching
- Anthropic, "Prompt caching" developer documentation — platform.claude.com/en/docs/build-with-claude/prompt-caching
- OpenAI, "Prompt caching" developer guide — developers.openai.com/api/docs/guides/prompt-caching
- OpenAI, "Prompt Caching in the API" (original 2024 announcement) — openai.com/index/api-prompt-caching
- Google, "Context caching" (Gemini API documentation) — ai.google.dev/gemini-api/docs/caching
- Microsoft, "Prompt caching" (Azure OpenAI / Azure AI Foundry documentation) — learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/prompt-caching
Additional practitioner background reading (used only to confirm general industry practice, not as a source of quoted or closely-followed text):
- General practitioner discussion of the distinction between provider-side prefix/KV caching and application-level semantic caching, and of common cache-hit-rate monitoring practices.
All product and company names referenced (Anthropic, Claude, OpenAI, Google/Gemini, Microsoft/Azure OpenAI, Notion) are trademarks of their respective owners.
📝 Summary
- Prompt caching stores the already-computed processing of a stable prompt prefix so later requests can reuse it, cutting both cost and latency without changing the model's actual answers.
- Multi-turn conversations, RAG pipelines, and agent loops all quietly resend large amounts of identical content on every call, which is exactly what caching is designed to eliminate.
- Caching works through an exact, positional prefix match on stored key-value computations — any change inside the cached portion invalidates everything after that point.
- Anthropic uses explicit breakpoints (plus an automatic option), OpenAI and Azure OpenAI now support both automatic and explicit caching on current models (with older models relying on automatic-only placement), and Gemini offers both automatic implicit caching and guaranteed explicit caching.
- Prompt structure directly determines cache-friendliness: stable content first, ever-changing content last, with careful attention to anything small and "harmless" that might sneak in before the breakpoint.
- Cache TTL is a real tradeoff between capturing more reuse and paying to store content that never gets reused — the right window depends on actual request cadence, not intuition.
- Semantic caching is a different, application-level technique that can skip generating a new answer entirely, and it carries a real correctness risk that exact-match prefix caching does not.
- Cache hit rate deserves the same ongoing monitoring as cost and latency generally, and there's a concrete floor to know: with Anthropic's standard pricing, caching only pays off once your hit rate clears roughly 22% — below that, the write premium can cost more than never caching at all.
- Scaling this across a company needs data governance for cached content, cost forecasting that accounts for realistic hit rates, and template governance so breakpoint placement doesn't silently drift.
- Most common mistakes trace back to invisible, easy-to-miss placement errors or to treating a one-time successful test as permanent proof the setup keeps working.
If you take away one thing: the words your model sees rarely all change at the same rate, and prompt caching is just the discipline of noticing which parts don't and letting them stay put. Happy caching! 🚀
Comments
Post a Comment