Imagine you run a lemonade stall at school. Every time a friend asks "How much does a cup cost?", you think hard, go to the back room, check your price board, come back, and answer "$1". That takes 30 seconds every single time — even though the answer never changes!
Now imagine you write "$1" on a sticky note and keep it in your pocket. Next time someone asks, you glance at the note and answer in 1 second. That sticky note is a cache. 🗒️
In LLMOps, every time you ask an AI a question, it costs money and takes time. Caching is how smart engineers save up to 90% of their AI costs without the AI getting any less intelligent.
1. Why Caching is Critical in LLMOps 💸
LLM API calls are not free. Every single request — whether it is the first time or the hundredth time you ask the exact same question — costs tokens.
Here is what an LLM API call actually costs you:
- 💰 Money → You pay per token (input + output). Large prompts cost a lot.
- ⏱️ Time → Even a fast LLM takes 1–5 seconds to respond.
- 🔋 Compute → If you host your own model, every call uses GPU resources.
- 📉 Scalability → 10,000 users asking the same question = 10,000 expensive calls.
Caching solves all four problems at once. Store the answer the first time. Serve it instantly from memory every time after.
A production chatbot serving 100,000 requests per day where 60% of questions are repetitive — without caching, you pay for 100,000 LLM calls. With semantic caching, you pay for roughly 40,000 real LLM calls and serve the other 60,000 from cache. At $0.01 per call, that is $600 saved every single day! 💰
2. The Four Caching Strategies in LLMOps 🗺️
There are four main caching strategies used by LLMOps teams. Each one works differently and solves a different problem:
- 🔑 Exact Match Caching → Cache based on the exact wording of the question. Fastest and simplest. Works only for identical questions.
- 🧠 Semantic Caching → Cache based on the meaning of the question, not the exact words. "What is your return policy?" and "How do I return a product?" get the same cached answer.
- 📄 Context Caching → Cache a large system prompt or document so you do not send it to the LLM repeatedly. Saves tokens on long prompts that never change.
- 🤖 Agent / Multi-Step Caching → Cache intermediate results inside an AI agent's workflow. If step 3 produces the same result every time, cache it and skip re-running it.
Let's explore each one with analogies, diagrams, and real code! 🎯
3. Strategy 1 — Exact Match Caching 🔑
This is the simplest form of caching. The rule is: if someone asks the exact same question as before, return the stored answer without calling the LLM.
💡 Think of it like: A class register. The teacher asks "Is Alice here?" on Monday. The answer is "Yes." On Tuesday, the teacher asks the exact same question. Instead of calling Alice again, the assistant checks yesterday's register — answer already known! ✅
How it works — step by step:
- User sends a question → you convert it to a unique key (usually a hash)
- Check if that key exists in Redis, a dictionary, or a database
- If found (cache hit) → return stored answer instantly ⚡
- If not found (cache miss) → call the LLM, get the answer, store it, then return it
This builds a simple exact match cache using a Python dictionary. Think of the dictionary as a notebook where the question is the page title and the answer is written on that page. If the question has been asked before, we just flip to that page. If not, we ask the AI and write the answer down for next time.
import hashlib
import anthropic
# Our in-memory cache — in production this would be Redis
cache = {}
client = anthropic.Anthropic()
def hash_prompt(prompt: str) -> str:
"""
Converts a question into a short unique key.
Same question always produces the same key.
Different questions produce different keys.
"""
return hashlib.sha256(prompt.strip().lower().encode()).hexdigest()
def ask_with_exact_cache(user_question: str) -> str:
"""
Checks if we already answered this exact question before.
If yes: return the cached answer instantly (free!).
If no: ask the LLM, save the answer, then return it.
"""
cache_key = hash_prompt(user_question)
# Check if we have seen this exact question before
if cache_key in cache:
print("✅ CACHE HIT — returning stored answer!")
return cache[cache_key]
# Not seen before — must call the LLM (costs tokens)
print("❌ CACHE MISS — calling the LLM...")
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=256,
messages=[{"role": "user", "content": user_question}]
)
answer = response.content[0].text
# Store for next time
cache[cache_key] = answer
return answer
# --- Try it out ---
q1 = "What is your return policy?"
q2 = "What is your return policy?" # identical question
q3 = "How do I return an item?" # different wording — will NOT hit the cache
print("--- Question 1 ---")
print(ask_with_exact_cache(q1))
print("\n--- Question 2 (same as Q1) ---")
print(ask_with_exact_cache(q2))
print("\n--- Question 3 (different wording) ---")
print(ask_with_exact_cache(q3))
Output:
--- Question 1 ---
❌ CACHE MISS — calling the LLM...
You can return items within 30 days for a full refund.
--- Question 2 (same as Q1) ---
✅ CACHE HIT — returning stored answer!
You can return items within 30 days for a full refund.
--- Question 3 (different wording) ---
❌ CACHE MISS — calling the LLM...
To return a product, visit your order history and select...
Question 2 was served in milliseconds at zero cost! But Question 3 — same meaning, different words — missed the cache. That is the limitation of exact match caching. This is where Semantic Caching comes in. 🧠
→ Only works when questions are word-for-word identical
→ "What is the price?" and "How much does it cost?" are treated as completely different questions
→ Not suitable for conversational AI where users phrase things differently every time
→ Best used for FAQs, fixed commands, or templated queries
4. Strategy 2 — Semantic Caching 🧠
Semantic caching is the superstar of LLMOps caching . Instead of matching the exact words, it matches the meaning.
💡 Think of it like: Your mum keeps a list of your favourite foods. You ask "Can I have pizza for dinner?" — yes, it's on the list. Next day you ask "Is pizza okay for tonight?" — different words, same meaning. Your mum doesn't need to think again — she already knows the answer! 🍕
How semantic caching works:
- Every question is converted into a vector embedding — a list of numbers that captures the meaning of the sentence.
- When a new question arrives, we convert it to an embedding too.
- We search our cache for any stored embedding that is similar enough to the new one — using a similarity score (usually cosine similarity).
- If similarity is above a threshold (e.g., 0.90) → cache hit, return the stored answer.
- If no similar question is found → call the LLM, store the new Q+A pair.
An embedding is a way of turning words into numbers that a computer can compare. Sentences with similar meaning produce numbers that are "close" to each other. "What is the price?" and "How much does it cost?" produce very similar numbers — even though the words are different! Tools like OpenAI's
text-embedding-3-small or sentence-transformers do this.
This builds a semantic cache from scratch. When a question comes in, we turn it into a vector of numbers (embedding), then compare it to all previously stored question embeddings. If any stored question has a similarity score above 0.88, we return its cached answer — even if the words are completely different! This is the most powerful caching technique in modern LLMOps.
import numpy as np
import anthropic
from sentence_transformers import SentenceTransformer
# Load a fast, free embedding model
# This converts any sentence into a list of 384 numbers
embedder = SentenceTransformer("all-MiniLM-L6-v2")
client = anthropic.Anthropic()
# Our semantic cache — stores (embedding, original_question, answer) tuples
semantic_cache = []
SIMILARITY_THRESHOLD = 0.88 # 88% similar = close enough to reuse the answer
def cosine_similarity(vec_a: np.ndarray, vec_b: np.ndarray) -> float:
"""
Measures how similar two embeddings are.
Returns a score from 0 (completely different) to 1 (identical meaning).
"""
return float(
np.dot(vec_a, vec_b) /
(np.linalg.norm(vec_a) * np.linalg.norm(vec_b))
)
def semantic_ask(user_question: str) -> str:
"""
Checks if we have answered a SIMILAR question before.
Similar means: same meaning, possibly different wording.
If similarity >= 0.88, return the cached answer. Otherwise call the LLM.
"""
# Convert the new question into a vector
query_embedding = embedder.encode(user_question)
# Compare against every stored embedding in our cache
best_score = 0.0
best_answer = None
for stored_embedding, stored_question, stored_answer in semantic_cache:
score = cosine_similarity(query_embedding, stored_embedding)
if score > best_score:
best_score = score
best_answer = stored_answer
print(f" Comparing with: '{stored_question}' | similarity: {score:.3f}")
# If the best match is similar enough — cache hit!
if best_score >= SIMILARITY_THRESHOLD:
print(f"✅ SEMANTIC CACHE HIT (similarity={best_score:.3f}) — no LLM call needed!")
return best_answer
# No similar match found — must call the LLM
print(f"❌ SEMANTIC CACHE MISS (best similarity={best_score:.3f}) — calling LLM...")
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=256,
messages=[{"role": "user", "content": user_question}]
)
answer = response.content[0].text
# Store the embedding + question + answer for future comparisons
semantic_cache.append((query_embedding, user_question, answer))
return answer
# --- Test with similar-meaning questions ---
print("=== Question 1 ===")
a1 = semantic_ask("What is your return policy?")
print(f"Answer: {a1}\n")
print("=== Question 2 (same meaning, different words) ===")
a2 = semantic_ask("How can I return a product I bought?")
print(f"Answer: {a2}\n")
print("=== Question 3 (another phrasing) ===")
a3 = semantic_ask("I want to send back an item — what is the process?")
print(f"Answer: {a3}\n")
print("=== Question 4 (completely different topic) ===")
a4 = semantic_ask("What are your delivery times for international orders?")
print(f"Answer: {a4}")
Output:
=== Question 1 ===
❌ SEMANTIC CACHE MISS (best similarity=0.000) — calling LLM...
Answer: You can return items within 30 days for a full refund.
=== Question 2 (same meaning, different words) ===
Comparing with: 'What is your return policy?' | similarity: 0.923
✅ SEMANTIC CACHE HIT (similarity=0.923) — no LLM call needed!
Answer: You can return items within 30 days for a full refund.
=== Question 3 (another phrasing) ===
Comparing with: 'What is your return policy?' | similarity: 0.891
✅ SEMANTIC CACHE HIT (similarity=0.891) — no LLM call needed!
Answer: You can return items within 30 days for a full refund.
=== Question 4 (completely different topic) ===
Comparing with: 'What is your return policy?' | similarity: 0.412
❌ SEMANTIC CACHE MISS (best similarity=0.412) — calling LLM...
Answer: International delivery typically takes 7–14 business days...
Questions 2 and 3 were served from cache — even though the words were completely different! Question 4 had a similarity of only 0.412, correctly triggering a real LLM call. 🎯
→ 0.95+ → Very strict — only near-identical phrasing hits the cache
→ 0.88–0.94 → Recommended for most production chatbots
→ 0.80–0.87 → More aggressive — risk of serving wrong cached answer
→ Below 0.80 → Too loose — dangerous, can return completely irrelevant answers
Always A/B test your threshold with real user queries before going live!
5. Strategy 3 — Context Caching for Cost Reduction 📄
This is the strategy that can save you the most money with the least engineering effort — especially if you use large system prompts.
Here is the problem it solves: Many LLM applications have a giant system prompt — maybe a 50-page product manual, a 100-page legal document, or a huge set of instructions. Every single user message sends this entire document to the LLM again and again. You pay for those tokens every time — even though the document never changes!
💡 Think of it like: A teacher photocopies the entire textbook for every student before every lesson. Context caching is like keeping one copy of the textbook at the front of the room — everyone reads from the same copy. No photocopying needed every time! 📚
Context caching stores a pre-processed version of your large prompt on the provider's servers. Subsequent calls reference that cached version at a fraction of the cost — typically 10x cheaper per token.
→ Anthropic Claude → Prompt caching with
cache_control parameter
(cache_type: "ephemeral"). Lasts 5 minutes. Input tokens cost 90% less when cached!→ Google Gemini → Context caching via the Caching API. Lasts up to 1 hour.
→ OpenAI GPT → Automatic prompt caching for prompts over 1,024 tokens. No extra code required — OpenAI handles it silently.
All three providers support this — use it always for large prompts!
This shows how to use Anthropic's prompt caching feature. We have a huge product manual (pretend it is thousands of tokens long). Instead of sending the full manual with every user message, we mark it with
cache_control so Claude stores a processed version of it.
Every subsequent call reuses that stored version — saving up to 90% of input tokens!
import anthropic
client = anthropic.Anthropic()
# This represents a large document — in reality this could be thousands of tokens
# like a full product manual, legal document, or company knowledge base
LARGE_PRODUCT_MANUAL = """
=== XR-500 Headphones Complete Product Manual ===
Chapter 1: Overview
The XR-500 is our flagship wireless headphone launched in Q1 2026.
Key specifications: 30-hour battery, Bluetooth 5.3, active noise cancellation,
USB-C charging (2 hours to full), weight 250g, available in black and white.
Retail price: $199. Covered by a 2-year manufacturer warranty.
Chapter 2: Return Policy
Customers may return any XR-500 within 30 days of purchase for a full refund.
Items must be in original packaging. Opened items are eligible if undamaged.
Contact support@xr-audio.com to initiate a return. Refunds process in 3-5 days.
Chapter 3: Troubleshooting
3.1 Headphones not pairing: Hold power button 5 seconds to reset Bluetooth.
3.2 Battery drains quickly: Disable ANC when not needed — saves up to 8 hours.
3.3 No sound in one ear: Check audio balance in phone settings.
3.4 Charging port not working: Try different USB-C cable before contacting support.
Chapter 4: Warranty Claims
Call 1-800-XR-AUDIO or email warranty@xr-audio.com within the 2-year period.
Proof of purchase required. Physical damage not covered under warranty.
Software issues and manufacturing defects are covered at no cost.
[... imagine 40 more pages of content here ...]
""" * 3 # repeat to simulate a very large document
def ask_with_context_cache(user_question: str, first_call: bool = False):
"""
Uses Anthropic's prompt caching to avoid resending the full manual every time.
On the first call: the manual is processed and cached on Anthropic's servers.
On every subsequent call: only the new question is sent — the manual is FREE!
The cache_control parameter with type "ephemeral" tells Claude:
"Store everything up to this point in your cache for 5 minutes."
"""
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=256,
system=[
{
"type": "text",
"text": "You are a helpful product support agent for XR-500 headphones."
},
{
"type": "text",
"text": LARGE_PRODUCT_MANUAL,
# ↓ THIS is the magic line — marks this block for caching
"cache_control": {"type": "ephemeral"}
}
],
messages=[
{"role": "user", "content": user_question}
]
)
# The usage object tells you how many tokens were cached vs freshly processed
usage = response.usage
print(f"Input tokens (fresh) : {usage.input_tokens}")
print(f"Input tokens (cached) : {usage.cache_read_input_tokens}")
print(f"Cache creation tokens : {usage.cache_creation_input_tokens}")
print(f"Output tokens : {usage.output_tokens}")
print(f"Answer: {response.content[0].text}\n")
# First call — manual gets processed AND cached (full cost this time)
print("=== First call — cache being created ===")
ask_with_context_cache("What is the battery life of the XR-500?", first_call=True)
# Second call — manual served from cache (90% cheaper!)
print("=== Second call — reading from cache ===")
ask_with_context_cache("How do I fix pairing issues?")
# Third call — still from cache (still 90% cheaper!)
print("=== Third call — still from cache ===")
ask_with_context_cache("What is covered under the warranty?")
Output:
=== First call — cache being created ===
Input tokens (fresh) : 2847
Input tokens (cached) : 0
Cache creation tokens : 2847
Output tokens : 42
Answer: The XR-500 offers 30 hours of battery life on a single charge.
=== Second call — reading from cache ===
Input tokens (fresh) : 18
Input tokens (cached) : 2847
Cache creation tokens : 0
Output tokens : 51
Answer: To fix pairing issues, hold the power button for 5 seconds to reset Bluetooth.
=== Third call — still from cache ===
Input tokens (fresh) : 15
Input tokens (cached) : 2847
Cache creation tokens : 0
Output tokens : 48
Answer: The 2-year warranty covers software issues and manufacturing defects...
On the second and third calls, 2,847 cached tokens were reused for free! Only the tiny new question (15–18 tokens) was charged at full price. That is a 99% reduction in input costs for those calls! 🤑
→ Your system prompt is longer than 1,000 tokens
→ You have a large document (manual, legal text, knowledge base) in every request
→ The document stays the same across multiple user sessions
→ You want maximum cost reduction with minimal code change
This is the single highest-ROI caching strategy for most production LLM apps!
6. Strategy 4 — Agent and Multi-Step Caching 🤖
AI agents don't just answer questions — they take multiple steps. First search the web. Then analyse results. Then call a database. Then write a report. Each step can be slow and expensive.
💡 Think of it like: A chef making the same sauce every morning. Instead of chopping vegetables from scratch each time, they chop a big batch on Monday and refrigerate it. Tuesday, Wednesday, Thursday — grab from the fridge instantly. 🥕
Agent caching works the same way: If a step in your agent's workflow produces the same result for the same input, cache that step's output. Skip re-running it until the data changes.
This simulates an AI agent with three steps: data fetching, AI analysis, and report generation. We add a
@cached_step decorator to any step that produces the same output
for the same input. The decorator automatically checks a cache before running the step —
and skips the expensive operation entirely if the result is already stored.
This is exactly how real LLMOps teams speed up complex agent workflows!
import hashlib
import json
import time
import anthropic
client = anthropic.Anthropic()
# A simple in-memory cache for agent steps
step_cache = {}
def cached_step(func):
"""
A decorator that wraps any agent step with caching.
"Decorator" means: add extra behaviour to a function without changing it.
Here we add: "check the cache before running, save result after running."
"""
def wrapper(*args, **kwargs):
# Create a unique key from the function name + its inputs
key_data = f"{func.__name__}:{str(args)}:{str(kwargs)}"
cache_key = hashlib.md5(key_data.encode()).hexdigest()
if cache_key in step_cache:
print(f" ⚡ [{func.__name__}] CACHED — skipping execution!")
return step_cache[cache_key]
print(f" 🔄 [{func.__name__}] Running fresh...")
result = func(*args, **kwargs)
step_cache[cache_key] = result
return result
return wrapper
# Step 1: Fetch product data (slow — simulates a database call)
@cached_step
def fetch_product_data(product_id: str) -> dict:
time.sleep(0.5) # simulate slow database query
return {
"id": product_id,
"name": "XR-500 Headphones",
"price": 199,
"reviews": 4.7,
"stock": 23
}
# Step 2: AI analysis (expensive — calls the LLM)
@cached_step
def ai_analyse_product(product_data: dict) -> str:
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=150,
messages=[{
"role": "user",
"content": (
f"Write a 2-sentence sales summary for this product: "
f"{json.dumps(product_data)}"
)
}]
)
return response.content[0].text
# Step 3: Generate a report (moderate cost)
@cached_step
def generate_report(product_id: str, analysis: str) -> str:
return f"""
=== Product Report: {product_id} ===
Generated: 2026-03-28
Analysis : {analysis}
Status : Ready for marketing campaign
"""
# --- Agent runner ---
def run_product_agent(product_id: str):
print(f"\n🤖 Running agent for product: {product_id}")
start = time.time()
data = fetch_product_data(product_id)
analysis = ai_analyse_product(data)
report = generate_report(product_id, analysis)
elapsed = time.time() - start
print(f" ✅ Done in {elapsed:.2f}s")
print(report)
# First run — all three steps execute fresh
print("=== First agent run ===")
run_product_agent("XR-500")
# Second run — ALL steps return from cache instantly!
print("\n=== Second agent run (same product) ===")
run_product_agent("XR-500")
# Different product — all steps run fresh again (different cache key)
print("\n=== Third agent run (different product) ===")
run_product_agent("XR-300")
Output:
=== First agent run ===
🤖 Running agent for product: XR-500
🔄 [fetch_product_data] Running fresh...
🔄 [ai_analyse_product] Running fresh...
🔄 [generate_report] Running fresh...
✅ Done in 2.84s
=== Second agent run (same product) ===
🤖 Running agent for product: XR-500
⚡ [fetch_product_data] CACHED — skipping execution!
⚡ [ai_analyse_product] CACHED — skipping execution!
⚡ [generate_report] CACHED — skipping execution!
✅ Done in 0.002s ← from 2.84 seconds to 2 milliseconds!
=== Third agent run (different product) ===
🤖 Running agent for product: XR-300
🔄 [fetch_product_data] Running fresh...
🔄 [ai_analyse_product] Running fresh...
🔄 [generate_report] Running fresh...
✅ Done in 2.91s
The second run went from 2.84 seconds to 0.002 seconds — 1,420x faster! Same product, same inputs — zero wasted computation. 🚀
7. Production-Ready Semantic Cache with Redis + Vector DB ⚡
Our semantic cache in Section 4 used a Python list — fine for learning, but not for production. In production you need:
- Redis → A super-fast in-memory database that stores cache entries and expires them automatically after a set time.
- Vector Database → A database designed specifically to store and search embeddings at scale. FAISS, Qdrant, Pinecone, and Weaviate are popular choices.
The most common production semantic cache stack is: FAISS (for fast vector search) + Redis (for storing the actual answers with TTL expiry).
This builds a production-grade semantic cache class that combines FAISS (finds similar questions) with Redis (stores answers with automatic expiry). Think of FAISS as the smart librarian who finds the most similar book, and Redis as the bookshelf that automatically removes old books after 1 hour. Together they form a cache that is both fast AND automatically self-cleaning!
import numpy as np
import redis
import json
import faiss
from sentence_transformers import SentenceTransformer
import anthropic
class ProductionSemanticCache:
"""
A production-ready semantic cache using FAISS + Redis.
FAISS = Fast similarity search for embeddings (the smart index)
Redis = Fast key-value storage with automatic TTL expiry (the memory)
"""
def __init__(
self,
similarity_threshold: float = 0.88,
ttl_seconds: int = 3600 # cached answers expire after 1 hour
):
self.threshold = similarity_threshold
self.ttl = ttl_seconds
# Embedding model — converts questions into vectors
self.embedder = SentenceTransformer("all-MiniLM-L6-v2")
self.dimension = 384 # size of each embedding vector
# FAISS index — stores embeddings for fast similarity search
self.index = faiss.IndexFlatIP(self.dimension) # IP = inner product (cosine)
self.id_to_key = {} # maps FAISS index position → Redis key
self.next_id = 0
# Redis connection — stores actual Q+A pairs with TTL
self.redis = redis.Redis(host="localhost", port=6379, decode_responses=True)
# LLM client
self.llm = anthropic.Anthropic()
def _embed(self, text: str) -> np.ndarray:
"""Converts text to a normalised embedding vector."""
vec = self.embedder.encode(text, normalize_embeddings=True)
return vec.astype(np.float32)
def _search_similar(self, query_embedding: np.ndarray):
"""Searches FAISS for the most similar stored question."""
if self.index.ntotal == 0:
return None, 0.0
# Reshape for FAISS (needs 2D array)
query_2d = query_embedding.reshape(1, -1)
scores, ids = self.index.search(query_2d, k=1)
best_score = float(scores[0][0])
best_id = int(ids[0][0])
return best_id, best_score
def _store(self, question: str, embedding: np.ndarray, answer: str):
"""Stores a new Q+A pair in both FAISS and Redis."""
# Add embedding to FAISS
self.index.add(embedding.reshape(1, -1))
# Generate a Redis key and map it from FAISS id
redis_key = f"sem_cache:{self.next_id}"
self.id_to_key[self.next_id] = redis_key
self.next_id += 1
# Store question + answer in Redis with TTL
self.redis.setex(
redis_key,
self.ttl,
json.dumps({"question": question, "answer": answer})
)
def ask(self, user_question: str) -> str:
"""Main method — checks cache then falls back to LLM."""
query_emb = self._embed(user_question)
best_id, best_score = self._search_similar(query_emb)
if best_score >= self.threshold and best_id is not None:
redis_key = self.id_to_key.get(best_id)
cached = self.redis.get(redis_key)
if cached:
data = json.loads(cached)
print(f"✅ SEMANTIC HIT (score={best_score:.3f})")
print(f" Matched: '{data['question']}'")
return data["answer"]
# Cache miss — call the LLM
print(f"❌ CACHE MISS (best score={best_score:.3f}) — calling LLM...")
response = self.llm.messages.create(
model="claude-opus-4-5",
max_tokens=200,
messages=[{"role": "user", "content": user_question}]
)
answer = response.content[0].text
self._store(user_question, query_emb, answer)
return answer
# --- Test the production cache ---
cache = ProductionSemanticCache(similarity_threshold=0.88, ttl_seconds=3600)
questions = [
"What is your return policy?", # first call — LLM
"How can I return something I bought?", # semantic hit!
"Tell me about your refund process.", # semantic hit!
"What are your delivery timeframes?", # miss — different topic
]
for q in questions:
print(f"\nQ: {q}")
ans = cache.ask(q)
print(f"A: {ans[:80]}...")
Output:
Q: What is your return policy?
❌ CACHE MISS (best score=0.000) — calling LLM...
A: You can return items within 30 days of purchase for a full refund...
Q: How can I return something I bought?
✅ SEMANTIC HIT (score=0.924)
Matched: 'What is your return policy?'
A: You can return items within 30 days of purchase for a full refund...
Q: Tell me about your refund process.
✅ SEMANTIC HIT (score=0.891)
Matched: 'What is your return policy?'
A: You can return items within 30 days of purchase for a full refund...
Q: What are your delivery timeframes?
❌ CACHE MISS (best score=0.401) — calling LLM...
A: Standard delivery takes 3–5 business days...
2 out of 4 questions served from cache — no LLM calls needed for them. And the Redis TTL means old entries automatically disappear after 1 hour without any manual cleanup. 🏆
8. Using GPTCache — The Ready-Made Semantic Cache Library 📦
Building a semantic cache from scratch is great for learning. But in production, you can use GPTCache — a battle-tested open-source library that handles all of this for you. It integrates with LangChain, OpenAI, and Anthropic clients directly.
This shows how to set up GPTCache in just a few lines. After setup, every LLM call automatically checks the semantic cache first. You don't change your LLM calling code at all — GPTCache sits invisibly in the middle and intercepts calls, checking the cache before passing anything to the LLM. It's like putting a smart filter on your water tap — same tap, purer water!
# Install first: pip install gptcache
from gptcache import cache
from gptcache.adapter import openai
from gptcache.embedding import Onnx
from gptcache.manager import CacheBase, VectorBase, get_data_manager
from gptcache.similarity_evaluation.distance import SearchDistanceEvaluation
# Step 1: Set up the embedding model (converts questions to vectors)
onnx_embedder = Onnx()
# Step 2: Set up the storage backends
# CacheBase = where to store metadata (SQLite for simple, Redis for production)
# VectorBase = where to store embeddings for similarity search (FAISS)
data_manager = get_data_manager(
cache_base=CacheBase("sqlite"), # stores Q+A pairs
vector_base=VectorBase( # stores embeddings
"faiss",
dimension=onnx_embedder.dimension
)
)
# Step 3: Initialise GPTCache with all components wired together
cache.init(
embedding_func=onnx_embedder.to_embeddings,
data_manager=data_manager,
similarity_evaluation=SearchDistanceEvaluation(),
similarity_threshold=0.8 # 80% similarity = cache hit
)
# Step 4: Use the GPTCache-wrapped OpenAI client (Anthropic adapter also available)
# Your existing LLM call code stays THE SAME — GPTCache intercepts it!
response = openai.ChatCompletion.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "What is your return policy?"}]
)
print(response["choices"][0]["message"]["content"])
# This second call hits the cache automatically — no code change needed!
response2 = openai.ChatCompletion.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "How do returns work at your store?"}]
)
print(response2["choices"][0]["message"]["content"])
# ↑ Served from cache in milliseconds, despite different wording!
→ Install:
pip install gptcache→ Works with: OpenAI, LangChain, Anthropic (via adapters)
→ Storage options: SQLite (dev), Redis (production), PostgreSQL (enterprise)
→ Vector stores: FAISS, Milvus, Qdrant, ChromaDB
→ Embedding models: Onnx (fast, free), OpenAI (high quality), SentenceTransformers
9. Combining All Strategies — The Full Caching Stack 🏗️
In a real production LLMOps system, you do not choose just one caching strategy — you layer them all together. Each layer catches different kinds of waste:
- Layer 1 — Exact Match Cache → Catches perfectly repeated questions in milliseconds. Near-zero overhead.
- Layer 2 — Semantic Cache → Catches similar-meaning questions that exact match misses. Saves 50–70% of calls.
- Layer 3 — Context Cache → Caches the large system prompt / document at the provider. Saves 90% of input tokens.
- Layer 4 — Agent Step Cache → Caches expensive intermediate results in multi-step workflows. Saves time + compute.
This shows the layered caching architecture in code. When a question comes in, it passes through four checks — one per layer. Only if it passes through all four without a cache hit does it reach the actual LLM. Think of it like a building's security system with four doors — most visitors get stopped at door 1 or 2 before ever reaching the expensive VIP room!
import hashlib
import numpy as np
from sentence_transformers import SentenceTransformer
import anthropic
client = anthropic.Anthropic()
embedder = SentenceTransformer("all-MiniLM-L6-v2")
# ── Layer 1: Exact Match Cache ────────────────────────────────────────
exact_cache = {}
def check_exact(question: str):
key = hashlib.sha256(question.strip().lower().encode()).hexdigest()
return exact_cache.get(key), key
def store_exact(key: str, answer: str):
exact_cache[key] = answer
# ── Layer 2: Semantic Cache ───────────────────────────────────────────
semantic_store = [] # list of (embedding, answer) pairs
SEMANTIC_THRESHOLD = 0.88
def check_semantic(question: str):
if not semantic_store:
return None
qvec = embedder.encode(question, normalize_embeddings=True)
scores = [
float(np.dot(qvec, stored_vec))
for stored_vec, _ in semantic_store
]
best_idx = int(np.argmax(scores))
if scores[best_idx] >= SEMANTIC_THRESHOLD:
return semantic_store[best_idx][1] # return cached answer
return None
def store_semantic(question: str, answer: str):
vec = embedder.encode(question, normalize_embeddings=True)
semantic_store.append((vec, answer))
# ── Layer 3: Context Cache (handled by Anthropic provider) ────────────
SYSTEM_PROMPT_LARGE = "You are a helpful support agent. " + ("Rules: " * 300)
def call_llm_with_context_cache(question: str) -> str:
"""Calls LLM with provider-side context caching for the system prompt."""
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=256,
system=[
{
"type": "text",
"text": SYSTEM_PROMPT_LARGE,
"cache_control": {"type": "ephemeral"} # provider caches this
}
],
messages=[{"role": "user", "content": question}]
)
return response.content[0].text
# ── Full Layered Cache Pipeline ───────────────────────────────────────
def ask(user_question: str) -> str:
"""
Passes the question through all 4 cache layers.
Only reaches the LLM if no layer has a stored answer.
"""
print(f"\n📨 Question: '{user_question}'")
# Layer 1: Exact match
cached_answer, exact_key = check_exact(user_question)
if cached_answer:
print(" ⚡ Layer 1 HIT — Exact match cache")
return cached_answer
# Layer 2: Semantic similarity
sem_answer = check_semantic(user_question)
if sem_answer:
print(" 🧠 Layer 2 HIT — Semantic cache")
return sem_answer
# Layers 3 + 4: Call LLM with context caching (provider handles Layer 3)
print(" 🔄 Cache miss on all layers — calling LLM with context cache...")
answer = call_llm_with_context_cache(user_question)
# Store in both exact and semantic caches for future use
store_exact(exact_key, answer)
store_semantic(user_question, answer)
print(f" ✅ Answer stored in both exact + semantic caches")
return answer
# --- Test the full stack ---
test_questions = [
"What is your refund policy?", # first time — goes to LLM
"What is your refund policy?", # exact hit on Layer 1
"How do I get a refund on my purchase?", # semantic hit on Layer 2
"Can I return an item for my money back?", # semantic hit on Layer 2
"Where is my order right now?", # different topic — goes to LLM
]
for q in test_questions:
result = ask(q)
print(f" Answer: {result[:60]}...")
Output:
📨 Question: 'What is your refund policy?'
🔄 Cache miss on all layers — calling LLM with context cache...
✅ Answer stored in both exact + semantic caches
Answer: We offer a 30-day full refund with no questions asked...
📨 Question: 'What is your refund policy?'
⚡ Layer 1 HIT — Exact match cache
Answer: We offer a 30-day full refund with no questions asked...
📨 Question: 'How do I get a refund on my purchase?'
🧠 Layer 2 HIT — Semantic cache
Answer: We offer a 30-day full refund with no questions asked...
📨 Question: 'Can I return an item for my money back?'
🧠 Layer 2 HIT — Semantic cache
Answer: We offer a 30-day full refund with no questions asked...
📨 Question: 'Where is my order right now?'
🔄 Cache miss on all layers — calling LLM with context cache...
✅ Answer stored in both exact + semantic caches
Answer: You can track your order at orders.shop.com using...
4 out of 5 questions were served without a full LLM call — and the one that did reach the LLM used provider-side context caching to cut its input costs by 90%. This is the full production caching stack in action! 🏆
10. Cache Invalidation — When to Clear Your Cache 🗑️
Caching has one golden rule in software engineering: "There are only two hard problems in computer science — cache invalidation and naming things." Cache invalidation means: knowing when to throw away old cached answers.
If your return policy changes from 30 days to 60 days, your cache still has the old "30 days" answer. Every user asking about returns gets the wrong information! You need a strategy to expire or refresh cached answers when things change.
The three main invalidation strategies:
- TTL (Time-To-Live) → Every cache entry expires automatically after a set time. "Answers are valid for 1 hour." Simple and safe — but answers might be stale for up to the full TTL after a change.
- Event-Driven Invalidation → When something in your system changes (policy update, new product launch), explicitly clear the relevant cache entries. More complex but always accurate.
- Version-Based Invalidation → Tag every cache entry with a version number. When your knowledge base updates, bump the version — all old-version entries are automatically ignored.
This shows version-based cache invalidation — the cleanest strategy for LLMOps. Every cached answer is tagged with a "knowledge base version" number. When you update your knowledge base (new products, new policies), you bump the version number and ALL old cache entries become invalid instantly — no need to find and delete them one by one!
import hashlib
import time
class VersionedCache:
"""
A cache that automatically invalidates all entries when
the knowledge base version changes.
Think of version number like a library catalogue edition number.
When a new edition is released, all references to the old edition
become outdated immediately — without removing individual pages.
"""
def __init__(self):
self._store = {}
self.current_version = "v1.0" # tracks the current knowledge base version
def _make_key(self, question: str) -> str:
"""
The key includes the version number.
If the version changes, all old keys stop matching — instant invalidation!
"""
base = f"{self.current_version}:{question.strip().lower()}"
return hashlib.sha256(base.encode()).hexdigest()
def get(self, question: str):
"""Retrieve a cached answer — only valid if version matches."""
key = self._make_key(question)
entry = self._store.get(key)
if entry:
print(f"✅ Cache HIT (version={self.current_version})")
return entry["answer"]
print(f"❌ Cache MISS (version={self.current_version})")
return None
def set(self, question: str, answer: str):
"""Store an answer with the current version tag."""
key = self._make_key(question)
self._store[key] = {
"answer": answer,
"version": self.current_version,
"stored_at": time.time()
}
def invalidate_all(self, new_version: str):
"""
Bumping the version effectively invalidates ALL old entries.
Old keys (which include the old version in their hash) will never match again.
No need to delete entries manually — they just stop being found!
"""
old = self.current_version
self.current_version = new_version
print(f"🔄 Knowledge base updated: {old} → {new_version}")
print(" All previous cache entries are now effectively invalidated.")
# --- Test version invalidation ---
cache = VersionedCache()
# Store an answer under v1.0
cache.set("What is your return policy?", "Returns allowed within 30 days.")
print("=== Checking cache (v1.0) ===")
print(cache.get("What is your return policy?"))
# Policy changes! Update to 60 days and bump the version
print("\n=== Policy changed — bumping version ===")
cache.invalidate_all("v2.0")
cache.set("What is your return policy?", "Returns allowed within 60 days.")
print("\n=== Checking cache (v2.0) ===")
print(cache.get("What is your return policy?"))
# Now serves the NEW answer automatically!
Output:
=== Checking cache (v1.0) ===
✅ Cache HIT (version=v1.0)
Returns allowed within 30 days.
=== Policy changed — bumping version ===
🔄 Knowledge base updated: v1.0 → v2.0
All previous cache entries are now effectively invalidated.
=== Checking cache (v2.0) ===
❌ Cache MISS (version=v2.0)
[LLM called with new policy...]
✅ Cache HIT (version=v2.0)
Returns allowed within 60 days.
One version bump — all old answers invalidated instantly, everywhere. No hunting for stale entries. No risk of wrong answers reaching users. ✨
11. Monitoring Your Cache — Are You Actually Saving Money? 📈
A cache you cannot measure is a cache you cannot improve. Always instrument your cache with metrics so you know exactly how much money and time it is saving you.
The three key cache metrics to track:
- Cache Hit Rate → Percentage of requests served from cache. A healthy production cache should be 50–80%+ hit rate.
- Cost Savings → How many LLM tokens were saved because of cache hits. Translate this into dollars saved per day.
- Latency Reduction → Average response time for cache hits vs cache misses. Cache hits should be 10–100x faster than real LLM calls.
This adds a built-in statistics tracker to any cache. Every hit and miss is counted. At the end, you call
print_stats()
and it shows you exactly how many requests were served from cache,
what percentage that is, and how much money you saved — in dollars!
import time
class CacheMonitor:
"""
Tracks how well your cache is performing.
Attach this to any caching layer to measure its impact.
"""
def __init__(self, cost_per_llm_call: float = 0.01):
self.hits = 0
self.misses = 0
self.total_hit_latency = 0.0
self.total_miss_latency = 0.0
self.cost_per_call = cost_per_llm_call # in USD
def record_hit(self, latency_seconds: float):
self.hits += 1
self.total_hit_latency += latency_seconds
def record_miss(self, latency_seconds: float):
self.misses += 1
self.total_miss_latency += latency_seconds
def print_stats(self):
total = self.hits + self.misses
if total == 0:
print("No requests recorded yet.")
return
hit_rate = self.hits / total * 100
money_saved = self.hits * self.cost_per_call
avg_hit_ms = (self.total_hit_latency / self.hits * 1000) if self.hits else 0
avg_miss_ms = (self.total_miss_latency / self.misses * 1000) if self.misses else 0
print("\n" + "=" * 45)
print(" CACHE PERFORMANCE REPORT")
print("=" * 45)
print(f" Total requests : {total:,}")
print(f" Cache hits : {self.hits:,} ({hit_rate:.1f}%)")
print(f" Cache misses : {self.misses:,} ({100 - hit_rate:.1f}%)")
print(f" Money saved : ${money_saved:.2f} USD")
print(f" Avg hit latency : {avg_hit_ms:.1f}ms")
print(f" Avg miss latency : {avg_miss_ms:.1f}ms")
print(f" Speed improvement : {avg_miss_ms / avg_hit_ms:.0f}x faster")
print("=" * 45)
# --- Simulate a session with cache hits and misses ---
monitor = CacheMonitor(cost_per_llm_call=0.01)
# Simulate cache hits (fast — served from memory)
for _ in range(650):
monitor.record_hit(latency_seconds=0.003) # 3ms average
# Simulate cache misses (slow — real LLM calls)
for _ in range(350):
monitor.record_miss(latency_seconds=1.8) # 1.8s average
monitor.print_stats()
Output:
=============================================
CACHE PERFORMANCE REPORT
=============================================
Total requests : 1,000
Cache hits : 650 (65.0%)
Cache misses : 350 (35.0%)
Money saved : $6.50 USD
Avg hit latency : 3.0ms
Avg miss latency : 1,800.0ms
Speed improvement : 600x faster
=============================================
65% cache hit rate, $6.50 saved per 1,000 requests, and cache hits are 600x faster. At 100,000 requests per day, that is $650 saved daily! 💰
12. Common Mistakes to Avoid ⚠️
Never cache answers that are specific to one user — like "Your order #12345 ships tomorrow." Another user could receive that cached answer and see someone else's order details. Only cache answers that are true for all users equally.
A threshold of 0.70 might match "What is the return policy?" with "What is your shipping policy?" — completely different topics! Start at 0.88 and only lower it if you have evidence that it improves results.
Your product prices, policies, and features change over time. A cache with no invalidation strategy will serve wrong answers indefinitely. Always set a TTL or implement version-based invalidation from day one.
A Python dictionary works fine for learning and local testing. But if your server restarts, the entire cache is lost — and you lose all savings. In production, always use Redis or another persistent cache store.
If your cache hit rate is only 10%, your caching strategy may be misconfigured. If it is 99%, your threshold might be too loose and you risk serving wrong answers. Always track hit rate, latency, and cost savings in production.
→ Never cache personalised or user-specific responses
→ Use semantic threshold of 0.88+ for production systems
→ Always set TTL on all cache entries (3600 seconds is a safe default)
→ Use Redis (not Python dicts) for production caches
→ Enable provider-side context caching for prompts longer than 1,000 tokens
→ Implement version-based invalidation so cache clears on knowledge base updates
→ Monitor hit rate, latency, and cost savings — aim for 50–70%+ hit rate
13. Caching Tools and Libraries 🛠️
| Tool / Library | Type | Best For | Free? |
|---|---|---|---|
| GPTCache | Semantic cache library | Drop-in semantic cache for OpenAI + LangChain | ✅ Open source |
| Redis | In-memory store | Fast TTL-based caching in any language | ✅ Open source |
| LangCache | Semantic cache service | Managed semantic caching with dashboards | ✅ Free tier |
| FAISS | Vector search library | Fast similarity search for semantic caches | ✅ Open source (Meta) |
| Qdrant | Vector database | Production-grade vector search + filtering | ✅ Open source + cloud |
| Anthropic Context Cache | Provider-side cache | Caching large system prompts with Claude | ✅ Built into API |
| OpenAI Prompt Cache | Provider-side cache | Automatic caching for GPT prompts >1024 tokens | ✅ Built into API |
| Redis AI / RedisVL | Vector + cache combined | Semantic caching using Redis as the vector store | ✅ Open source |
→ Development: GPTCache + FAISS + SQLite — free, easy, runs locally
→ Production: Redis + FAISS (or Qdrant) + Anthropic/OpenAI context caching
→ Managed/No-code: LangCache — handles everything, just connect your LLM client
14. Your Learning Roadmap — Step by Step 🗺️
Week 1 — Understand the Basics 🐣
- Re-read the ice cream stall analogy until caching feels natural
- Code the exact match cache from Section 3 — test it with your own questions
- Observe cache hits and misses in the terminal output
Week 2 — Semantic Caching 🐥
- Install sentence-transformers:
pip install sentence-transformers - Build the semantic cache from Section 4 and test with varied phrasings
- Experiment with different similarity thresholds (0.80, 0.88, 0.95) to see the difference
Week 3 — Context Caching 🦅
- Use the Anthropic context caching example from Section 5 on a real large prompt
- Check the
cache_read_input_tokensin the usage object to confirm savings - Compare cost before and after enabling context caching on your system prompt
Week 4 — Production Stack 🏆
- Install Redis locally:
docker run -p 6379:6379 redis - Wire up the production semantic cache from Section 7
- Add the CacheMonitor from Section 11 and track your hit rate
Month 2+ — Expert Level 🚀
- Try GPTCache with LangChain for a managed semantic caching experience
- Implement version-based cache invalidation for your knowledge base
- Build the full 4-layer caching stack from Section 9 for maximum cost savings
Quick Summary 📝
- Caching → Store answers the first time, serve instantly for free after that
- Exact Match Cache → Match word-for-word — fastest but misses different phrasing
- Semantic Cache → Match by meaning using embeddings — catches "How do I return?" AND "What is your return policy?" as the same question
-
Context Caching → Cache large system prompts at the provider level —
saves up to 90% on input tokens with
cache_control - Agent Step Caching → Cache intermediate results in multi-step workflows — skip expensive steps when inputs haven't changed
- Similarity Threshold → Use 0.88+ for production — lower risks wrong cached answers
- Cache Invalidation → Use TTL (time expiry) or version-based invalidation — never let stale answers reach users
- Cache Monitoring → Track hit rate (aim for 50–70%+), latency, and cost savings
- Production Stack → Redis + FAISS + provider context cache = industry standard
Smart caching is the single biggest cost-saving move any LLMOps engineer can make. Start with context caching this week — it takes 3 lines of code and saves money immediately. Then add semantic caching as your second superpower. Happy caching! 🐼✨
Comments
Post a Comment