Skip to main content

Caching Strategies in LLMOps: Reduce Latency, Cost, and LLM Inference Time

Calculating read time…

Imagine you run a lemonade stall at school. Every time a friend asks "How much does a cup cost?", you think hard, go to the back room, check your price board, come back, and answer "$1". That takes 30 seconds every single time — even though the answer never changes!

Now imagine you write "$1" on a sticky note and keep it in your pocket. Next time someone asks, you glance at the note and answer in 1 second. That sticky note is a cache. 🗒️

In LLMOps, every time you ask an AI a question, it costs money and takes time. Caching is how smart engineers save up to 90% of their AI costs without the AI getting any less intelligent.

1. Why Caching is Critical in LLMOps 💸

LLM API calls are not free. Every single request — whether it is the first time or the hundredth time you ask the exact same question — costs tokens.

Here is what an LLM API call actually costs you:

  • 💰 Money → You pay per token (input + output). Large prompts cost a lot.
  • ⏱️ Time → Even a fast LLM takes 1–5 seconds to respond.
  • 🔋 Compute → If you host your own model, every call uses GPU resources.
  • 📉 Scalability → 10,000 users asking the same question = 10,000 expensive calls.

Caching solves all four problems at once. Store the answer the first time. Serve it instantly from memory every time after.

💡 Real Numbers — Why This Matters:
A production chatbot serving 100,000 requests per day where 60% of questions are repetitive — without caching, you pay for 100,000 LLM calls. With semantic caching, you pay for roughly 40,000 real LLM calls and serve the other 60,000 from cache. At $0.01 per call, that is $600 saved every single day! 💰

2. The Four Caching Strategies in LLMOps 🗺️

There are four main caching strategies used by LLMOps teams. Each one works differently and solves a different problem:

  • 🔑 Exact Match Caching → Cache based on the exact wording of the question. Fastest and simplest. Works only for identical questions.
  • 🧠 Semantic Caching → Cache based on the meaning of the question, not the exact words. "What is your return policy?" and "How do I return a product?" get the same cached answer.
  • 📄 Context Caching → Cache a large system prompt or document so you do not send it to the LLM repeatedly. Saves tokens on long prompts that never change.
  • 🤖 Agent / Multi-Step Caching → Cache intermediate results inside an AI agent's workflow. If step 3 produces the same result every time, cache it and skip re-running it.

Let's explore each one with analogies, diagrams, and real code! 🎯

3. Strategy 1 — Exact Match Caching 🔑

This is the simplest form of caching. The rule is: if someone asks the exact same question as before, return the stored answer without calling the LLM.

💡 Think of it like: A class register. The teacher asks "Is Alice here?" on Monday. The answer is "Yes." On Tuesday, the teacher asks the exact same question. Instead of calling Alice again, the assistant checks yesterday's register — answer already known! ✅

How it works — step by step:

  • User sends a question → you convert it to a unique key (usually a hash)
  • Check if that key exists in Redis, a dictionary, or a database
  • If found (cache hit) → return stored answer instantly ⚡
  • If not found (cache miss) → call the LLM, get the answer, store it, then return it
🎯 What this code block will do:
This builds a simple exact match cache using a Python dictionary. Think of the dictionary as a notebook where the question is the page title and the answer is written on that page. If the question has been asked before, we just flip to that page. If not, we ask the AI and write the answer down for next time.
import hashlib
import anthropic

# Our in-memory cache — in production this would be Redis
cache = {}
client = anthropic.Anthropic()

def hash_prompt(prompt: str) -> str:
    """
    Converts a question into a short unique key.
    Same question always produces the same key.
    Different questions produce different keys.
    """
    return hashlib.sha256(prompt.strip().lower().encode()).hexdigest()

def ask_with_exact_cache(user_question: str) -> str:
    """
    Checks if we already answered this exact question before.
    If yes: return the cached answer instantly (free!).
    If no: ask the LLM, save the answer, then return it.
    """
    cache_key = hash_prompt(user_question)

    # Check if we have seen this exact question before
    if cache_key in cache:
        print("✅ CACHE HIT — returning stored answer!")
        return cache[cache_key]

    # Not seen before — must call the LLM (costs tokens)
    print("❌ CACHE MISS — calling the LLM...")
    response = client.messages.create(
        model="claude-opus-4-5",
        max_tokens=256,
        messages=[{"role": "user", "content": user_question}]
    )
    answer = response.content[0].text

    # Store for next time
    cache[cache_key] = answer
    return answer


# --- Try it out ---
q1 = "What is your return policy?"
q2 = "What is your return policy?"   # identical question
q3 = "How do I return an item?"       # different wording — will NOT hit the cache

print("--- Question 1 ---")
print(ask_with_exact_cache(q1))

print("\n--- Question 2 (same as Q1) ---")
print(ask_with_exact_cache(q2))

print("\n--- Question 3 (different wording) ---")
print(ask_with_exact_cache(q3))

Output:

--- Question 1 ---
❌ CACHE MISS — calling the LLM...
You can return items within 30 days for a full refund.

--- Question 2 (same as Q1) ---
✅ CACHE HIT — returning stored answer!
You can return items within 30 days for a full refund.

--- Question 3 (different wording) ---
❌ CACHE MISS — calling the LLM...
To return a product, visit your order history and select...

Question 2 was served in milliseconds at zero cost! But Question 3 — same meaning, different words — missed the cache. That is the limitation of exact match caching. This is where Semantic Caching comes in. 🧠

❌ Exact Match Caching Limitations:
→ Only works when questions are word-for-word identical
→ "What is the price?" and "How much does it cost?" are treated as completely different questions
→ Not suitable for conversational AI where users phrase things differently every time
→ Best used for FAQs, fixed commands, or templated queries

4. Strategy 2 — Semantic Caching 🧠

Semantic caching is the superstar of LLMOps caching . Instead of matching the exact words, it matches the meaning.

💡 Think of it like: Your mum keeps a list of your favourite foods. You ask "Can I have pizza for dinner?" — yes, it's on the list. Next day you ask "Is pizza okay for tonight?" — different words, same meaning. Your mum doesn't need to think again — she already knows the answer! 🍕

How semantic caching works:

  • Every question is converted into a vector embedding — a list of numbers that captures the meaning of the sentence.
  • When a new question arrives, we convert it to an embedding too.
  • We search our cache for any stored embedding that is similar enough to the new one — using a similarity score (usually cosine similarity).
  • If similarity is above a threshold (e.g., 0.90) → cache hit, return the stored answer.
  • If no similar question is found → call the LLM, store the new Q+A pair.
💡 What is a Vector Embedding?
An embedding is a way of turning words into numbers that a computer can compare. Sentences with similar meaning produce numbers that are "close" to each other. "What is the price?" and "How much does it cost?" produce very similar numbers — even though the words are different! Tools like OpenAI's text-embedding-3-small or sentence-transformers do this.
🎯 What this code block will do:
This builds a semantic cache from scratch. When a question comes in, we turn it into a vector of numbers (embedding), then compare it to all previously stored question embeddings. If any stored question has a similarity score above 0.88, we return its cached answer — even if the words are completely different! This is the most powerful caching technique in modern LLMOps.
import numpy as np import anthropic from sentence_transformers import SentenceTransformer # Load a fast, free embedding model # This converts any sentence into a list of 384 numbers embedder = SentenceTransformer("all-MiniLM-L6-v2") client = anthropic.Anthropic() # Our semantic cache — stores (embedding, original_question, answer) tuples semantic_cache = [] SIMILARITY_THRESHOLD = 0.88 # 88% similar = close enough to reuse the answer def cosine_similarity(vec_a: np.ndarray, vec_b: np.ndarray) -> float: """ Measures how similar two embeddings are. Returns a score from 0 (completely different) to 1 (identical meaning). """ return float( np.dot(vec_a, vec_b) / (np.linalg.norm(vec_a) * np.linalg.norm(vec_b)) ) def semantic_ask(user_question: str) -> str: """ Checks if we have answered a SIMILAR question before. Similar means: same meaning, possibly different wording. If similarity >= 0.88, return the cached answer. Otherwise call the LLM. """ # Convert the new question into a vector query_embedding = embedder.encode(user_question) # Compare against every stored embedding in our cache best_score = 0.0 best_answer = None for stored_embedding, stored_question, stored_answer in semantic_cache: score = cosine_similarity(query_embedding, stored_embedding) if score > best_score: best_score = score best_answer = stored_answer print(f" Comparing with: '{stored_question}' | similarity: {score:.3f}") # If the best match is similar enough — cache hit! if best_score >= SIMILARITY_THRESHOLD: print(f"✅ SEMANTIC CACHE HIT (similarity={best_score:.3f}) — no LLM call needed!") return best_answer # No similar match found — must call the LLM print(f"❌ SEMANTIC CACHE MISS (best similarity={best_score:.3f}) — calling LLM...") response = client.messages.create( model="claude-opus-4-5", max_tokens=256, messages=[{"role": "user", "content": user_question}] ) answer = response.content[0].text # Store the embedding + question + answer for future comparisons semantic_cache.append((query_embedding, user_question, answer)) return answer # --- Test with similar-meaning questions --- print("=== Question 1 ===") a1 = semantic_ask("What is your return policy?") print(f"Answer: {a1}\n") print("=== Question 2 (same meaning, different words) ===") a2 = semantic_ask("How can I return a product I bought?") print(f"Answer: {a2}\n") print("=== Question 3 (another phrasing) ===") a3 = semantic_ask("I want to send back an item — what is the process?") print(f"Answer: {a3}\n") print("=== Question 4 (completely different topic) ===") a4 = semantic_ask("What are your delivery times for international orders?") print(f"Answer: {a4}")

Output:

=== Question 1 ===
❌ SEMANTIC CACHE MISS (best similarity=0.000) — calling LLM...
Answer: You can return items within 30 days for a full refund.

=== Question 2 (same meaning, different words) ===
   Comparing with: 'What is your return policy?' | similarity: 0.923
✅ SEMANTIC CACHE HIT (similarity=0.923) — no LLM call needed!
Answer: You can return items within 30 days for a full refund.

=== Question 3 (another phrasing) ===
   Comparing with: 'What is your return policy?' | similarity: 0.891
✅ SEMANTIC CACHE HIT (similarity=0.891) — no LLM call needed!
Answer: You can return items within 30 days for a full refund.

=== Question 4 (completely different topic) ===
   Comparing with: 'What is your return policy?' | similarity: 0.412
❌ SEMANTIC CACHE MISS (best similarity=0.412) — calling LLM...
Answer: International delivery typically takes 7–14 business days...

Questions 2 and 3 were served from cache — even though the words were completely different! Question 4 had a similarity of only 0.412, correctly triggering a real LLM call. 🎯

✅ Choosing the Right Similarity Threshold:
→ 0.95+ → Very strict — only near-identical phrasing hits the cache
→ 0.88–0.94 → Recommended for most production chatbots
→ 0.80–0.87 → More aggressive — risk of serving wrong cached answer
→ Below 0.80 → Too loose — dangerous, can return completely irrelevant answers
Always A/B test your threshold with real user queries before going live!

5. Strategy 3 — Context Caching for Cost Reduction 📄

This is the strategy that can save you the most money with the least engineering effort — especially if you use large system prompts.

Here is the problem it solves: Many LLM applications have a giant system prompt — maybe a 50-page product manual, a 100-page legal document, or a huge set of instructions. Every single user message sends this entire document to the LLM again and again. You pay for those tokens every time — even though the document never changes!

💡 Think of it like: A teacher photocopies the entire textbook for every student before every lesson. Context caching is like keeping one copy of the textbook at the front of the room — everyone reads from the same copy. No photocopying needed every time! 📚

Context caching stores a pre-processed version of your large prompt on the provider's servers. Subsequent calls reference that cached version at a fraction of the cost — typically 10x cheaper per token.

💡 Context Caching — Provider Support:
→ Anthropic Claude → Prompt caching with cache_control parameter (cache_type: "ephemeral"). Lasts 5 minutes. Input tokens cost 90% less when cached!
→ Google Gemini → Context caching via the Caching API. Lasts up to 1 hour.
→ OpenAI GPT → Automatic prompt caching for prompts over 1,024 tokens. No extra code required — OpenAI handles it silently.
All three providers support this — use it always for large prompts!
🎯 What this code block will do:
This shows how to use Anthropic's prompt caching feature. We have a huge product manual (pretend it is thousands of tokens long). Instead of sending the full manual with every user message, we mark it with cache_control so Claude stores a processed version of it. Every subsequent call reuses that stored version — saving up to 90% of input tokens!
import anthropic

client = anthropic.Anthropic()

# This represents a large document — in reality this could be thousands of tokens
# like a full product manual, legal document, or company knowledge base
LARGE_PRODUCT_MANUAL = """
=== XR-500 Headphones Complete Product Manual ===

Chapter 1: Overview
The XR-500 is our flagship wireless headphone launched in Q1 2026.
Key specifications: 30-hour battery, Bluetooth 5.3, active noise cancellation,
USB-C charging (2 hours to full), weight 250g, available in black and white.
Retail price: $199. Covered by a 2-year manufacturer warranty.

Chapter 2: Return Policy
Customers may return any XR-500 within 30 days of purchase for a full refund.
Items must be in original packaging. Opened items are eligible if undamaged.
Contact support@xr-audio.com to initiate a return. Refunds process in 3-5 days.

Chapter 3: Troubleshooting
3.1 Headphones not pairing: Hold power button 5 seconds to reset Bluetooth.
3.2 Battery drains quickly: Disable ANC when not needed — saves up to 8 hours.
3.3 No sound in one ear: Check audio balance in phone settings.
3.4 Charging port not working: Try different USB-C cable before contacting support.

Chapter 4: Warranty Claims
Call 1-800-XR-AUDIO or email warranty@xr-audio.com within the 2-year period.
Proof of purchase required. Physical damage not covered under warranty.
Software issues and manufacturing defects are covered at no cost.

[... imagine 40 more pages of content here ...]
""" * 3   # repeat to simulate a very large document

def ask_with_context_cache(user_question: str, first_call: bool = False):
    """
    Uses Anthropic's prompt caching to avoid resending the full manual every time.

    On the first call: the manual is processed and cached on Anthropic's servers.
    On every subsequent call: only the new question is sent — the manual is FREE!

    The cache_control parameter with type "ephemeral" tells Claude:
    "Store everything up to this point in your cache for 5 minutes."
    """
    response = client.messages.create(
        model="claude-opus-4-5",
        max_tokens=256,

        system=[
            {
                "type": "text",
                "text": "You are a helpful product support agent for XR-500 headphones."
            },
            {
                "type": "text",
                "text": LARGE_PRODUCT_MANUAL,
                # ↓ THIS is the magic line — marks this block for caching
                "cache_control": {"type": "ephemeral"}
            }
        ],

        messages=[
            {"role": "user", "content": user_question}
        ]
    )

    # The usage object tells you how many tokens were cached vs freshly processed
    usage = response.usage
    print(f"Input tokens (fresh)  : {usage.input_tokens}")
    print(f"Input tokens (cached) : {usage.cache_read_input_tokens}")
    print(f"Cache creation tokens : {usage.cache_creation_input_tokens}")
    print(f"Output tokens         : {usage.output_tokens}")
    print(f"Answer: {response.content[0].text}\n")


# First call — manual gets processed AND cached (full cost this time)
print("=== First call — cache being created ===")
ask_with_context_cache("What is the battery life of the XR-500?", first_call=True)

# Second call — manual served from cache (90% cheaper!)
print("=== Second call — reading from cache ===")
ask_with_context_cache("How do I fix pairing issues?")

# Third call — still from cache (still 90% cheaper!)
print("=== Third call — still from cache ===")
ask_with_context_cache("What is covered under the warranty?")

Output:

=== First call — cache being created ===
Input tokens (fresh)  : 2847
Input tokens (cached) : 0
Cache creation tokens : 2847
Output tokens         : 42
Answer: The XR-500 offers 30 hours of battery life on a single charge.

=== Second call — reading from cache ===
Input tokens (fresh)  : 18
Input tokens (cached) : 2847
Cache creation tokens : 0
Output tokens         : 51
Answer: To fix pairing issues, hold the power button for 5 seconds to reset Bluetooth.

=== Third call — still from cache ===
Input tokens (fresh)  : 15
Input tokens (cached) : 2847
Cache creation tokens : 0
Output tokens         : 48
Answer: The 2-year warranty covers software issues and manufacturing defects...

On the second and third calls, 2,847 cached tokens were reused for free! Only the tiny new question (15–18 tokens) was charged at full price. That is a 99% reduction in input costs for those calls! 🤑

✅ When to Use Context Caching:
→ Your system prompt is longer than 1,000 tokens
→ You have a large document (manual, legal text, knowledge base) in every request
→ The document stays the same across multiple user sessions
→ You want maximum cost reduction with minimal code change
This is the single highest-ROI caching strategy for most production LLM apps!

6. Strategy 4 — Agent and Multi-Step Caching 🤖

AI agents don't just answer questions — they take multiple steps. First search the web. Then analyse results. Then call a database. Then write a report. Each step can be slow and expensive.

💡 Think of it like: A chef making the same sauce every morning. Instead of chopping vegetables from scratch each time, they chop a big batch on Monday and refrigerate it. Tuesday, Wednesday, Thursday — grab from the fridge instantly. 🥕

Agent caching works the same way: If a step in your agent's workflow produces the same result for the same input, cache that step's output. Skip re-running it until the data changes.

🎯 What this code block will do:
This simulates an AI agent with three steps: data fetching, AI analysis, and report generation. We add a @cached_step decorator to any step that produces the same output for the same input. The decorator automatically checks a cache before running the step — and skips the expensive operation entirely if the result is already stored. This is exactly how real LLMOps teams speed up complex agent workflows!
import hashlib
import json
import time
import anthropic

client = anthropic.Anthropic()

# A simple in-memory cache for agent steps
step_cache = {}

def cached_step(func):
    """
    A decorator that wraps any agent step with caching.
    "Decorator" means: add extra behaviour to a function without changing it.
    Here we add: "check the cache before running, save result after running."
    """
    def wrapper(*args, **kwargs):
        # Create a unique key from the function name + its inputs
        key_data = f"{func.__name__}:{str(args)}:{str(kwargs)}"
        cache_key = hashlib.md5(key_data.encode()).hexdigest()

        if cache_key in step_cache:
            print(f"   ⚡ [{func.__name__}] CACHED — skipping execution!")
            return step_cache[cache_key]

        print(f"   🔄 [{func.__name__}] Running fresh...")
        result = func(*args, **kwargs)
        step_cache[cache_key] = result
        return result

    return wrapper


# Step 1: Fetch product data (slow — simulates a database call)
@cached_step
def fetch_product_data(product_id: str) -> dict:
    time.sleep(0.5)   # simulate slow database query
    return {
        "id": product_id,
        "name": "XR-500 Headphones",
        "price": 199,
        "reviews": 4.7,
        "stock": 23
    }


# Step 2: AI analysis (expensive — calls the LLM)
@cached_step
def ai_analyse_product(product_data: dict) -> str:
    response = client.messages.create(
        model="claude-opus-4-5",
        max_tokens=150,
        messages=[{
            "role": "user",
            "content": (
                f"Write a 2-sentence sales summary for this product: "
                f"{json.dumps(product_data)}"
            )
        }]
    )
    return response.content[0].text


# Step 3: Generate a report (moderate cost)
@cached_step
def generate_report(product_id: str, analysis: str) -> str:
    return f"""
    === Product Report: {product_id} ===
    Generated: 2026-03-28
    Analysis : {analysis}
    Status   : Ready for marketing campaign
    """


# --- Agent runner ---
def run_product_agent(product_id: str):
    print(f"\n🤖 Running agent for product: {product_id}")
    start = time.time()

    data     = fetch_product_data(product_id)
    analysis = ai_analyse_product(data)
    report   = generate_report(product_id, analysis)

    elapsed = time.time() - start
    print(f"   ✅ Done in {elapsed:.2f}s")
    print(report)


# First run — all three steps execute fresh
print("=== First agent run ===")
run_product_agent("XR-500")

# Second run — ALL steps return from cache instantly!
print("\n=== Second agent run (same product) ===")
run_product_agent("XR-500")

# Different product — all steps run fresh again (different cache key)
print("\n=== Third agent run (different product) ===")
run_product_agent("XR-300")

Output:

=== First agent run ===
🤖 Running agent for product: XR-500
   🔄 [fetch_product_data] Running fresh...
   🔄 [ai_analyse_product] Running fresh...
   🔄 [generate_report] Running fresh...
   ✅ Done in 2.84s

=== Second agent run (same product) ===
🤖 Running agent for product: XR-500
   ⚡ [fetch_product_data] CACHED — skipping execution!
   ⚡ [ai_analyse_product] CACHED — skipping execution!
   ⚡ [generate_report] CACHED — skipping execution!
   ✅ Done in 0.002s   ← from 2.84 seconds to 2 milliseconds!

=== Third agent run (different product) ===
🤖 Running agent for product: XR-300
   🔄 [fetch_product_data] Running fresh...
   🔄 [ai_analyse_product] Running fresh...
   🔄 [generate_report] Running fresh...
   ✅ Done in 2.91s

The second run went from 2.84 seconds to 0.002 seconds — 1,420x faster! Same product, same inputs — zero wasted computation. 🚀

7. Production-Ready Semantic Cache with Redis + Vector DB ⚡

Our semantic cache in Section 4 used a Python list — fine for learning, but not for production. In production you need:

  • Redis → A super-fast in-memory database that stores cache entries and expires them automatically after a set time.
  • Vector Database → A database designed specifically to store and search embeddings at scale. FAISS, Qdrant, Pinecone, and Weaviate are popular choices.

The most common production semantic cache stack is: FAISS (for fast vector search) + Redis (for storing the actual answers with TTL expiry).

🎯 What this code block will do:
This builds a production-grade semantic cache class that combines FAISS (finds similar questions) with Redis (stores answers with automatic expiry). Think of FAISS as the smart librarian who finds the most similar book, and Redis as the bookshelf that automatically removes old books after 1 hour. Together they form a cache that is both fast AND automatically self-cleaning!
import numpy as np
import redis
import json
import faiss
from sentence_transformers import SentenceTransformer
import anthropic

class ProductionSemanticCache:
    """
    A production-ready semantic cache using FAISS + Redis.

    FAISS  = Fast similarity search for embeddings (the smart index)
    Redis  = Fast key-value storage with automatic TTL expiry (the memory)
    """

    def __init__(
        self,
        similarity_threshold: float = 0.88,
        ttl_seconds: int = 3600          # cached answers expire after 1 hour
    ):
        self.threshold = similarity_threshold
        self.ttl = ttl_seconds

        # Embedding model — converts questions into vectors
        self.embedder = SentenceTransformer("all-MiniLM-L6-v2")
        self.dimension = 384             # size of each embedding vector

        # FAISS index — stores embeddings for fast similarity search
        self.index = faiss.IndexFlatIP(self.dimension)   # IP = inner product (cosine)
        self.id_to_key = {}              # maps FAISS index position → Redis key
        self.next_id = 0

        # Redis connection — stores actual Q+A pairs with TTL
        self.redis = redis.Redis(host="localhost", port=6379, decode_responses=True)

        # LLM client
        self.llm = anthropic.Anthropic()

    def _embed(self, text: str) -> np.ndarray:
        """Converts text to a normalised embedding vector."""
        vec = self.embedder.encode(text, normalize_embeddings=True)
        return vec.astype(np.float32)

    def _search_similar(self, query_embedding: np.ndarray):
        """Searches FAISS for the most similar stored question."""
        if self.index.ntotal == 0:
            return None, 0.0

        # Reshape for FAISS (needs 2D array)
        query_2d = query_embedding.reshape(1, -1)
        scores, ids = self.index.search(query_2d, k=1)

        best_score = float(scores[0][0])
        best_id    = int(ids[0][0])

        return best_id, best_score

    def _store(self, question: str, embedding: np.ndarray, answer: str):
        """Stores a new Q+A pair in both FAISS and Redis."""
        # Add embedding to FAISS
        self.index.add(embedding.reshape(1, -1))

        # Generate a Redis key and map it from FAISS id
        redis_key = f"sem_cache:{self.next_id}"
        self.id_to_key[self.next_id] = redis_key
        self.next_id += 1

        # Store question + answer in Redis with TTL
        self.redis.setex(
            redis_key,
            self.ttl,
            json.dumps({"question": question, "answer": answer})
        )

    def ask(self, user_question: str) -> str:
        """Main method — checks cache then falls back to LLM."""
        query_emb = self._embed(user_question)
        best_id, best_score = self._search_similar(query_emb)

        if best_score >= self.threshold and best_id is not None:
            redis_key = self.id_to_key.get(best_id)
            cached = self.redis.get(redis_key)

            if cached:
                data = json.loads(cached)
                print(f"✅ SEMANTIC HIT (score={best_score:.3f})")
                print(f"   Matched: '{data['question']}'")
                return data["answer"]

        # Cache miss — call the LLM
        print(f"❌ CACHE MISS (best score={best_score:.3f}) — calling LLM...")
        response = self.llm.messages.create(
            model="claude-opus-4-5",
            max_tokens=200,
            messages=[{"role": "user", "content": user_question}]
        )
        answer = response.content[0].text
        self._store(user_question, query_emb, answer)
        return answer


# --- Test the production cache ---
cache = ProductionSemanticCache(similarity_threshold=0.88, ttl_seconds=3600)

questions = [
    "What is your return policy?",           # first call — LLM
    "How can I return something I bought?",  # semantic hit!
    "Tell me about your refund process.",    # semantic hit!
    "What are your delivery timeframes?",    # miss — different topic
]

for q in questions:
    print(f"\nQ: {q}")
    ans = cache.ask(q)
    print(f"A: {ans[:80]}...")

Output:

Q: What is your return policy?
❌ CACHE MISS (best score=0.000) — calling LLM...
A: You can return items within 30 days of purchase for a full refund...

Q: How can I return something I bought?
✅ SEMANTIC HIT (score=0.924)
   Matched: 'What is your return policy?'
A: You can return items within 30 days of purchase for a full refund...

Q: Tell me about your refund process.
✅ SEMANTIC HIT (score=0.891)
   Matched: 'What is your return policy?'
A: You can return items within 30 days of purchase for a full refund...

Q: What are your delivery timeframes?
❌ CACHE MISS (best score=0.401) — calling LLM...
A: Standard delivery takes 3–5 business days...

2 out of 4 questions served from cache — no LLM calls needed for them. And the Redis TTL means old entries automatically disappear after 1 hour without any manual cleanup. 🏆

8. Using GPTCache — The Ready-Made Semantic Cache Library 📦

Building a semantic cache from scratch is great for learning. But in production, you can use GPTCache — a battle-tested open-source library that handles all of this for you. It integrates with LangChain, OpenAI, and Anthropic clients directly.

🎯 What this code block will do:
This shows how to set up GPTCache in just a few lines. After setup, every LLM call automatically checks the semantic cache first. You don't change your LLM calling code at all — GPTCache sits invisibly in the middle and intercepts calls, checking the cache before passing anything to the LLM. It's like putting a smart filter on your water tap — same tap, purer water!
# Install first: pip install gptcache
from gptcache import cache
from gptcache.adapter import openai
from gptcache.embedding import Onnx
from gptcache.manager import CacheBase, VectorBase, get_data_manager
from gptcache.similarity_evaluation.distance import SearchDistanceEvaluation

# Step 1: Set up the embedding model (converts questions to vectors)
onnx_embedder = Onnx()

# Step 2: Set up the storage backends
# CacheBase = where to store metadata (SQLite for simple, Redis for production)
# VectorBase = where to store embeddings for similarity search (FAISS)
data_manager = get_data_manager(
    cache_base=CacheBase("sqlite"),     # stores Q+A pairs
    vector_base=VectorBase(             # stores embeddings
        "faiss",
        dimension=onnx_embedder.dimension
    )
)

# Step 3: Initialise GPTCache with all components wired together
cache.init(
    embedding_func=onnx_embedder.to_embeddings,
    data_manager=data_manager,
    similarity_evaluation=SearchDistanceEvaluation(),
    similarity_threshold=0.8            # 80% similarity = cache hit
)

# Step 4: Use the GPTCache-wrapped OpenAI client (Anthropic adapter also available)
# Your existing LLM call code stays THE SAME — GPTCache intercepts it!
response = openai.ChatCompletion.create(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "What is your return policy?"}]
)
print(response["choices"][0]["message"]["content"])

# This second call hits the cache automatically — no code change needed!
response2 = openai.ChatCompletion.create(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "How do returns work at your store?"}]
)
print(response2["choices"][0]["message"]["content"])
# ↑ Served from cache in milliseconds, despite different wording!
✅ GPTCache Quick Reference:
→ Install: pip install gptcache
→ Works with: OpenAI, LangChain, Anthropic (via adapters)
→ Storage options: SQLite (dev), Redis (production), PostgreSQL (enterprise)
→ Vector stores: FAISS, Milvus, Qdrant, ChromaDB
→ Embedding models: Onnx (fast, free), OpenAI (high quality), SentenceTransformers

9. Combining All Strategies — The Full Caching Stack 🏗️

In a real production LLMOps system, you do not choose just one caching strategy — you layer them all together. Each layer catches different kinds of waste:

  • Layer 1 — Exact Match Cache → Catches perfectly repeated questions in milliseconds. Near-zero overhead.
  • Layer 2 — Semantic Cache → Catches similar-meaning questions that exact match misses. Saves 50–70% of calls.
  • Layer 3 — Context Cache → Caches the large system prompt / document at the provider. Saves 90% of input tokens.
  • Layer 4 — Agent Step Cache → Caches expensive intermediate results in multi-step workflows. Saves time + compute.
🎯 What this code block will do:
This shows the layered caching architecture in code. When a question comes in, it passes through four checks — one per layer. Only if it passes through all four without a cache hit does it reach the actual LLM. Think of it like a building's security system with four doors — most visitors get stopped at door 1 or 2 before ever reaching the expensive VIP room!
import hashlib
import numpy as np
from sentence_transformers import SentenceTransformer
import anthropic

client = anthropic.Anthropic()
embedder = SentenceTransformer("all-MiniLM-L6-v2")

# ── Layer 1: Exact Match Cache ────────────────────────────────────────
exact_cache = {}

def check_exact(question: str):
    key = hashlib.sha256(question.strip().lower().encode()).hexdigest()
    return exact_cache.get(key), key

def store_exact(key: str, answer: str):
    exact_cache[key] = answer


# ── Layer 2: Semantic Cache ───────────────────────────────────────────
semantic_store = []  # list of (embedding, answer) pairs

SEMANTIC_THRESHOLD = 0.88

def check_semantic(question: str):
    if not semantic_store:
        return None
    qvec = embedder.encode(question, normalize_embeddings=True)
    scores = [
        float(np.dot(qvec, stored_vec))
        for stored_vec, _ in semantic_store
    ]
    best_idx = int(np.argmax(scores))
    if scores[best_idx] >= SEMANTIC_THRESHOLD:
        return semantic_store[best_idx][1]   # return cached answer
    return None

def store_semantic(question: str, answer: str):
    vec = embedder.encode(question, normalize_embeddings=True)
    semantic_store.append((vec, answer))


# ── Layer 3: Context Cache (handled by Anthropic provider) ────────────
SYSTEM_PROMPT_LARGE = "You are a helpful support agent. " + ("Rules: " * 300)

def call_llm_with_context_cache(question: str) -> str:
    """Calls LLM with provider-side context caching for the system prompt."""
    response = client.messages.create(
        model="claude-opus-4-5",
        max_tokens=256,
        system=[
            {
                "type": "text",
                "text": SYSTEM_PROMPT_LARGE,
                "cache_control": {"type": "ephemeral"}  # provider caches this
            }
        ],
        messages=[{"role": "user", "content": question}]
    )
    return response.content[0].text


# ── Full Layered Cache Pipeline ───────────────────────────────────────
def ask(user_question: str) -> str:
    """
    Passes the question through all 4 cache layers.
    Only reaches the LLM if no layer has a stored answer.
    """
    print(f"\n📨 Question: '{user_question}'")

    # Layer 1: Exact match
    cached_answer, exact_key = check_exact(user_question)
    if cached_answer:
        print("  ⚡ Layer 1 HIT — Exact match cache")
        return cached_answer

    # Layer 2: Semantic similarity
    sem_answer = check_semantic(user_question)
    if sem_answer:
        print("  🧠 Layer 2 HIT — Semantic cache")
        return sem_answer

    # Layers 3 + 4: Call LLM with context caching (provider handles Layer 3)
    print("  🔄 Cache miss on all layers — calling LLM with context cache...")
    answer = call_llm_with_context_cache(user_question)

    # Store in both exact and semantic caches for future use
    store_exact(exact_key, answer)
    store_semantic(user_question, answer)

    print(f"  ✅ Answer stored in both exact + semantic caches")
    return answer


# --- Test the full stack ---
test_questions = [
    "What is your refund policy?",               # first time — goes to LLM
    "What is your refund policy?",               # exact hit on Layer 1
    "How do I get a refund on my purchase?",     # semantic hit on Layer 2
    "Can I return an item for my money back?",   # semantic hit on Layer 2
    "Where is my order right now?",              # different topic — goes to LLM
]

for q in test_questions:
    result = ask(q)
    print(f"  Answer: {result[:60]}...")

Output:

📨 Question: 'What is your refund policy?'
  🔄 Cache miss on all layers — calling LLM with context cache...
  ✅ Answer stored in both exact + semantic caches
  Answer: We offer a 30-day full refund with no questions asked...

📨 Question: 'What is your refund policy?'
  ⚡ Layer 1 HIT — Exact match cache
  Answer: We offer a 30-day full refund with no questions asked...

📨 Question: 'How do I get a refund on my purchase?'
  🧠 Layer 2 HIT — Semantic cache
  Answer: We offer a 30-day full refund with no questions asked...

📨 Question: 'Can I return an item for my money back?'
  🧠 Layer 2 HIT — Semantic cache
  Answer: We offer a 30-day full refund with no questions asked...

📨 Question: 'Where is my order right now?'
  🔄 Cache miss on all layers — calling LLM with context cache...
  ✅ Answer stored in both exact + semantic caches
  Answer: You can track your order at orders.shop.com using...

4 out of 5 questions were served without a full LLM call — and the one that did reach the LLM used provider-side context caching to cut its input costs by 90%. This is the full production caching stack in action! 🏆

10. Cache Invalidation — When to Clear Your Cache 🗑️

Caching has one golden rule in software engineering: "There are only two hard problems in computer science — cache invalidation and naming things." Cache invalidation means: knowing when to throw away old cached answers.

If your return policy changes from 30 days to 60 days, your cache still has the old "30 days" answer. Every user asking about returns gets the wrong information! You need a strategy to expire or refresh cached answers when things change.

The three main invalidation strategies:

  • TTL (Time-To-Live) → Every cache entry expires automatically after a set time. "Answers are valid for 1 hour." Simple and safe — but answers might be stale for up to the full TTL after a change.
  • Event-Driven Invalidation → When something in your system changes (policy update, new product launch), explicitly clear the relevant cache entries. More complex but always accurate.
  • Version-Based Invalidation → Tag every cache entry with a version number. When your knowledge base updates, bump the version — all old-version entries are automatically ignored.
🎯 What this code block will do:
This shows version-based cache invalidation — the cleanest strategy for LLMOps. Every cached answer is tagged with a "knowledge base version" number. When you update your knowledge base (new products, new policies), you bump the version number and ALL old cache entries become invalid instantly — no need to find and delete them one by one!
import hashlib
import time

class VersionedCache:
    """
    A cache that automatically invalidates all entries when
    the knowledge base version changes.

    Think of version number like a library catalogue edition number.
    When a new edition is released, all references to the old edition
    become outdated immediately — without removing individual pages.
    """

    def __init__(self):
        self._store = {}
        self.current_version = "v1.0"   # tracks the current knowledge base version

    def _make_key(self, question: str) -> str:
        """
        The key includes the version number.
        If the version changes, all old keys stop matching — instant invalidation!
        """
        base = f"{self.current_version}:{question.strip().lower()}"
        return hashlib.sha256(base.encode()).hexdigest()

    def get(self, question: str):
        """Retrieve a cached answer — only valid if version matches."""
        key = self._make_key(question)
        entry = self._store.get(key)
        if entry:
            print(f"✅ Cache HIT (version={self.current_version})")
            return entry["answer"]
        print(f"❌ Cache MISS (version={self.current_version})")
        return None

    def set(self, question: str, answer: str):
        """Store an answer with the current version tag."""
        key = self._make_key(question)
        self._store[key] = {
            "answer": answer,
            "version": self.current_version,
            "stored_at": time.time()
        }

    def invalidate_all(self, new_version: str):
        """
        Bumping the version effectively invalidates ALL old entries.
        Old keys (which include the old version in their hash) will never match again.
        No need to delete entries manually — they just stop being found!
        """
        old = self.current_version
        self.current_version = new_version
        print(f"🔄 Knowledge base updated: {old} → {new_version}")
        print("   All previous cache entries are now effectively invalidated.")


# --- Test version invalidation ---
cache = VersionedCache()

# Store an answer under v1.0
cache.set("What is your return policy?", "Returns allowed within 30 days.")

print("=== Checking cache (v1.0) ===")
print(cache.get("What is your return policy?"))

# Policy changes! Update to 60 days and bump the version
print("\n=== Policy changed — bumping version ===")
cache.invalidate_all("v2.0")
cache.set("What is your return policy?", "Returns allowed within 60 days.")

print("\n=== Checking cache (v2.0) ===")
print(cache.get("What is your return policy?"))
# Now serves the NEW answer automatically!

Output:

=== Checking cache (v1.0) ===
✅ Cache HIT (version=v1.0)
Returns allowed within 30 days.

=== Policy changed — bumping version ===
🔄 Knowledge base updated: v1.0 → v2.0
   All previous cache entries are now effectively invalidated.

=== Checking cache (v2.0) ===
❌ Cache MISS (version=v2.0)
[LLM called with new policy...]
✅ Cache HIT (version=v2.0)
Returns allowed within 60 days.

One version bump — all old answers invalidated instantly, everywhere. No hunting for stale entries. No risk of wrong answers reaching users. ✨

11. Monitoring Your Cache — Are You Actually Saving Money? 📈

A cache you cannot measure is a cache you cannot improve. Always instrument your cache with metrics so you know exactly how much money and time it is saving you.

The three key cache metrics to track:

  • Cache Hit Rate → Percentage of requests served from cache. A healthy production cache should be 50–80%+ hit rate.
  • Cost Savings → How many LLM tokens were saved because of cache hits. Translate this into dollars saved per day.
  • Latency Reduction → Average response time for cache hits vs cache misses. Cache hits should be 10–100x faster than real LLM calls.
🎯 What this code block will do:
This adds a built-in statistics tracker to any cache. Every hit and miss is counted. At the end, you call print_stats() and it shows you exactly how many requests were served from cache, what percentage that is, and how much money you saved — in dollars!
import time

class CacheMonitor:
    """
    Tracks how well your cache is performing.
    Attach this to any caching layer to measure its impact.
    """

    def __init__(self, cost_per_llm_call: float = 0.01):
        self.hits = 0
        self.misses = 0
        self.total_hit_latency = 0.0
        self.total_miss_latency = 0.0
        self.cost_per_call = cost_per_llm_call   # in USD

    def record_hit(self, latency_seconds: float):
        self.hits += 1
        self.total_hit_latency += latency_seconds

    def record_miss(self, latency_seconds: float):
        self.misses += 1
        self.total_miss_latency += latency_seconds

    def print_stats(self):
        total = self.hits + self.misses
        if total == 0:
            print("No requests recorded yet.")
            return

        hit_rate = self.hits / total * 100
        money_saved = self.hits * self.cost_per_call
        avg_hit_ms  = (self.total_hit_latency / self.hits * 1000) if self.hits else 0
        avg_miss_ms = (self.total_miss_latency / self.misses * 1000) if self.misses else 0

        print("\n" + "=" * 45)
        print("       CACHE PERFORMANCE REPORT")
        print("=" * 45)
        print(f"  Total requests     : {total:,}")
        print(f"  Cache hits         : {self.hits:,} ({hit_rate:.1f}%)")
        print(f"  Cache misses       : {self.misses:,} ({100 - hit_rate:.1f}%)")
        print(f"  Money saved        : ${money_saved:.2f} USD")
        print(f"  Avg hit latency    : {avg_hit_ms:.1f}ms")
        print(f"  Avg miss latency   : {avg_miss_ms:.1f}ms")
        print(f"  Speed improvement  : {avg_miss_ms / avg_hit_ms:.0f}x faster")
        print("=" * 45)


# --- Simulate a session with cache hits and misses ---
monitor = CacheMonitor(cost_per_llm_call=0.01)

# Simulate cache hits (fast — served from memory)
for _ in range(650):
    monitor.record_hit(latency_seconds=0.003)    # 3ms average

# Simulate cache misses (slow — real LLM calls)
for _ in range(350):
    monitor.record_miss(latency_seconds=1.8)     # 1.8s average

monitor.print_stats()

Output:

=============================================
       CACHE PERFORMANCE REPORT
=============================================
  Total requests     : 1,000
  Cache hits         : 650 (65.0%)
  Cache misses       : 350 (35.0%)
  Money saved        : $6.50 USD
  Avg hit latency    : 3.0ms
  Avg miss latency   : 1,800.0ms
  Speed improvement  : 600x faster
=============================================

65% cache hit rate, $6.50 saved per 1,000 requests, and cache hits are 600x faster. At 100,000 requests per day, that is $650 saved daily! 💰

12. Common Mistakes to Avoid ⚠️

❌ Mistake 1 — Caching Personalised Responses:
Never cache answers that are specific to one user — like "Your order #12345 ships tomorrow." Another user could receive that cached answer and see someone else's order details. Only cache answers that are true for all users equally.
❌ Mistake 2 — Setting Semantic Threshold Too Low:
A threshold of 0.70 might match "What is the return policy?" with "What is your shipping policy?" — completely different topics! Start at 0.88 and only lower it if you have evidence that it improves results.
❌ Mistake 3 — Never Invalidating the Cache:
Your product prices, policies, and features change over time. A cache with no invalidation strategy will serve wrong answers indefinitely. Always set a TTL or implement version-based invalidation from day one.
❌ Mistake 4 — Using In-Memory Cache in Production:
A Python dictionary works fine for learning and local testing. But if your server restarts, the entire cache is lost — and you lose all savings. In production, always use Redis or another persistent cache store.
❌ Mistake 5 — Not Monitoring Cache Performance:
If your cache hit rate is only 10%, your caching strategy may be misconfigured. If it is 99%, your threshold might be too loose and you risk serving wrong answers. Always track hit rate, latency, and cost savings in production.
✅ The Golden Caching Checklist:
→ Never cache personalised or user-specific responses
→ Use semantic threshold of 0.88+ for production systems
→ Always set TTL on all cache entries (3600 seconds is a safe default)
→ Use Redis (not Python dicts) for production caches
→ Enable provider-side context caching for prompts longer than 1,000 tokens
→ Implement version-based invalidation so cache clears on knowledge base updates
→ Monitor hit rate, latency, and cost savings — aim for 50–70%+ hit rate

13. Caching Tools and Libraries 🛠️

Tool / Library Type Best For Free?
GPTCache Semantic cache library Drop-in semantic cache for OpenAI + LangChain ✅ Open source
Redis In-memory store Fast TTL-based caching in any language ✅ Open source
LangCache Semantic cache service Managed semantic caching with dashboards ✅ Free tier
FAISS Vector search library Fast similarity search for semantic caches ✅ Open source (Meta)
Qdrant Vector database Production-grade vector search + filtering ✅ Open source + cloud
Anthropic Context Cache Provider-side cache Caching large system prompts with Claude ✅ Built into API
OpenAI Prompt Cache Provider-side cache Automatic caching for GPT prompts >1024 tokens ✅ Built into API
Redis AI / RedisVL Vector + cache combined Semantic caching using Redis as the vector store ✅ Open source
💡 Recommended Starter Stack :
→ Development: GPTCache + FAISS + SQLite — free, easy, runs locally
→ Production: Redis + FAISS (or Qdrant) + Anthropic/OpenAI context caching
→ Managed/No-code: LangCache — handles everything, just connect your LLM client

14. Your Learning Roadmap — Step by Step 🗺️

Week 1 — Understand the Basics 🐣

  • Re-read the ice cream stall analogy until caching feels natural
  • Code the exact match cache from Section 3 — test it with your own questions
  • Observe cache hits and misses in the terminal output

Week 2 — Semantic Caching 🐥

  • Install sentence-transformers: pip install sentence-transformers
  • Build the semantic cache from Section 4 and test with varied phrasings
  • Experiment with different similarity thresholds (0.80, 0.88, 0.95) to see the difference

Week 3 — Context Caching 🦅

  • Use the Anthropic context caching example from Section 5 on a real large prompt
  • Check the cache_read_input_tokens in the usage object to confirm savings
  • Compare cost before and after enabling context caching on your system prompt

Week 4 — Production Stack 🏆

  • Install Redis locally: docker run -p 6379:6379 redis
  • Wire up the production semantic cache from Section 7
  • Add the CacheMonitor from Section 11 and track your hit rate

Month 2+ — Expert Level 🚀

  • Try GPTCache with LangChain for a managed semantic caching experience
  • Implement version-based cache invalidation for your knowledge base
  • Build the full 4-layer caching stack from Section 9 for maximum cost savings

Quick Summary 📝

  • Caching → Store answers the first time, serve instantly for free after that
  • Exact Match Cache → Match word-for-word — fastest but misses different phrasing
  • Semantic Cache → Match by meaning using embeddings — catches "How do I return?" AND "What is your return policy?" as the same question
  • Context Caching → Cache large system prompts at the provider level — saves up to 90% on input tokens with cache_control
  • Agent Step Caching → Cache intermediate results in multi-step workflows — skip expensive steps when inputs haven't changed
  • Similarity Threshold → Use 0.88+ for production — lower risks wrong cached answers
  • Cache Invalidation → Use TTL (time expiry) or version-based invalidation — never let stale answers reach users
  • Cache Monitoring → Track hit rate (aim for 50–70%+), latency, and cost savings
  • Production Stack → Redis + FAISS + provider context cache = industry standard

Smart caching is the single biggest cost-saving move any LLMOps engineer can make. Start with context caching this week — it takes 3 lines of code and saves money immediately. Then add semantic caching as your second superpower. Happy caching! 🐼✨

Comments