Skip to main content

Caching in System Design: Cache Hits, TTL, Redis and Performance Explained

Calculating read time…

Your AI fraud detection system is brilliant. Accurate. Reliable. But it takes 800 milliseconds to answer each request — because every single time, it goes all the way to the database to fetch the customer profile.

What if the same customer makes 50 transactions in one day? That is 50 trips to the database. 50 × 800ms = 40 seconds of wasted database time. The customer profile did not even change!

Caching is the solution. And by the end of this post, you will understand it so deeply that you will never build a slow system again. ⚡

💡 Did You Know?
Netflix serves 250 million users worldwide. If every video request hit their database, it would collapse instantly. Their entire streaming platform runs on aggressive, multi-layer caching. Caching is not an optimisation — it is the foundation of every fast system.


🧠 What is Caching?

Let me tell you about your brain. 🧠 Every morning you wake up and you know your home address. You know your mum's phone number. You know 2 + 2 = 4.

You did NOT look these up in a book this morning. They are stored in your short-term memory — fast, instantly accessible. If you needed to look everything up in a book every single time — life would be impossibly slow.

A computer cache is exactly this — a fast, nearby copy of frequently needed data. Instead of going to the slow database every time, the system checks the cache first. If the data is there — it returns in 0.1 milliseconds instead of 800ms. That is 8,000 times faster. 


✅ Simple Definition:
Caching = Storing a copy of frequently-needed data in a fast, nearby location so you never have to fetch it from the slow original source again.

Cache Hit = Data is in the cache. Answer returned instantly. ⚡
Cache Miss = Data is NOT in the cache. Must go to the database. 🐢

🏦 Our Real-World Example: AnomalyAI Fraud Detection

AnomalyAI checks every bank transaction for fraud in real-time. For every transaction, it needs the customer's behaviour profile: average spend, usual merchants, home country, recent transaction velocity.

Without caching: 500,000 transactions/minute × 800ms DB lookup = catastrophe. 💥
With caching: 500,000 transactions/minute × 0.1ms cache hit = totally fine. ✅

  ANOMALYAI — CACHING IMPACT:

  WITHOUT CACHE:
  ┌──────────────────────────────────────────────────────────────┐
  │  500,000 transactions/min                                    │
  │  Each needs a customer profile from DB                       │
  │  DB can handle: ~5,000 queries/min maximum                   │
  │  Result: DB overwhelmed at 10× capacity 💥 CRASH!            │
  └──────────────────────────────────────────────────────────────┘

  WITH CACHE:
  ┌──────────────────────────────────────────────────────────────┐
  │  500,000 transactions/min                                    │
  │  Cache hit rate: 95% (repeat customers)                      │
  │  475,000 → served from Cache in 0.1ms ⚡                     │
  │   25,000 → first-time today, goes to DB                      │
  │  DB load: 25,000/min = very manageable ✅                     │
  │  Response time: 0.1ms average (was 800ms)                    │
  │  Speed improvement: 8,000× faster! 🚀                        │
  └──────────────────────────────────────────────────────────────┘

⚡ Cache Hit vs Cache Miss — The Full Journey

⚡ Cache HIT vs Cache MISS — Watch the Journey! ✅ SCENARIO A: Cache HIT (data already in cache!) 🏦 AnomalyAI "Get profile customer #42" Check cache ⚡ CACHE customer #42 ✅ Found! → HIT! Returns in 0.1ms ⚡ ← Profile data 🎉 Database NOT touched! Zero DB load. Pure speed. ⚡ ⚠️ SCENARIO B: Cache MISS (data NOT in cache — first time!) 🏦 AnomalyAI "Get profile customer #99" (new) ⚡ CACHE customer #99 ❌ MISS — not found go to DB 🗄️ DATABASE Found customer #99 ✅ Returns in 800ms 🐢 ✨ Result saved to cache for next time! Next call for #99 → CACHE HIT! ⚡ Cache Miss is rare after the first call. Subsequent calls = instant ⚡

🏗️ Caching Strategies — HOW to Read and Write

There are different strategies for when and how to put data into the cache. Each has strengths and weaknesses. Choosing the wrong one can cause stale data or cache stampedes.

🥇 Strategy 1 — Cache-Aside (Lazy Loading)

This is the most popular strategy . The application is in charge of managing the cache itself. Think of it like a smart student who only writes on their sticky notes when they actually need to look something up. 📝


🥈 Strategy 2 — Write-Through

Every time data is written, it is simultaneously written to both the cache AND the database at the same time. The cache is always up to date. Zero stale data. ✅

  WRITE-THROUGH FLOW:

  App writes: "Update customer #42 risk score to HIGH"
         │
         ├──► ① Write to CACHE immediately
         │         (cache updated in 0.1ms)
         │
         └──► ② Write to DATABASE simultaneously
                   (persisted to disk in 800ms)

  ✅ Cache always fresh
  ✅ No stale reads possible
  ⚠️  Every write is slightly slower (must wait for both)
  ⚠️  Cache fills with rarely-read data (written but never read again)

  ANOMALYAI USE: When an anomaly is flagged, write the alert
                 to BOTH cache and DB simultaneously.
                 Dashboard reads it from cache instantly. ⚡

🥉 Strategy 3 — Write-Behind (Write-Back)

The application writes to the cache only. The cache asynchronously writes to the database later — in batches. This is the fastest write strategy, but has a small risk. ⚠️

  WRITE-BEHIND FLOW:

  App writes: "Update transaction count: +1"
         │
         └──► ① Write to CACHE only (done! ⚡ 0.1ms)
                   (App continues — does NOT wait for DB)

  Meanwhile, asynchronously:
         Cache ──► ② Batch writes to DB every 1 second

  ✅ Fastest write performance
  ✅ Batching reduces DB write pressure by 100×
  ⚠️  If cache crashes BEFORE batch write → data lost!

  ANOMALYAI USE: Transaction counters (velocity tracking).
                 Losing 1 second of count data is acceptable.
                 But never use this for fraud alerts! 🚨
💡 Which Strategy Does AnomalyAI Use?
  • 📖 Customer profiles (reads) → Cache-Aside (lazy load, 5-min TTL)
  • 🚨 Fraud alerts (critical writes) → Write-Through (always persisted immediately)
  • 📊 Transaction counters (velocity) → Write-Behind (batch, performance first)
  • 🤖 AI prediction results (repeat queries) → Cache-Aside with Semantic Cache

⏳ Cache Eviction — What Gets Thrown Out and Why

A cache has limited memory. It cannot hold everything forever. When the cache is full and new data needs to be stored, it must decide what to throw out to make room. This is called Eviction.

Think of your desk. 🖥️ It only has space for 10 folders. When an 11th folder arrives, you must decide which of the 10 to put back in the filing cabinet. Which one do you choose? That decision is your eviction policy.


💡 AnomalyAI Uses Which Eviction Policy?
  • 👤 Customer profiles → TTL of 5 minutes (profiles may change — never serve data older than 5 min)
  • 🧠 AI prediction results → LRU (keep recently active customers' predictions in cache)
  • 📊 Transaction counters → TTL of 60 seconds (velocity window is 60 seconds — old counts are useless)

🧠 Semantic Caching — The AI Superpower

Traditional caching is exact match only. If you ask "What is the fraud risk for transaction #9923?" and then ask "What is the risk score for txn 9923?" — a traditional cache returns a miss even though it is the same question! 😤

Semantic Caching uses AI embeddings to cache by meaning, not exact text. Two questions that mean the same thing get the same cached answer. This is a game-changer for LLM-powered applications. 🤖


📝 What the code below does (in simple words):
This is the Semantic Cache for AnomalyAI's LLM-powered fraud explanation feature. When the AI gives a fraud explanation, it saves the question AND the answer. Next time anyone asks a similar question (even with different words), it finds it using vector similarity and returns the cached answer — no LLM call needed! 🧠⚡
# ── FILE: semantic_cache.py ────────────────────────────────────
# PURPOSE: A Semantic Cache for AnomalyAI's LLM fraud explanations.
#          Stores answers by MEANING, not exact text.
#          Similar questions get the same cached answer instantly.
#          This saves LLM API calls and makes responses 20000× faster.
# ───────────────────────────────────────────────────────────────

import oracledb
import numpy as np
import array
import hashlib
from typing import Optional

class SemanticCache:
    """
    AI-powered cache that matches questions by meaning, not exact text.
    Uses OCI Generative AI embeddings + Oracle 23ai Vector Search.
    """

    def __init__(self, similarity_threshold: float = 0.92):
        """
        similarity_threshold = 0.92 means:
        "If two questions are 92% similar in meaning → return the cached answer"
        Higher value = stricter matching (less cache hits but more accurate)
        Lower value = looser matching (more cache hits but may match wrong answers)
        """
        self.threshold = similarity_threshold
        self.conn = oracledb.connect(
            user     = "ANOMALYAI_USER",
            password = "SecurePass123!",
            dsn      = "anomalyai_cache_high"
        )
        self._create_table()
        print(f"✅ SemanticCache ready (threshold: {similarity_threshold})")

    def _create_table(self):
        """Create the cache table with VECTOR column in Oracle 23ai."""
        cursor = self.conn.cursor()
        cursor.execute("""
            CREATE TABLE IF NOT EXISTS semantic_cache (
                id              NUMBER GENERATED ALWAYS AS IDENTITY,
                question_text   CLOB,                        -- Original question stored for reference
                question_vector VECTOR(1536, FLOAT32),       -- The meaning-vector of the question
                answer_text     CLOB,                        -- The LLM's answer (cached)
                hit_count       NUMBER DEFAULT 0,            -- How many times this was served from cache
                created_at      TIMESTAMP DEFAULT SYSTIMESTAMP,
                expires_at      TIMESTAMP DEFAULT SYSTIMESTAMP + INTERVAL '1' HOUR
            )
        """)
        self.conn.commit()
        cursor.close()

    def _embed(self, text: str) -> list:
        """
        Convert a question into a meaning-vector using OCI Generative AI.
        This is what makes the cache 'semantic' — it understands meaning!
        """
        import oci
        from oci.generative_ai_inference import GenerativeAiInferenceClient
        from oci.generative_ai_inference.models import EmbedTextDetails, OnDemandServingMode

        config = oci.config.from_file("~/.oci/config")
        client = GenerativeAiInferenceClient(config)
        req    = EmbedTextDetails(
            inputs         = [text],
            serving_mode   = OnDemandServingMode(model_id="cohere.embed-multilingual-v3"),
            compartment_id = "ocid1.compartment.oc1..xxxxx",
            input_type     = "SEARCH_QUERY"
        )
        response = client.embed_text(req)
        return response.data.embeddings[0]

    def get(self, question: str) -> Optional[str]:
        """
        Look up the cache by meaning.
        Returns the cached answer if a similar question was asked before.
        Returns None if no similar question is found (cache miss).
        """
        # Step 1: Convert this question to a meaning-vector
        question_vector = self._embed(question)
        q_array = array.array("f", question_vector)

        cursor = self.conn.cursor()

        # Step 2: Search for similar questions using VECTOR_DISTANCE
        # COSINE distance: 0.0 = identical meaning, 2.0 = completely opposite meaning
        # We want distance < (1 - threshold), so for 0.92 → distance < 0.08
        cursor.execute("""
            SELECT question_text,
                   answer_text,
                   id,
                   VECTOR_DISTANCE(question_vector, :qv, COSINE) AS dist
            FROM   semantic_cache
            WHERE  expires_at > SYSTIMESTAMP              -- Not expired
              AND  VECTOR_DISTANCE(question_vector, :qv, COSINE) < :max_dist
            ORDER  BY dist ASC                            -- Closest match first
            FETCH  FIRST 1 ROW ONLY
        """, qv=q_array, max_dist=(1 - self.threshold))

        row = cursor.fetchone()

        if row:
            _, answer, cache_id, distance = row
            similarity = 1 - distance

            # Increment the hit counter
            cursor.execute("UPDATE semantic_cache SET hit_count = hit_count + 1 WHERE id = :1",
                           [cache_id])
            self.conn.commit()

            print(f"⚡ Semantic Cache HIT! Similarity: {similarity:.1%} | Saved LLM call!")
            cursor.close()
            return answer   # Return the cached answer instantly!

        cursor.close()
        print(f"❌ Cache MISS — will call LLM and cache the result.")
        return None         # Cache miss — caller must get answer from LLM

    def set(self, question: str, answer: str, ttl_minutes: int = 60):
        """
        Save a new question-answer pair to the semantic cache.
        Called after every LLM response so future similar questions are cached.
        """
        question_vector = self._embed(question)
        q_array = array.array("f", question_vector)

        cursor = self.conn.cursor()
        cursor.execute("""
            INSERT INTO semantic_cache (question_text, question_vector, answer_text, expires_at)
            VALUES (:1, :2, :3, SYSTIMESTAMP + NUMTODSINTERVAL(:4, 'MINUTE'))
        """, [question, q_array, answer, ttl_minutes])
        self.conn.commit()
        cursor.close()
        print(f"✨ Cached! Future similar questions will be answered in 0.1ms.")


# ── HOW ANOMALYAI USES THE SEMANTIC CACHE ──────────────────────

semantic_cache = SemanticCache(similarity_threshold=0.92)

def explain_fraud(transaction_id: str, user_question: str) -> str:
    """
    Explain a fraud detection result in plain English.
    Uses Semantic Cache so similar questions don't hit the expensive LLM.
    """
    # Step 1: Check semantic cache first
    cached_answer = semantic_cache.get(user_question)
    if cached_answer:
        return cached_answer   # Returned in 0.1ms! LLM not called. 💰

    # Step 2: Cache miss — call the LLM (expensive: ~2000ms, ~$0.002)
    # ... (call OCI GenAI LLM here)
    llm_answer = call_oci_genai_llm(user_question, transaction_id)

    # Step 3: Save to semantic cache for next time
    semantic_cache.set(user_question, llm_answer, ttl_minutes=60)

    return llm_answer


# TEST IT:
q1 = "What is the fraud risk for transaction #9923?"
q2 = "Risk score of transaction 9923?"          # Different words, same meaning
q3 = "Is txn number 9923 suspicious?"           # Even more different

ans1 = explain_fraud("TXN-9923", q1)   # Cache MISS → calls LLM
ans2 = explain_fraud("TXN-9923", q2)   # Cache HIT ⚡ (similarity: 97%)
ans3 = explain_fraud("TXN-9923", q3)   # Cache HIT ⚡ (similarity: 93%)

🌍 Cache Layers — The Full Stack

In production, AnomalyAI uses multiple layers of caching working together like a relay race. The fastest layer is checked first.

  ANOMALYAI CACHE LAYERS (Fastest to Slowest):

  ┌──────────────────────────────────────────────────────────────────────────┐
  │                                                                          │
  │  Layer 1: IN-PROCESS MEMORY CACHE (Python dictionary)                   │
  │  ──────────────────────────────────────────────────────                 │
  │  Speed: 0.001ms (nanoseconds!)                                          │
  │  Size: ~100 MB (limited by RAM of one server)                           │
  │  Content: Model weights, config, the last 1000 profiles                 │
  │  Eviction: LRU, cleared when server restarts                            │
  │                                                                          │
  │  Layer 2: REDIS / OCI CACHE WITH REDIS                                  │
  │  ──────────────────────────────────────────────────────                 │
  │  Speed: 0.1ms                                                           │
  │  Size: ~50 GB (shared across all 50 servers)                            │
  │  Content: Customer profiles, transaction history, AI predictions         │
  │  Eviction: LRU + TTL (5 min for profiles, 1 hour for predictions)       │
  │                                                                          │
  │  Layer 3: SEMANTIC CACHE (Oracle 23ai Vector Search)                    │
  │  ──────────────────────────────────────────────────────                 │
  │  Speed: 5ms (vector similarity search)                                  │
  │  Size: Unlimited (disk-backed)                                          │
  │  Content: LLM fraud explanations, similar question answers              │
  │  Eviction: TTL (60 minutes)                                             │
  │                                                                          │
  │  Layer 4: DATABASE (Oracle 23ai)                                        │
  │  ──────────────────────────────────────────────────────                 │
  │  Speed: 800ms                                                           │
  │  Size: Unlimited                                                        │
  │  Content: Everything (source of truth)                                  │
  │  Eviction: Never (permanent)                                            │
  │                                                                          │
  └──────────────────────────────────────────────────────────────────────────┘

  REQUEST FLOW:
  Check L1 → HIT? Return 0.001ms ⚡
  Check L2 → HIT? Return 0.1ms ⚡
  Check L3 → HIT? Return 5ms ⚡
  None hit → Go to DB, populate L1+L2+L3 → Return 800ms 🐢
📝 What the code below does (in simple words):
This is the multi-layer cache manager for AnomalyAI. It checks Layer 1 (Python dict) first, then Layer 2 (Redis), then the database. On a miss, it populates ALL layers so future requests are fast. Think of it like checking your pocket, then your bag, then going to the shop! 🛍️
# ── FILE: multi_layer_cache.py ─────────────────────────────────
# PURPOSE: A multi-layer cache manager for AnomalyAI.
#          Checks the fastest cache first. If missed, checks the
#          next layer. On DB fetch, populates all cache layers
#          so the next request is much faster. 🚀
# ───────────────────────────────────────────────────────────────

import redis
import json
import time
from functools import lru_cache
from typing import Optional, Any

# ── LAYER 2: Redis connection (shared across all servers) ───────
redis_client = redis.Redis(
    host     = "anomalyai-cache.cache.oci.oraclecloud.com",
    port     = 6380,
    password = "RedisSecurePass!",
    ssl      = True,
    decode_responses = True
)

# ── LAYER 1: In-process LRU cache (per server, ultra fast) ─────
# @lru_cache turns a function into a cached function automatically!
# maxsize=1000 means: keep at most 1000 results in memory
@lru_cache(maxsize=1000)
def _local_cache_get(key: str) -> Optional[str]:
    """Layer 1: Python's built-in in-memory LRU cache."""
    return None   # This is overridden by the actual cached value


class MultiLayerCache:
    """
    Checks 3 cache layers in order of speed.
    Populates all layers on DB fetch for future speed.
    """

    def __init__(self):
        self._local = {}          # Layer 1: Simple dict in memory
        self._local_ttl = {}      # Track expiry for local dict
        print("✅ Multi-layer cache ready: Memory → Redis → Database")

    def get(self, key: str) -> Optional[Any]:
        """
        Try to get a value from the fastest available cache layer.
        """
        # ── LAYER 1: Check local memory (0.001ms) ───────────────
        if key in self._local:
            if time.time() < self._local_ttl.get(key, 0):
                print(f"⚡⚡ L1 HIT (memory): {key}")
                return self._local[key]
            else:
                del self._local[key]          # Expired, clean up

        # ── LAYER 2: Check Redis (0.1ms) ─────────────────────────
        try:
            redis_val = redis_client.get(key)
            if redis_val:
                data = json.loads(redis_val)
                # Promote to L1 for next time (5-min local TTL)
                self._set_local(key, data, ttl_seconds=300)
                print(f"⚡ L2 HIT (Redis): {key}")
                return data
        except redis.RedisError as e:
            print(f"⚠️ Redis error (falling through to DB): {e}")

        # ── All cache layers missed ─────────────────────────────
        print(f"❌ Cache MISS — going to database: {key}")
        return None

    def set(self, key: str, value: Any,
            l1_ttl: int = 300,     # 5 minutes in local memory
            l2_ttl: int = 1800):   # 30 minutes in Redis
        """
        Save a value to ALL cache layers simultaneously.
        Called after every database fetch.
        """
        serialised = json.dumps(value)

        # ── Save to Layer 1 (local memory) ──────────────────────
        self._set_local(key, value, l1_ttl)

        # ── Save to Layer 2 (Redis) ──────────────────────────────
        try:
            redis_client.setex(key, l2_ttl, serialised)
        except redis.RedisError as e:
            print(f"⚠️ Could not write to Redis: {e}")   # Non-fatal

        print(f"✨ Cached in all layers: {key}")

    def _set_local(self, key: str, value: Any, ttl_seconds: int):
        """Save to local in-memory dict with expiry tracking."""
        self._local[key] = value
        self._local_ttl[key] = time.time() + ttl_seconds

    def invalidate(self, key: str):
        """
        Remove a key from ALL cache layers.
        Called when the underlying data changes in the database.
        (This is the famous 'cache invalidation' problem!)
        """
        self._local.pop(key, None)
        self._local_ttl.pop(key, None)
        try:
            redis_client.delete(key)
        except redis.RedisError:
            pass
        print(f"🗑️ Invalidated from all layers: {key}")


# ── HOW ANOMALYAI USES IT ───────────────────────────────────────

cache = MultiLayerCache()

def get_customer_profile(customer_id: str) -> dict:
    """
    Get customer profile with multi-layer caching.
    First call: ~800ms (DB). Every call after: ~0.001ms (L1 memory).
    """
    cache_key = f"profile:{customer_id}"

    # Step 1: Check all cache layers
    cached = cache.get(cache_key)
    if cached:
        return cached   # Instant return from cache! ⚡

    # Step 2: Cache miss — fetch from database (~800ms)
    profile = fetch_from_database(customer_id)

    # Step 3: Populate ALL cache layers for next time
    cache.set(
        key    = cache_key,
        value  = profile,
        l1_ttl = 300,   # Keep in local memory for 5 minutes
        l2_ttl = 1800   # Keep in Redis for 30 minutes
    )

    return profile

💣 Cache Invalidation — "The Hardest Problem in Computer Science"

Phil Karlton famously said: "There are only two hard things in Computer Science: cache invalidation and naming things."

Cache Invalidation means: when the real data changes in the database, you must update or delete the cached copy too. If you forget — users see stale (outdated) data. 📅

  ANOMALYAI CACHE INVALIDATION SCENARIOS:

  SCENARIO 1: Customer updates their home address
  ────────────────────────────────────────────────
  1. Customer changes address in the app
  2. New address saved to database ✅
  3. But cache still has OLD address! ❌
  4. AnomalyAI uses old address for geo-checking → wrong result!
  SOLUTION: After DB update → call cache.invalidate("profile:42")

  SCENARIO 2: Fraud model is retrained with new patterns
  ────────────────────────────────────────────────────────
  1. New model version deployed
  2. Old predictions in Semantic Cache are now potentially wrong!
  3. SOLUTION: Flush the entire Semantic Cache on model deployment

  SCENARIO 3: Customer's risk level changes
  ──────────────────────────────────────────
  1. Compliance team manually upgrades customer to HIGH RISK
  2. Cache still shows NORMAL risk
  3. Fraud checks using cached risk level → dangerous!
  4. SOLUTION: Publish a cache invalidation event via OCI Streaming

  INVALIDATION STRATEGIES:
  ┌─────────────────────────┬──────────────────────────────────────────┐
  │  Strategy               │  How                                     │
  ├─────────────────────────┼──────────────────────────────────────────┤
  │  TTL (Time-to-Live)     │  Data expires after N seconds. Simple.   │
  │  Event-based            │  DB change → message → cache deleted     │
  │  Write-Through          │  Every write updates cache AND DB        │
  │  Manual invalidation    │  Code explicitly deletes cache key       │
  └─────────────────────────┴──────────────────────────────────────────┘

⚠️ The Cache Stampede Problem

Imagine 10,000 users all request the same data at the exact same moment. The cache has just expired (TTL). All 10,000 requests miss the cache simultaneously and ALL of them go to the database at once. The database crashes. 💥

This is a Cache Stampede (also called Thundering Herd). It is one of the most dangerous production failures in high-traffic systems.

📝 What the code below does (in simple words):
This prevents cache stampedes using a mutex lock. When 10,000 requests arrive and the cache is empty, only the first request is allowed to go to the database. The other 9,999 wait briefly, then read from the cache that was just populated. Like a single door that only one person can open at a time. 🚪
# ── FILE: stampede_protection.py ───────────────────────────────
# PURPOSE: Prevent Cache Stampede (Thundering Herd).
#          When cache expires and 10,000 requests arrive at once,
#          only ONE goes to the database. The others wait and then
#          get the result from cache. Saves the database from crash!
# ───────────────────────────────────────────────────────────────

import redis
import time
import json

redis_client = redis.Redis(host="anomalyai-cache.oci.oraclecloud.com",
                           port=6380, ssl=True, decode_responses=True)

def get_with_stampede_protection(
    key: str,
    fetch_from_db,          # A function that gets data from database
    ttl_seconds: int = 300,
    lock_timeout: int = 10  # Max seconds to hold the lock
) -> dict:
    """
    Safe cache get that prevents stampedes.
    Only 1 request fetches from DB when cache is empty.
    All others wait, then read from the freshly populated cache.
    """

    # Step 1: Try to get from cache (happy path — most requests end here)
    cached = redis_client.get(key)
    if cached:
        return json.loads(cached)   # Cache hit! Instant return ⚡

    # Step 2: Cache miss! Try to acquire a distributed lock
    # Only ONE server/request can hold this lock at a time
    lock_key     = f"lock:{key}"
    lock_acquired = redis_client.set(
        lock_key, "1",
        nx  = True,            # Only set if NOT EXISTS (atomic operation)
        ex  = lock_timeout     # Lock auto-expires after 10 seconds
    )

    if lock_acquired:
        # WE got the lock! We are the chosen one to fetch from DB 🏆
        try:
            print(f"🔒 Lock acquired! Fetching from DB: {key}")
            data = fetch_from_db()   # Go to database (one time only!)

            # Save to cache so others can use it
            redis_client.setex(key, ttl_seconds, json.dumps(data))
            print(f"✅ Data cached! Releasing lock.")
            return data

        finally:
            redis_client.delete(lock_key)   # Always release the lock!

    else:
        # Someone else has the lock — they are already fetching from DB
        # We wait briefly and try reading from cache
        print(f"⏳ Another request is fetching... waiting briefly.")
        for attempt in range(20):                 # Wait up to 2 seconds
            time.sleep(0.1)
            cached = redis_client.get(key)
            if cached:
                print(f"⚡ Got it from cache after {(attempt+1)*100}ms wait!")
                return json.loads(cached)

        # Still not in cache after waiting (unusual) — go to DB as fallback
        print(f"⚠️ Fallback: Going to DB directly.")
        return fetch_from_db()


# ── USAGE IN ANOMALYAI ──────────────────────────────────────────
def get_fraud_model_config() -> dict:
    """
    Get the fraud model configuration.
    This is read by EVERY transaction check — millions of times per day.
    Perfect candidate for stampede protection!
    """
    return get_with_stampede_protection(
        key          = "anomalyai:model:config:v2",
        fetch_from_db= lambda: fetch_model_config_from_db(),
        ttl_seconds  = 300    # Cache for 5 minutes
    )

⚠️ Common Caching Mistakes

🚫 DON'T #1 — Cache everything blindly.
Not all data benefits from caching. Only cache data that is read often, changes rarely, and is expensive to compute or fetch. Caching data that changes every second wastes memory and causes stale data nightmares. 🗑️
🚫 DON'T #2 — Set TTL to infinity (or forget TTL entirely).
Without expiry, stale data lives forever in your cache. Always set a TTL appropriate to how quickly your data changes. Customer profiles: 5 min. Currency exchange rates: 30 sec. Country list: 24 hours. ⏱️
🚫 DON'T #3 — Use the same TTL for all cache keys.
Different data has different freshness requirements. A fraud alert must be seconds-fresh. A country name list can be hours old. Design TTL values specifically for each type of data. 🎯
🚫 DON'T #4 — Cache sensitive data without encryption.
Redis cache contents can be read by anyone with network access. Always encrypt sensitive data (PII, financial details) before caching. Or better: cache only non-sensitive identifiers and fetch sensitive fields separately. 🔐
✅ DO #1 — Monitor your Cache Hit Rate continuously. A healthy cache should have a hit rate above 90%. Below 80% means your TTL is too short, your cache is too small, or your data access patterns are too random. 📊
✅ DO #2 — Design your cache keys carefully. Use a consistent, hierarchical naming scheme: service:entity:id:version e.g. anomalyai:customer:42:profile. Good keys make debugging, invalidation, and monitoring 10× easier. 🔑
✅ DO #3 — Always design your system to work correctly even if the cache is empty. Never assume the cache is always available. Redis can crash, OCI Cache can restart, network can drop. The database is always the safety net. 🛡️

  • 🧠 What caching is — your computer's short-term memory
  • ⚡ Cache Hit vs Miss — 0.1ms vs 800ms, 8,000× difference
  • 🔄 Cache-Aside, Write-Through, Write-Behind — three strategies, three use cases
  • 📤 Eviction policies — LRU, LFU, TTL: which one and why
  • 🧠 Semantic Caching — AI superpower for LLM caching
  • 🏗️ Multi-layer caching — Memory → Redis → DB waterfall
  • 💣 Cache stampede — the hidden danger and how to prevent it
  • 💣 Cache invalidation — why it is the hardest problem, and how to solve it

happy caching! ⚡✨

Comments