Skip to main content

Caching Strategies in System Design: Cache-Aside, Read-Through, Write-Through & Write-Behind

Calculating read time…

Your database query takes 200ms. You have 10,000 users hitting the same endpoint per second. That's 2,000,000 ms of database time per second — your database crumbles. Add a cache in front and that same query takes 1ms from Redis. Same 10,000 users, same endpoint, 200x less database load.

Caching is the single most impactful performance optimisation in all of software engineering. But it is also one of the most subtle — implement the wrong strategy and you get stale data, data loss, thundering herd failures, or a cache that doesn't actually help.

Today we will master every caching strategy — from the simplest pattern to the subtle art of preventing cache stampedes that can bring down production systems at peak load. 


💡 Caching Powers Everything You Use

🔴 Redis: The world's most popular cache — used by Twitter, GitHub, Airbnb, Snapchat
🟣 Memcached: Facebook's original cache — served billions of requests per second
🌐 CDN Cache: Cloudflare, Akamai — cache web pages and media at the edge
🐘 PostgreSQL: Has its own shared buffer cache internally
💻 Your Browser: Caches HTTP responses, images, scripts locally
📱 Mobile Apps: Cache API responses offline for better UX

"There are only two hard things in Computer Science: cache invalidation and naming things." — Phil Karlton. After this post, cache invalidation won't be hard anymore. 

🧊 Section 1: Think of Caching Like Your Kitchen Fridge

Before diving into strategies, let's nail the mental model with a perfect analogy.

Imagine you love a specific dish. To make it, you need a particular spice. The spice shop is 2 km away (your database). Your kitchen fridge is right next to you (your cache).

You keep a small supply of that spice in the fridge. Need it? Grab it from the fridge instantly (cache HIT, ~1ms). Ran out? Walk to the spice shop, buy some, bring extra home and store in the fridge (cache MISS, ~100ms).

🧊 Kitchen Analogy 💻 Engineering Reality ⚡ Performance
Fridge (next to you) Cache (Redis / Memcached) ~1ms
Spice shop (2km away) Database (PostgreSQL / MySQL) ~100ms
Found spice in fridge! Cache HIT → return instantly ⚡ 100x faster
Fridge empty, go to shop Cache MISS → query database 🐌 Slow (one time)
Stale food in fridge Stale data (TTL expired or not updated) ⚠️ Correctness issue
Everyone ran out of the spice at the same time Cache key expires → 10,000 requests hit DB simultaneously 💥 Cache Stampede!

⚡ Animated: Cache HIT vs Cache MISS — Feel the Speed Difference

📱 Request
⚡
🟢 Cache HIT!
~1ms
← No DB needed! 🎉
📱 Request
❌
🔴 Cache MISS
...waiting...
🐌
🗄️ DB Query
~100ms

↑ Cache HIT = 100x faster. A 95%+ cache hit rate means 95% of requests never touch the database. That's how Instagram serves billions of requests daily.


📦 Section 2: Cache-Aside (Lazy Loading) — The Most Popular Pattern

Cache-Aside is the most widely used caching pattern. You'll find it in almost every production system — from Instagram's feed service to Twitter's user profile cache.

💡 The "On-Demand Grocery" Analogy

Your fridge is empty when you move in. You only stock it as you need things. Want milk? Check fridge → empty → go buy milk → put some in the fridge for next time.

You only store what you actually need. Items you never cook with never waste fridge space. This is Cache-Aside (Lazy Loading) — the cache is populated on demand, only with data that was actually requested.

📦 Cache-Aside — Complete Flow (Both HIT and MISS)

✅ CACHE HIT Path

1️⃣ App checks Redis: GET user:42
2️⃣ Redis: DATA FOUND ✅
3️⃣ Return data directly to caller
🚀 ~1ms. DB never contacted!

⚠️ CACHE MISS Path

1️⃣ App checks Redis: GET user:42
2️⃣ Redis: NOT FOUND ❌
3️⃣ App queries DB: SELECT * FROM users WHERE id=42
4️⃣ App writes result to Redis: SET user:42 {data} EX 300
5️⃣ Return data to caller
🐌 ~100ms (but next request = HIT!)
✅ Advantages
Cache only stores requested data (no waste)
Resilient to cache failure (just slower)
Works with any cache store (Redis, Memcached)
App has full control over caching logic
❌ Disadvantages
Cold start: first request always misses
Risk of stale data (TTL must be chosen carefully)
Cache Stampede risk on key expiry
App code handles cache logic (complexity)
✅ Who Uses Cache-Aside?

Instagram: User profiles, post metadata, follower counts — all Cache-Aside with Redis.
Twitter: User timelines and tweet metadata — Cache-Aside with a 5-minute TTL.
GitHub: Repository metadata — Cache-Aside with Memcached.

Best for: Read-heavy workloads where reads vastly outnumber writes. User profiles, product catalogues, article content, configuration data.

📖 Section 3: Read-Through — The Smart Librarian Pattern

Read-Through looks very similar to Cache-Aside but with one key difference: the cache itself is responsible for loading data from the database — not the application code.

💡 The Smart Librarian Analogy

You ask a librarian for a book. If it's on their desk (cache), they hand it to you immediately. If not, they go to the storeroom (database), get the book, put a copy on their desk for next time, and hand it to you — all without you doing anything.

In Cache-Aside: YOU go to the storeroom when the librarian doesn't have it.
In Read-Through: The librarian does it all — you always ask the librarian only.

📖 Read-Through vs Cache-Aside — Key Difference

📦 Cache-Aside — App controls everything
App → Cache (miss) → App → DB → App stores in cache → App returns
⚠️ App knows about both cache AND database
📖 Read-Through — Cache controls DB fetch
App → Cache (miss) → Cache → DB → Cache stores → Cache returns to App
✅ App only knows about the cache layer

📖 Read-Through Request Flow

📱 App calls
cache.get("user:42")
→
📚 Cache Layer
Not in cache!
I'll get it...
→
🗄️ Database
SELECT * FROM users
WHERE id=42
→
📚 Cache stores
result + returns to App

↑ The application ONLY interacts with the cache — never directly with the database. The cache library handles DB fetching internally (e.g., AWS ElastiCache with DAX for DynamoDB).

✅ Best for:
Apps using caching-native services like DynamoDB with DAX, or NCache. Simpler application code since DB logic lives in the cache layer.
❌ Downsides:
Cache must support "loader" functions. Less flexible. Still suffers from cold start and cache stampede problems.

✍️ Section 4: Write-Through — Always Keep Cache and DB in Sync

Write-Through takes a different approach: instead of only caring about reads, it intercepts writes too. Every write goes to the cache and the database at the same time.

💡 The Twin Notebooks Analogy

Imagine you always write everything in two notebooks simultaneously: your pocket notebook (cache) and your main journal (database). Any time you write, both get updated at the same instant.

This means your pocket notebook is always perfectly up-to-date with your main journal. Reading from the pocket notebook (cache) is always safe — no stale data! ✅

✍️ Write-Through — Every Write Goes to Both Cache AND Database

📱 App writes
user profile update
→
📚 Cache Layer
Step 1: Update cache
SET user:42 {new data}
↘️
↗️
🗄️ Database
Step 2: Write to DB
UPDATE users SET...
✅ Cache always consistent with DB
Reads are always fresh — no stale data problems. Perfect for data that must be accurate (user settings, account balances, permissions).
❌ Every write has DB latency
Writes are slower (must wait for DB confirmation). Cache can fill with data that is never read (wrote it, no one ever reads that cache key).
💡 The "Write Amplification" Problem in Write-Through

Imagine a batch job that updates 10,000 user records. With Write-Through, every single update hits both the cache AND the database. Most of those cached values are never read before they expire. You wasted cache memory and database write operations on data that was never served from cache.

Solution: Combine Write-Through with a TTL. Only cache data you're likely to read soon. Or use Write-Through only for frequently-read data, and skip caching on bulk operations.

⏱️ Section 5: Write-Behind (Write-Back) — Fast Writes, Async Persistence

Write-Behind is the most performance-optimised write strategy. Writes go to the cache immediately — the database is updated later, asynchronously. This is how your computer's SSD writes work at the hardware level!

💡 The "Draft Auto-Save" Analogy

When you type in Google Docs, your words appear instantly (written to memory/cache). Google saves to their servers every few seconds in the background (async database write). You don't wait for the server save to see your words appear. It feels instant because the write-behind buffer is doing the heavy lifting.

But if your computer crashes before the background save — you lose the last few seconds of typing. That's the risk of Write-Behind: potential data loss if the cache fails before DB is updated.

⏱️ Write-Behind — Instant Cache Write, Async Database Write

📱 App writes data
→
📚 Cache updated INSTANTLY
→
✅ ACK to App (<1ms!)
Meanwhile, in background:
⏳ Write queue accumulates changes
→
🗄️ Batched DB write (every 5s or when buffer full)
⏱️ Timeline: Write-Behind vs Write-Through for 1000 writes/second
Write-Through:
1000 DB writes/sec 😰 HIGH LOAD
Write-Behind:
1 batched DB write/5sec 😊 LOW LOAD

↑ Write-Behind batches 5000 cache writes into 1 DB batch write. 5000x reduction in DB load!

🚫 The Critical Risk: Data Loss on Cache Failure

In Write-Behind, the database is NOT immediately updated. If Redis crashes or is restarted before the async write completes, those buffered writes are LOST FOREVER. Your database is now out of sync with what users saw in the application.

Mitigations:
🔹 Use Redis AOF (Append-Only File) persistence so the write queue survives restarts
🔹 Use Kafka or SQS as the write queue — far more durable than Redis
🔹 Keep Write-Behind for non-critical data only (view counts, analytics, likes)
🔹 Never use Write-Behind for financial transactions or user authentication data

📊 Section 6: All Four Strategies Side by Side

Property 📦 Cache-Aside 📖 Read-Through ✍️ Write-Through ⏱️ Write-Behind
Who loads cache? Application code Cache library Application + Cache Cache (async)
Read latency Fast (on HIT) Fast (on HIT) Fast (always cached) Fast (always cached)
Write latency N/A (reads only) N/A (reads only) Slow (wait for DB) Instant (async DB)
Stale data risk ⚠️ Yes (TTL based) ⚠️ Yes (TTL based) ✅ No (always fresh) ✅ No (cache is source of truth)
Data loss risk ✅ No ✅ No ✅ No ❌ Yes (if cache dies)
Best for Read-heavy. General purpose. Simplifying app code. DAX, NCache. Write + read. Critical data. Write-heavy. Analytics. Counters.
Examples Instagram, Twitter, GitHub DynamoDB + DAX User settings, auth tokens Like counts, view counts

🐂 Section 7: Cache Stampede — The Most Dangerous Cache Bug

You've implemented caching perfectly. Your cache hit rate is 99%. Then, at 8 PM on a Friday, a popular cache key expires. Every user who was viewing that page simultaneously — maybe 50,000 people — sends a request. All 50,000 requests miss the cache at the same instant. All 50,000 hit your database. Your database falls over. Your site goes down. That's a Cache Stampede.

💡 The Concert Ticket Analogy

Imagine concert tickets go on sale at exactly 10:00 AM. Millions of fans are watching the clock. The moment it hits 10:00 — millions of requests slam the ticketing server simultaneously. The server crashes. Nobody gets tickets.

Cache Stampede is the same thing but triggered by a cache key expiring at an exact timestamp. Everyone cached the page with TTL = 300 seconds. They all got the cached version at 8:00 PM. At exactly 8:05 PM — every single one of those keys expires simultaneously. Everyone rushes to the database at the same moment.

🐂 Cache Stampede — The Chain of Disaster

⏱️ Cache key for "/homepage" was set with TTL = 300s. Cached at 8:00:00 PM.
← TTL counting down (300s) → EXPIRED at 8:05:00 PM
💥 8:05:00 PM exactly — 50,000 users see the expired key. ALL get cache MISS simultaneously!
🐂 Req 1
🐂 Req 2
🐂 Req 3
... 49,997 more requests ...
→ ALL HIT DATABASE! 💀
🗄️ Database: 50,000 concurrent queries for same data. Overwhelmed. DOWN. 💀

🛡️ Section 8: Cache Stampede Prevention — Four Battle-Tested Solutions

🔐 Strategy 1: Mutex / Distributed Lock

💡 The Single Chef Analogy

When the fridge is empty and 50 people want milk, instead of all 50 people running to the shop (stampede!), they use a sign-up sheet. The first person puts their name on the sheet, goes to the shop, and brings back enough milk for everyone. Everyone else waits for that one person to return. Only ONE trip to the shop. This is a Mutex / Distributed Lock.

🔐 Mutex Strategy — Only One Request Hits the DB

🐂 50,000 requests
→
Cache MISS! Try to acquire lock...
Request #1 — WINS the lock 🔐
Queries DB → gets result → writes to cache → releases lock ✅
Requests #2–50,000 — Wait for lock ⏳
Poll for lock release → lock released → read from cache → return ✅
🎉 Result: 1 DB query instead of 50,000. Stampede eliminated!
⚠️ Risk: Waiting requests add latency. If the DB query is slow, all 50,000 users wait. Always set a lock timeout (e.g., 5 seconds) so you don't block forever on a DB failure.

🎲 Strategy 2: Probabilistic Early Expiration (XFetch / PER Algorithm)

🎲 Probabilistic Early Expiration — The Brilliant Randomised Approach

Instead of waiting for the key to actually expire, requests occasionally "voluntarily" refresh the cache early — but only with a small probability that increases as the TTL approaches zero. Requests that are near the expiry have a higher chance of proactively fetching fresh data. This distributes the refresh over time, preventing simultaneous expiry!

The XFetch Formula:
refresh = (-beta × delta × log(random())) > remaining_ttl

Where: beta = tuning constant (e.g., 1.0), delta = time to recompute the cached value, remaining_ttl = seconds left before key expires.

As TTL → 0, this expression increasingly becomes True for random requests. Some requests refresh early. By the time TTL actually hits 0, the key is often already refreshed!

🌊 Strategy 3: Stale-While-Revalidate

When a key expires, return the stale (old) value immediately and trigger a background refresh. Users get an instant response (stale data), and the cache is refreshed in the background. Next request gets fresh data.

🌊 Stale-While-Revalidate — Serve Stale, Refresh Asynchronously

📱 Request arrives
after TTL expired
→
→
🔄 Background thread fetches fresh data from DB
↻
→
✅ Next request gets fresh data from cache!

HTTP Cache-Control: stale-while-revalidate=30 does exactly this in browsers and CDNs! Serve stale for up to 30 seconds while revalidating in the background. This is also the basis of Next.js ISR (Incremental Static Regeneration).

⏰ Strategy 4: Background Refresh (Proactive Cache Warming)

Instead of waiting for keys to expire, proactively refresh them before they do. A background job runs every N seconds and updates frequently-accessed cache keys before any user request would see a miss.

✅ Proactive Refresh Pattern
Set TTL = 300s on cache key
Background job runs every 240s (before expiry!)
Fetches fresh data → writes to cache → resets TTL
Cache never actually expires! Zero misses! 🎉
⚠️ When to Use
Only works for known "hot" keys
Can't proactively refresh millions of keys
Best for: homepage data, trending feeds, global configs

🗑️ Section 09: Cache Invalidation — The Hardest Problem in Computer Science

Phil Karlton famously said: "There are only two hard things in Computer Science: cache invalidation and naming things." Cache invalidation — deciding when and how to remove stale data from the cache — is genuinely difficult.

⏱️ 1. TTL (Time-To-Live) — Simplest Approach

Set an expiration time on every cache entry. After TTL expires, the key is automatically deleted. No need to track when data changes. Simple. But potentially serves stale data for up to TTL seconds. Best for: Data that changes infrequently or where eventual consistency is acceptable. User avatars (1 hour), product prices (5 minutes), news articles (30 seconds).

🎯 2. Event-Driven Invalidation — Most Accurate

When data changes in the database, immediately delete or update the corresponding cache key. No stale data ever. Requires coordination between write path and cache. Pattern: UPDATE users SET... → DELETE CACHE user:42. Use change data capture (CDC) from the WAL (write-ahead log) to trigger cache invalidation — this is what Facebook uses at massive scale.

🏷️ 3. Tag-Based Invalidation — Most Flexible

Tag cache entries with related entity IDs. When an entity changes, invalidate all tagged keys. user:profile:42 tagged with ["user:42", "account:99"]. When user 42 changes, invalidate all keys tagged with user:42 in one operation. Redis supports this via sets: maintain a set of keys per tag, then delete all on change. Used by Symfony, Laravel, and Drupal caching systems.


🗺️ Section 10: Everything Together — Complete Caching Architecture

⚡ Multi-Layer Caching Architecture — How Production Systems Cache

── CLIENT ──
📱 Browser / Mobile App
L1 Cache: HTTP Cache-Control headers (browser cache)
⬇️
🌐 CDN Layer (Cloudflare / Akamai)
L2 Cache: Static assets, HTML pages, API responses with s-maxage
⬇️
🛡️ Reverse Proxy Cache (NGINX, Varnish)
L3 Cache: Full-page cache, API response cache
⬇️
🖥️ Application Tier (Microservices)
📦 Cache-Aside (user profiles, products)
✍️ Write-Through (auth sessions)
⏱️ Write-Behind (like counts, views)
⬇️
🔴 Redis / Memcached — L4 Cache
In-memory cache. ~1ms. The workhorse of application caching.
Stampede prevention: Mutex Lock, Stale-While-Revalidate, XFetch
⬇️ only on cache miss
🗄️ Database — L5 (Source of Truth)
PostgreSQL / MySQL. Has its own internal buffer cache. ~100ms.

📐 Section 11: Core Caching Design Principles

📏 Cache What Is Expensive to Compute, Not Everything

Caching a value that takes 1ms to compute isn't worth the complexity. Cache values that take 100ms+ (complex DB queries, external API calls, report generation). Also consider caching computed aggregates: "total orders for user 42" is cheaper to cache than to recompute from thousands of rows every time.

⚠️ Always Plan for Cache Failure

What happens when Redis goes down? Your system must still work — just slower. Every cache read should have a fallback to the database. Every cache write failure should be logged but not crash the request. Cache is a performance optimisation, not a dependency. Design it as optional.

🎯 Choose TTL Wisely — It's a Business Decision

TTL = trade-off between freshness and performance. Ask: "How stale can this data be?" User avatar: 1 hour (stale avatar is fine). Product price: 5 minutes (stale price causes complaints). Bank balance: 0 seconds (never cache). News article: 30 seconds. The right TTL is decided by business requirements, not by engineers alone.

🐂 Always Design for Cache Stampede

Any popular cache key is a stampede waiting to happen. Use stale-while-revalidate for most cases. Use mutex for critical data. Add jitter to TTLs: instead of exactly 300 seconds, use 300 + random(0, 30). This spreads expiration across a 30-second window, preventing simultaneous expiry.

📊 Measure Your Cache Hit Rate — Aim for 90%+

A cache hit rate below 80% means your caching strategy isn't working — you're adding complexity without enough performance benefit. Monitor cache hits, misses, evictions, and memory usage. Redis INFO STATS gives you this data. Most production caches target 95%+ hit rate. If hit rate is low, investigate: are TTLs too short? Are keys poorly designed?


🎓 Section 12: Cheat Sheet

❓ "When would you use Cache-Aside vs Read-Through?"

Cache-Aside: when the application needs full control over caching logic, or when different code paths need different caching strategies. Most common choice. Read-Through: when using a cache-native service (DynamoDB + DAX, Redis with CacheLookup plugin) that handles DB fetching internally. Simplifies app code but less flexible. In practice: Cache-Aside is used ~90% of the time because most apps use Redis directly.

❓ "What is a Cache Stampede and how do you prevent it?"

Cache Stampede (thundering herd): popular cache key expires simultaneously for thousands of users. All requests miss cache, all hit DB at once, DB is overwhelmed. Prevention: (1) Mutex/lock — only 1 request fetches from DB, others wait. (2) Stale-while-revalidate — serve stale, refresh in background. (3) Probabilistic early expiration (XFetch) — proactively refresh before expiry. (4) TTL jitter — add random(0, 30s) to TTL to spread expiration.

❓ "When would you use Write-Behind vs Write-Through?"

Write-Through: when data consistency between cache and DB is critical, and write latency is acceptable. Use for: user settings, authentication sessions, permissions. Write-Behind: when write speed is critical and some data loss is acceptable. Use for: view counts, like counts, analytics events, session activity. NEVER use Write-Behind for: financial transactions, user accounts, inventory quantities.

❓ "How do you handle cache invalidation?"

Three approaches: (1) TTL — simple, eventual consistency, some staleness. (2) Event-driven — on any write, immediately delete the cache key. Fresh but requires coordination. (3) Change Data Capture (CDC) — read DB WAL events and invalidate cache asynchronously. Used by Facebook at scale. Most robust. In practice: use TTL for most data + event-driven invalidation for critical paths. Add TTL jitter to prevent stampede on batch expiry.

❓ "Design a caching strategy for a product catalogue with 10M products"

Use Cache-Aside with Redis. Key = "product:{id}". TTL = 10 minutes (prices may update). Cache only requested products (can't pre-load 10M). On write: invalidate specific key. For popular products (home page, trending): use background refresh with no TTL. For product search results: cache query hash → result list, TTL = 60s. Stampede prevention: stale-while-revalidate for popular products, TTL jitter for rest. Cache hit rate target: 95%+ (most users view the same popular products).

❓ "What is cache eviction and what policies exist?"

Cache eviction: when cache is full, old entries must be removed to make room for new ones. Policies: LRU (Least Recently Used — evict the entry not accessed for the longest time — most common), LFU (Least Frequently Used — evict the entry accessed fewest times), FIFO (evict oldest by insertion time — simple but not always optimal), Random (evict random entry — simple), TTL-based (always evict expired entries first). Redis default: LRU with configurable maxmemory-policy. Choose LFU for skewed access patterns.


🎉 Final Summary 

📦 Cache-Aside: App checks cache → miss → app queries DB → app writes to cache. Most flexible. App controls everything. Risk: cold start, stampede.
📖 Read-Through: App only talks to cache. Cache auto-loads from DB on miss. Simpler app code. Requires cache-native system (DAX, NCache).
✍️ Write-Through: Every write goes to cache AND DB simultaneously. Always consistent. Higher write latency. Best for critical, frequently-read data.
⏱️ Write-Behind: Write to cache instantly. DB updated asynchronously (batched). Fastest writes. Risk: data loss if cache dies before DB write. For counters, analytics, likes.
🐂 Cache Stampede: Popular key expires → thousands hit DB simultaneously → DB dies. Caused by synchronised TTL expiry.
🔐 Mutex Prevention: Redis SET NX (atomic lock). Only 1 request fetches from DB. Others wait. Eliminates stampede but adds latency.
🌊 Stale-While-Revalidate: Return stale data instantly. Background thread fetches fresh. Zero wait. Best for most applications. HTTP standard header too.
🎲 XFetch / Probabilistic: Random requests refresh cache before expiry. Probability increases as TTL → 0. Distributes refresh naturally without locks.
🏷️ Cache Invalidation: TTL (simple, some staleness), Event-driven (delete on write, always fresh), CDC from WAL (most robust, used by Facebook). Add TTL jitter to prevent batch expiry.
📊 Cache Hit Rate: Target 95%+. Monitor HIT, MISS, evictions. A cache with 80% hit rate may not be worth the complexity. Measure everything.
✅ The Most Important Lesson from Caching:

Caching is not one pattern — it is a toolkit of complementary strategies, each solving a different dimension of the performance-vs-consistency trade-off.

Read-heavy workloads → Cache-Aside or Read-Through.
Consistency critical → Write-Through.
Write-heavy, tolerates loss → Write-Behind.
Popular endpoints → Stampede prevention (always!).
Data changes → Invalidation strategy.

The master system designer chooses the right combination for each specific use case — not the same strategy for everything. That nuanced judgment is the mark of expertise.


Happy Learning! Keep Building! 🔥

Comments