Skip to main content

Back-of-the-Envelope Calculations for System Design Interviews: QPS, Storage and Capacity

Calculating read time…

A senior engineer walks into a system design interview. The interviewer says: "Design a fraud detection system that handles 500,000 transactions per minute."

Before writing a single line of architecture, the engineer grabs a napkin and starts scribbling numbers. 🖊️ "500K per minute... that's about 8,333 per second... each needs 10ms compute... so 83 CPU cores... round up to 100 servers..."

In under 5 minutes, they know exactly how big this system needs to be. No tools. No spreadsheet. Just smart, structured arithmetic. That is Back-of-the-Envelope Calculation — and it is one of the most valuable skills a systems architect can have. 🧮

💡 Did You Know?
Jeff Dean (Google's legendary engineer) is famous for being able to estimate the cost, performance, and scale of any system in his head within minutes. He published a list of "Numbers Every Engineer Should Know" that is still memorised by engineers worldwide . We cover it all in this post. 🧠


📝 What Is a Back-of-the-Envelope Calculation?

Imagine you are opening a pizza restaurant. 🍕 Before you sign a lease and buy 10 ovens, you want to know: "How busy will we actually be?"

You do not need a perfect business plan. You just think: "Our town has 50,000 people. Maybe 1% orders pizza on any given day = 500 pizzas. We can make 50 per hour, so we need 10 hours of cooking capacity. Two ovens. Four staff. Done."

That napkin estimate is good enough to make a confident decision. Back-of-the-envelope (BOTE) calculations do exactly the same for computer systems — quick, reasonably accurate estimates that guide big architectural choices.

✅ The Goal of BOTE:
Not to be perfectly accurate — but to be accurate enough to make good decisions. An estimate within 2× of reality is considered excellent. You are not building a spreadsheet — you are building intuition. 🎯

🔢 The Numbers Every Engineer Must Know Cold

Just like a pilot knows the aircraft's speed and fuel rate without calculating, a systems architect should know these numbers by heart. They are the building blocks of every BOTE calculation.

⚡ Latency Numbers — How Fast Is Fast?

⚡ Latency Numbers — Memorise These For Life! Speed bars race to the right — notice how HUGE the difference is between levels! L1 Cache (CPU) 1 ns ⚡⚡⚡ L2/L3 Cache 10 ns ⚡⚡ RAM (Memory) 100 ns ⚡ SSD Random Read 100 µs (0.1ms) 🟡 Same Datacenter 0.5 ms 🟠 HDD Seek 10 ms 🔴 Cross-Region (USA↔EU) 50 ms 🔴 USA ↔ Asia 150ms L1 Cache is 150,000× faster than cross-continental network! 🤯
  THE NUMBERS EVERY ENGINEER KNOWS BY HEART :

  ┌─────────────────────────────────────────────────────────────────────────┐
  │  LATENCY CHEAT SHEET                                                    │
  │  ─────────────────────────────────────────────────────────────────────  │
  │  L1 cache hit                    :     1 ns                             │
  │  L2 cache hit                    :     4 ns                             │
  │  RAM access                      :   100 ns = 0.0001 ms                 │
  │  SSD random read                 : 0.1 ms  = 100 µs                     │
  │  Network same datacenter         : 0.5 ms                               │
  │  HDD seek                        :  10 ms                               │
  │  Cross-region network (US↔EU)    :  50 ms                               │
  │  Intercontinental (US↔Asia)      : 150 ms                               │
  ├─────────────────────────────────────────────────────────────────────────┤
  │  THROUGHPUT CHEAT SHEET                                                  │
  │  ─────────────────────────────────────────────────────────────────────  │
  │  Typical web server (1 core)     : 1,000 req/sec                        │
  │  DB (SQL, indexed)               : 1,000 queries/sec                    │
  │  Redis / Cache                   : 100,000 ops/sec                      │
  │  SSD throughput                  : 500 MB/sec                            │
  │  RAM throughput                  : 20 GB/sec                             │
  │  Network (1 Gbps link)           : 125 MB/sec                            │
  ├─────────────────────────────────────────────────────────────────────────┤
  │  TIME CHEAT SHEET                                                        │
  │  ─────────────────────────────────────────────────────────────────────  │
  │  1 day                           : 86,400 sec ≈ 100,000 (10⁵)          │
  │  1 week                          : 604,800 sec ≈ 600,000                │
  │  1 month                         : 2,592,000 sec ≈ 2.5 million          │
  │  1 year                          : 31,536,000 sec ≈ 30 million (3×10⁷) │
  └─────────────────────────────────────────────────────────────────────────┘
💡 The Three Golden Approximations for BOTE:
  • 🔢 Round aggressively — 8,333 becomes 10,000. That is fine.
  • 📐 Work in powers of 10 — 1K, 10K, 100K, 1M, 10M, 1B
  • ✅ Favour round numbers — 100 servers, not 83. Always round up for safety.

🧮 The BOTE Framework — 5 Steps Every Time

Every great back-of-the-envelope calculation follows the same structured approach. Master these 5 steps and you can size any system confidently.

🧮 The 5-Step BOTE Framework ① TRAFFIC How many users/requests per second? DAU × actions ÷ 86,400 = QPS → ② STORAGE How much data per record? How long kept? records/day × size × years = TB / PB → ③ BANDWIDTH How much data in/out per sec? Network sizing QPS × record size = MB/s → ④ MEMORY How much cache do we need? Hot data set 20% data = 80% traffic (Pareto) = GB / TB → ⑤ SERVERS How many servers do we need? QPS ÷ server capacity ×3 for peak+ redundancy = N servers Steps appear in order — always follow this sequence for any system! 📋

🏦 Full Worked Example: AnomalyAI at Scale

Let us now work through a complete, real BOTE calculation for our AnomalyAI fraud detection system. An interviewer says: "Design the infrastructure for a bank processing 500,000 transactions per minute." Here is how a senior architect solves it — step by step, on the back of an envelope.

📊 Step 1 — Traffic Estimation

  TRAFFIC ESTIMATION — AnomalyAI

  Given:
  ─────
  • 500,000 transactions per minute (stated in the problem)
  • Peak factor: 3× normal (Black Friday, month-end)

  Calculate:
  ──────────
  Normal TPS   = 500,000 ÷ 60 = ~8,333 transactions/second
  Peak TPS     = 8,333 × 3 = ~25,000 TPS (design for this!)

  Round to nice numbers:
  Normal TPS ≈ 10,000 TPS
  Peak TPS   ≈ 25,000 TPS  ← Use this for sizing everything

  ✅ WRITE QPS: 25,000 (each transaction is written to DB)
  ✅ READ  QPS: 75,000 (each transaction reads 3 tables: customer, history, rules)

🗄️ Step 2 — Storage Estimation

  STORAGE ESTIMATION — AnomalyAI

  Per transaction record:
  ───────────────────────
  • Transaction ID:      16 bytes
  • Customer ID:          8 bytes
  • Amount:               8 bytes (float64)
  • Merchant info:      100 bytes
  • Timestamp:            8 bytes
  • Risk score:           4 bytes
  • Anomaly features:   200 bytes (JSON)
  • Audit metadata:     150 bytes
  ─────────────────────────────────
  Total per record:  ≈ 500 bytes → Round to 1 KB (always add buffer!)

  Daily storage:
  ──────────────
  Normal rate:  8,333 TPS × 86,400 sec/day = 720,000,000 records/day
  Round to:     700 million records/day

  Storage/day:  700M records × 1 KB = 700 GB/day

  Annual storage:
  ───────────────
  700 GB/day × 365 days = 255 TB/year
  Round to:   260 TB/year (raw transactions)

  With replication (3 copies):
  260 TB × 3 = 780 TB ≈ 1 PB/year

  With indexes and WAL logs (add 30%):
  1 PB × 1.3 = 1.3 PB/year

  ✅ ANSWER: Plan for ~1.5 PB/year of storage (with buffer)
             Use OCI Block Storage + Object Storage tiering

🌐 Step 3 — Bandwidth Estimation

  BANDWIDTH ESTIMATION — AnomalyAI

  INBOUND (data coming IN to the system):
  ────────────────────────────────────────
  Peak TPS:  25,000 transactions/sec
  Payload per transaction: ~0.5 KB (incoming JSON request)

  Inbound bandwidth = 25,000 × 0.5 KB = 12,500 KB/s = 12.5 MB/s
  Round to: 15 MB/s inbound bandwidth needed

  OUTBOUND (data going OUT of the system):
  ─────────────────────────────────────────
  Each fraud check returns: ~0.2 KB (risk score + reason)
  Notification events:      ~0.3 KB × 25,000 = 7,500 KB/s

  Total outbound: (25,000 × 0.2 KB) + 7,500 KB/s = 12,500 KB/s ≈ 13 MB/s
  Round to: 15 MB/s outbound

  TOTAL NETWORK: 15 + 15 = 30 MB/s = 240 Mbps

  ✅ ANSWER: Need at least 1 Gbps network links (240Mbps with good headroom)
             OCI FastConnect provides 1-10 Gbps dedicated circuits — perfect.

💾 Step 4 — Memory / Cache Estimation

  MEMORY / CACHE ESTIMATION — AnomalyAI

  Pareto Principle: 20% of customers make 80% of transactions
  ─────────────────────────────────────────────────────────────
  Total customers:          10 million
  Active customers cached:  10M × 20% = 2 million customers
  Customer profile size:    2 KB each

  Cache size needed = 2,000,000 × 2 KB = 4,000,000 KB = 4 GB

  Add:
  • Recent transaction history (top 2M customers × 1KB):  2 GB
  • AI model feature vectors (1M cached predictions × 1KB): 1 GB
  • Rules engine cache:                                      1 GB
  ─────────────────────────────────────────────────────────────
  Total cache needed: 4 + 2 + 1 + 1 = 8 GB
  With 2× safety buffer:                = 16 GB

  ✅ ANSWER: 16 GB Redis cluster (OCI Cache with Redis)
             With 2-node HA: 2 × 8 GB nodes
             Cache hit rate target: > 90%



📋 The Complete AnomalyAI BOTE Summary

  ANOMALYAI — COMPLETE BACK-OF-THE-ENVELOPE SUMMARY
  ════════════════════════════════════════════════════════════════════════════

  GIVEN:
  • Bank with 10M customers
  • 500,000 transactions/minute (normal)
  • 3× peak factor on Black Friday

  ┌─────────────────────────┬──────────────────────────────────────────────┐
  │  Dimension              │  Estimate                                    │
  ├─────────────────────────┼──────────────────────────────────────────────┤
  │  Normal TPS             │  8,333 ≈ 10,000 TPS                          │
  │  Peak TPS               │  25,000 TPS (design target)                  │
  │  Read QPS               │  75,000 QPS (3 reads per transaction)        │
  ├─────────────────────────┼──────────────────────────────────────────────┤
  │  Record size            │  ~1 KB per transaction                       │
  │  Records per day        │  700 million                                 │
  │  Storage per year       │  ~260 TB raw → 1.5 PB with replication       │
  ├─────────────────────────┼──────────────────────────────────────────────┤
  │  Inbound bandwidth      │  ~15 MB/s                                    │
  │  Outbound bandwidth     │  ~15 MB/s                                    │
  │  Network link needed    │  1 Gbps (plenty of headroom)                 │
  ├─────────────────────────┼──────────────────────────────────────────────┤
  │  Cache size             │  ~16 GB Redis (20% hot data)                 │
  │  Cache hit rate target  │  > 90%                                       │
  ├─────────────────────────┼──────────────────────────────────────────────┤
  │  API servers            │  50 × OCI E5.Flex (8 OCPU, 32 GB)           │
  │  Database servers       │  30 × OCI Oracle 23ai (HA + read replicas)  │
  │  Cache / Queue servers  │  20 × OCI Cache + OCI Streaming              │
  │  Total servers          │  ~100 instances                              │
  └─────────────────────────┴──────────────────────────────────────────────┘

  Monthly OCI cost estimate:
  100 servers × ~$500/month = ~$50,000/month
  Storage (1.5 PB) at OCI rates ≈ ~$15,000/month
  Total: ~$65,000/month → Round to $70,000/month for safety

  ════════════════════════════════════════════════════════════════════════════
  ✅ All done in under 10 minutes on a napkin!
  ✅ Accuracy: Within 2× of real-world numbers (acceptable for BOTE)

🎭 Common BOTE Mistakes to Avoid

🚫 DON'T #1 — Forget the peak multiplier.
Designing for average load is the #1 BOTE mistake. Always multiply normal load by 2–5× for peaks. Black Friday, month-end, viral moments — they will come. Size for them. 📈
🚫 DON'T #2 — Forget replication in storage estimates.
Every important database stores 3 copies (1 primary + 2 replicas). If you estimate 100 TB of data, provision 300 TB of storage. Single copies are a production disaster waiting to happen. 🗄️
🚫 DON'T #3 — Get paralysed by precision.
This is back-of-the-ENVELOPE, not a PhD thesis. If you spend 20 minutes getting an exact answer when a rough one was needed — you missed the entire point. Round aggressively and move on. 🎯
✅ DO #1 — Always state your assumptions out loud (in interviews) or in comments (in documents). "I'm assuming 1 KB per record" or "assuming 3× peak factor" shows architectural maturity. Assumptions drive everything. 📋
✅ DO #2 — Work in powers of 10 whenever possible. 1K, 10K, 100K, 1M, 10M, 1B. These round numbers are easy to multiply mentally and communicate clearly. 🔢
✅ DO #3 — Always add a 30–50% safety buffer to your final answer. Real systems always have hidden overhead: OS, monitoring agents, logging, indexing. If your calculation says 70 servers — provision 100. 🛡️
  • 📝 What BOTE is — structured approximation for big decisions
  • 🔢 The essential numbers — latency, storage units, throughput cheat sheet
  • 🧮 The 5-step framework — Traffic → Storage → Bandwidth → Memory → Servers
  • 🏦 Full AnomalyAI calculation — 25K TPS, 1.5PB/year, 100 servers

Happy estimating! 🧮✨

Comments