Back-of-the-Envelope Calculations for System Design Interviews: QPS, Storage and Capacity
A senior engineer walks into a system design interview. The interviewer says: "Design a fraud detection system that handles 500,000 transactions per minute."
Before writing a single line of architecture, the engineer grabs a napkin and starts scribbling numbers. 🖊️ "500K per minute... that's about 8,333 per second... each needs 10ms compute... so 83 CPU cores... round up to 100 servers..."
In under 5 minutes, they know exactly how big this system needs to be. No tools. No spreadsheet. Just smart, structured arithmetic. That is Back-of-the-Envelope Calculation — and it is one of the most valuable skills a systems architect can have. 🧮
Jeff Dean (Google's legendary engineer) is famous for being able to estimate the cost, performance, and scale of any system in his head within minutes. He published a list of "Numbers Every Engineer Should Know" that is still memorised by engineers worldwide . We cover it all in this post. 🧠
📝 What Is a Back-of-the-Envelope Calculation?
Imagine you are opening a pizza restaurant. 🍕 Before you sign a lease and buy 10 ovens, you want to know: "How busy will we actually be?"
You do not need a perfect business plan. You just think: "Our town has 50,000 people. Maybe 1% orders pizza on any given day = 500 pizzas. We can make 50 per hour, so we need 10 hours of cooking capacity. Two ovens. Four staff. Done."
That napkin estimate is good enough to make a confident decision. Back-of-the-envelope (BOTE) calculations do exactly the same for computer systems — quick, reasonably accurate estimates that guide big architectural choices.
Not to be perfectly accurate — but to be accurate enough to make good decisions. An estimate within 2× of reality is considered excellent. You are not building a spreadsheet — you are building intuition. 🎯
🔢 The Numbers Every Engineer Must Know Cold
Just like a pilot knows the aircraft's speed and fuel rate without calculating, a systems architect should know these numbers by heart. They are the building blocks of every BOTE calculation.
⚡ Latency Numbers — How Fast Is Fast?
THE NUMBERS EVERY ENGINEER KNOWS BY HEART : ┌─────────────────────────────────────────────────────────────────────────┐ │ LATENCY CHEAT SHEET │ │ ───────────────────────────────────────────────────────────────────── │ │ L1 cache hit : 1 ns │ │ L2 cache hit : 4 ns │ │ RAM access : 100 ns = 0.0001 ms │ │ SSD random read : 0.1 ms = 100 µs │ │ Network same datacenter : 0.5 ms │ │ HDD seek : 10 ms │ │ Cross-region network (US↔EU) : 50 ms │ │ Intercontinental (US↔Asia) : 150 ms │ ├─────────────────────────────────────────────────────────────────────────┤ │ THROUGHPUT CHEAT SHEET │ │ ───────────────────────────────────────────────────────────────────── │ │ Typical web server (1 core) : 1,000 req/sec │ │ DB (SQL, indexed) : 1,000 queries/sec │ │ Redis / Cache : 100,000 ops/sec │ │ SSD throughput : 500 MB/sec │ │ RAM throughput : 20 GB/sec │ │ Network (1 Gbps link) : 125 MB/sec │ ├─────────────────────────────────────────────────────────────────────────┤ │ TIME CHEAT SHEET │ │ ───────────────────────────────────────────────────────────────────── │ │ 1 day : 86,400 sec ≈ 100,000 (10⁵) │ │ 1 week : 604,800 sec ≈ 600,000 │ │ 1 month : 2,592,000 sec ≈ 2.5 million │ │ 1 year : 31,536,000 sec ≈ 30 million (3×10⁷) │ └─────────────────────────────────────────────────────────────────────────┘
- 🔢 Round aggressively — 8,333 becomes 10,000. That is fine.
- 📐 Work in powers of 10 — 1K, 10K, 100K, 1M, 10M, 1B
- ✅ Favour round numbers — 100 servers, not 83. Always round up for safety.
🧮 The BOTE Framework — 5 Steps Every Time
Every great back-of-the-envelope calculation follows the same structured approach. Master these 5 steps and you can size any system confidently.
🏦 Full Worked Example: AnomalyAI at Scale
Let us now work through a complete, real BOTE calculation for our AnomalyAI fraud detection system. An interviewer says: "Design the infrastructure for a bank processing 500,000 transactions per minute." Here is how a senior architect solves it — step by step, on the back of an envelope.
📊 Step 1 — Traffic Estimation
TRAFFIC ESTIMATION — AnomalyAI Given: ───── • 500,000 transactions per minute (stated in the problem) • Peak factor: 3× normal (Black Friday, month-end) Calculate: ────────── Normal TPS = 500,000 ÷ 60 = ~8,333 transactions/second Peak TPS = 8,333 × 3 = ~25,000 TPS (design for this!) Round to nice numbers: Normal TPS ≈ 10,000 TPS Peak TPS ≈ 25,000 TPS ← Use this for sizing everything ✅ WRITE QPS: 25,000 (each transaction is written to DB) ✅ READ QPS: 75,000 (each transaction reads 3 tables: customer, history, rules)
🗄️ Step 2 — Storage Estimation
STORAGE ESTIMATION — AnomalyAI
Per transaction record:
───────────────────────
• Transaction ID: 16 bytes
• Customer ID: 8 bytes
• Amount: 8 bytes (float64)
• Merchant info: 100 bytes
• Timestamp: 8 bytes
• Risk score: 4 bytes
• Anomaly features: 200 bytes (JSON)
• Audit metadata: 150 bytes
─────────────────────────────────
Total per record: ≈ 500 bytes → Round to 1 KB (always add buffer!)
Daily storage:
──────────────
Normal rate: 8,333 TPS × 86,400 sec/day = 720,000,000 records/day
Round to: 700 million records/day
Storage/day: 700M records × 1 KB = 700 GB/day
Annual storage:
───────────────
700 GB/day × 365 days = 255 TB/year
Round to: 260 TB/year (raw transactions)
With replication (3 copies):
260 TB × 3 = 780 TB ≈ 1 PB/year
With indexes and WAL logs (add 30%):
1 PB × 1.3 = 1.3 PB/year
✅ ANSWER: Plan for ~1.5 PB/year of storage (with buffer)
Use OCI Block Storage + Object Storage tiering
🌐 Step 3 — Bandwidth Estimation
BANDWIDTH ESTIMATION — AnomalyAI
INBOUND (data coming IN to the system):
────────────────────────────────────────
Peak TPS: 25,000 transactions/sec
Payload per transaction: ~0.5 KB (incoming JSON request)
Inbound bandwidth = 25,000 × 0.5 KB = 12,500 KB/s = 12.5 MB/s
Round to: 15 MB/s inbound bandwidth needed
OUTBOUND (data going OUT of the system):
─────────────────────────────────────────
Each fraud check returns: ~0.2 KB (risk score + reason)
Notification events: ~0.3 KB × 25,000 = 7,500 KB/s
Total outbound: (25,000 × 0.2 KB) + 7,500 KB/s = 12,500 KB/s ≈ 13 MB/s
Round to: 15 MB/s outbound
TOTAL NETWORK: 15 + 15 = 30 MB/s = 240 Mbps
✅ ANSWER: Need at least 1 Gbps network links (240Mbps with good headroom)
OCI FastConnect provides 1-10 Gbps dedicated circuits — perfect.
💾 Step 4 — Memory / Cache Estimation
MEMORY / CACHE ESTIMATION — AnomalyAI
Pareto Principle: 20% of customers make 80% of transactions
─────────────────────────────────────────────────────────────
Total customers: 10 million
Active customers cached: 10M × 20% = 2 million customers
Customer profile size: 2 KB each
Cache size needed = 2,000,000 × 2 KB = 4,000,000 KB = 4 GB
Add:
• Recent transaction history (top 2M customers × 1KB): 2 GB
• AI model feature vectors (1M cached predictions × 1KB): 1 GB
• Rules engine cache: 1 GB
─────────────────────────────────────────────────────────────
Total cache needed: 4 + 2 + 1 + 1 = 8 GB
With 2× safety buffer: = 16 GB
✅ ANSWER: 16 GB Redis cluster (OCI Cache with Redis)
With 2-node HA: 2 × 8 GB nodes
Cache hit rate target: > 90%
📋 The Complete AnomalyAI BOTE Summary
ANOMALYAI — COMPLETE BACK-OF-THE-ENVELOPE SUMMARY ════════════════════════════════════════════════════════════════════════════ GIVEN: • Bank with 10M customers • 500,000 transactions/minute (normal) • 3× peak factor on Black Friday ┌─────────────────────────┬──────────────────────────────────────────────┐ │ Dimension │ Estimate │ ├─────────────────────────┼──────────────────────────────────────────────┤ │ Normal TPS │ 8,333 ≈ 10,000 TPS │ │ Peak TPS │ 25,000 TPS (design target) │ │ Read QPS │ 75,000 QPS (3 reads per transaction) │ ├─────────────────────────┼──────────────────────────────────────────────┤ │ Record size │ ~1 KB per transaction │ │ Records per day │ 700 million │ │ Storage per year │ ~260 TB raw → 1.5 PB with replication │ ├─────────────────────────┼──────────────────────────────────────────────┤ │ Inbound bandwidth │ ~15 MB/s │ │ Outbound bandwidth │ ~15 MB/s │ │ Network link needed │ 1 Gbps (plenty of headroom) │ ├─────────────────────────┼──────────────────────────────────────────────┤ │ Cache size │ ~16 GB Redis (20% hot data) │ │ Cache hit rate target │ > 90% │ ├─────────────────────────┼──────────────────────────────────────────────┤ │ API servers │ 50 × OCI E5.Flex (8 OCPU, 32 GB) │ │ Database servers │ 30 × OCI Oracle 23ai (HA + read replicas) │ │ Cache / Queue servers │ 20 × OCI Cache + OCI Streaming │ │ Total servers │ ~100 instances │ └─────────────────────────┴──────────────────────────────────────────────┘ Monthly OCI cost estimate: 100 servers × ~$500/month = ~$50,000/month Storage (1.5 PB) at OCI rates ≈ ~$15,000/month Total: ~$65,000/month → Round to $70,000/month for safety ════════════════════════════════════════════════════════════════════════════ ✅ All done in under 10 minutes on a napkin! ✅ Accuracy: Within 2× of real-world numbers (acceptable for BOTE)
🎭 Common BOTE Mistakes to Avoid
Designing for average load is the #1 BOTE mistake. Always multiply normal load by 2–5× for peaks. Black Friday, month-end, viral moments — they will come. Size for them. 📈
Every important database stores 3 copies (1 primary + 2 replicas). If you estimate 100 TB of data, provision 300 TB of storage. Single copies are a production disaster waiting to happen. 🗄️
This is back-of-the-ENVELOPE, not a PhD thesis. If you spend 20 minutes getting an exact answer when a rough one was needed — you missed the entire point. Round aggressively and move on. 🎯
- 📝 What BOTE is — structured approximation for big decisions
- 🔢 The essential numbers — latency, storage units, throughput cheat sheet
- 🧮 The 5-step framework — Traffic → Storage → Bandwidth → Memory → Servers
- 🏦 Full AnomalyAI calculation — 25K TPS, 1.5PB/year, 100 servers
Happy estimating! 🧮✨
Comments
Post a Comment