Today we are going to learn about something that every single website and app you use every day depends on — but almost nobody explains it clearly.
It is called Rate Limiting.
In 2024 alone, Cloudflare blocked over 6.8 million DDoS attacks — most of them were stopped by Rate Limiting before they could do any damage. Without rate limiting, your favourite apps would crash every single day. 🛡️
🎠 Let's Start With a Story
Imagine your favourite amusement park on a summer weekend. 🎡 Every single person in the city wants to ride the roller coaster at the same time.
Without any control, thousands of people rush the gate at once. People get hurt. The machine breaks. The park shuts down. Nobody gets a ride.
So the smart park manager puts up a gate with a simple rule:
"Only 10 people may enter every 2 minutes. Wait your turn. Everyone gets a safe, fair ride."
That gate is Rate Limiting. Your API is the roller coaster. The gate keeps it safe, fair, and working for everyone.
Rate Limiting = A technique that controls how many requests a user or system can make to a service within a specific time window.
Think of it as: "You can only knock on this door 10 times per minute. After that, you have to wait." 🚪
🌍 Real-World Examples You Already Know
Rate limiting is everywhere around you. You just never had a name for it before!
- 📱 Instagram / Twitter / X — You can only like or post a certain number of times per hour. This stops spam bots from flooding the platform with fake likes.
- 🏧 ATM Machine — Only 3 wrong PIN attempts before your card is blocked for 24 hours. This stops thieves from guessing your PIN by trying thousands of combinations.
- 🤖 OpenAI / Claude API — Free tier users get 10 requests per minute. Pro users get 1,000 per minute. Everyone pays fairly for what they use.
- 🎮 Online Games — Your character can only perform 60 actions per second. This stops cheaters from writing bots that attack at superhuman speed.
- ✉️ Gmail / Outlook — You can send only 500 emails per day. Without this, spammers would send billions of emails from a single account.
Why Do We NEED Rate Limiting? What Happens Without It?
Let's say you build a brand-new API and forget to add any rate limiting. Here is what happens within the first few days:
WITHOUT Rate Limiting: User sends 1 request ──────────────────► ✅ Server responds fine User sends 10 requests ──────────────────► ✅ Server responds fine User sends 1,000 requests ──────────────────► ⚠️ Server slowing down... Hacker sends 1,000,000 req ──────────────────► 💥 SERVER CRASHES! Result: EVERYONE is down. Real users get nothing. WITH Rate Limiting: User sends 10 requests ──► ✅ Allowed (within limit) User sends request #11 ──► 🚫 Blocked (limit hit — wait 60s) Hacker sends 1,000,000 req ──► 🚫 Blocked after first 10. Server safe. Result: Server is healthy. Real users get fast, reliable service. ✅
🤖 DDoS Attacks — Hackers flood your server with fake requests to crash it.
💸 Cloud Bill Explosion — Unexpected traffic spikes cost thousands of dollars overnight.
🐢 Slow Service for Everyone — One greedy user hogs all the resources.
🔒 Brute Force Attacks — Bots try millions of passwords per second until one works.
📉 Database Crashes — Too many queries overwhelm your database.
⚡ Cascade Failures — One slow service makes 10 dependent services fail too.
🧠 The 5 Rate Limiting Algorithms — From Simple to Expert
There are five main ways engineers implement rate limiting. Each has a different trade-off between simplicity, accuracy, and memory usage. Let's learn them one by one — from the easiest to the most sophisticated!
⏰ Algorithm 1: Fixed Window Counter
This is the simplest algorithm. Divide time into fixed windows — for example, every 60 seconds. Count how many requests arrive in each window. If the count goes over the limit, reject new requests until the next window begins.
Your mum gives you a jar of 10 cookies every hour. Once you eat all 10, you must wait. At the top of the next hour — boom! — a fresh jar of 10 cookies. The clock resets completely and the count goes back to zero.
Fixed Window — Limit: 5 requests per minute
12:00:00 ─────────────────────────────── 12:01:00
│ │
│ Request 1 ──► ✅ Count = 1 │
│ Request 2 ──► ✅ Count = 2 │
│ Request 3 ──► ✅ Count = 3 │
│ Request 4 ──► ✅ Count = 4 │
│ Request 5 ──► ✅ Count = 5 ← LIMIT HIT │
│ Request 6 ──► 🚫 BLOCKED │
│ Request 7 ──► 🚫 BLOCKED │
│ │
└───────────────────────────────────────────────┘
↓ At 12:01:00 — counter resets to 0 ↓
12:01:00 ─────────────────────────────── 12:02:00
│ │
│ Request 8 ──► ✅ Count = 1 (fresh start!) │
└───────────────────────────────────────────────┘
A sneaky user sends 5 requests at 12:00:59 (just before reset) and then 5 more at 12:01:00 (just after reset). Result: 10 requests in 2 seconds — double the intended limit!
This is called the thundering herd at the boundary problem. It is the main reason Fixed Window is not used for security-sensitive endpoints.
✅ Best For: Simple use cases, prototypes, internal tools, non-critical APIs.
❌ Avoid For: Login endpoints, payments, or anywhere burst attacks matter.
📜 Algorithm 2: Sliding Window Log
More accurate than Fixed Window — but uses more memory. Instead of a hard clock reset, we keep a log (list) of the exact timestamp of every single request that comes in.
When a new request arrives, we ask: "How many requests happened in the LAST 60 seconds counting from RIGHT NOW?" — not from the start of a fixed window. Log entries older than 60 seconds are discarded.
Instead of a fixed "photos from this month only" album, imagine a rolling 30-day window. At any moment, you count only the photos taken in the last 30 days from today. Old photos naturally "expire" as time moves forward. Smooth and continuous — no hard resets!
Sliding Window Log — Limit: 3 requests per 60 seconds
Stored log for this user:
┌────────────────────────────────────────────────┐
│ [12:00:10] [12:00:35] [12:00:50] │
└────────────────────────────────────────────────┘
New request arrives at 12:01:05:
→ Window = [12:00:05 → 12:01:05]
→ [12:00:10] ✅ inside window (count = 1)
→ [12:00:35] ✅ inside window (count = 2)
→ [12:00:50] ✅ inside window (count = 3) ← limit hit!
→ 🚫 BLOCKED
New request arrives at 12:01:15:
→ Window = [12:00:15 → 12:01:15]
→ [12:00:10] ❌ expired! Removed from log.
→ [12:00:35] ✅ inside window (count = 1)
→ [12:00:50] ✅ inside window (count = 2)
→ count = 2 < limit = 3 → ✅ ALLOWED!
✅ Best For: High-accuracy requirements, smaller-scale APIs.
❌ Avoid For: Very high traffic — each request stores a log entry.
1 million requests = 1 million log entries. Memory gets expensive fast!
⚖️ Algorithm 3: Sliding Window Counter (The Sweet Spot!)
The best of both worlds. It combines the memory-efficiency of Fixed Window with the accuracy of Sliding Window Log. Instead of storing every timestamp, it uses a simple weighted math formula to approximate the sliding window.
This is what Cloudflare, Stripe, and GitHub actually use in production!
Estimated count = (Previous window count × % of previous window still in range) + Current window count
Example: Limit = 100 req/min. Previous window had 80 requests. We are 25% into the current window.
Estimated count = (80 × 75%) + current window count = 60 + current count.
Only 2 numbers stored per user — not thousands of timestamps. Memory-efficient!
Sliding Window Counter — Limit: 10 req/min
Previous window (11:00–12:00): 8 requests stored
Current window (12:00–13:00): 4 requests so far
We are 30% into the current window
Estimated count = (8 × 70%) + 4
= 5.6 + 4
= 9.6
9.6 < 10 → ✅ ALLOW this request
──────────────────────────────────────────────────
Memory used: just 2 numbers per user 🎉
vs Log algo: 1 entry per request ❌
✅ Best For: Production systems at scale — great accuracy with very low memory.
⭐ Used By: Cloudflare, Stripe, GitHub, Redis's built-in rate limiter.
🪣 Algorithm 4: Token Bucket (Industry Favourite!)
Imagine a bucket that holds tokens (coins 🪙). New tokens are added at a steady rate — say, 2 tokens per second. Each request consumes one token. No token? Request is rejected.
The key magic: unused tokens accumulate (up to the bucket's capacity). So if you don't make requests for a while, you build up a reserve — and you can spend that reserve in a short burst whenever you need to.
Think of your character's stamina bar. It refills slowly over time. You can sprint (use multiple actions quickly) and drain the bar fast. But if you sprint too long, you run out and must wait for it to refill. Standing still lets you build a full bar for the next sprint. This is exactly how Token Bucket works!
Token Bucket — Capacity: 5, Refill rate: 1 token/second Start: bucket = [ 🪙 🪙 🪙 🪙 🪙 ] (full — 5 tokens) Request 1 → consume 1 → bucket = [ 🪙 🪙 🪙 🪙 ] → ✅ ALLOW Request 2 → consume 1 → bucket = [ 🪙 🪙 🪙 ] → ✅ ALLOW Request 3 → consume 1 → bucket = [ 🪙 🪙 ] → ✅ ALLOW Request 4 → consume 1 → bucket = [ 🪙 ] → ✅ ALLOW Request 5 → consume 1 → bucket = [ ] → ✅ ALLOW Request 6 → no tokens! bucket = [ ] → 🚫 BLOCK ... wait 3 seconds (3 tokens refill) ... bucket = [ 🪙 🪙 🪙 ] Request 7 → consume 1 → bucket = [ 🪙 🪙 ] → ✅ ALLOW KEY INSIGHT: Tokens saved during idle time = burst capacity for later.
✅ Best For: APIs where users legitimately need occasional bursts of activity.
⭐ Used By: AWS API Gateway, Nginx, Shopify, most mobile-first APIs.
💧 Algorithm 5: Leaky Bucket (Steady-Flow Algorithm)
Imagine a real bucket with a small hole at the bottom. 🪣 Water (requests) pours in from the top at any rate — fast or slow. But water drips out from the bottom at a constant, fixed rate. If the bucket overflows (too many requests poured in too fast), the extra water spills away and is rejected.
This creates a perfectly smooth, constant output rate no matter how chaotic the input traffic is.
No matter how fast you pour water into a drip coffee machine, the coffee comes out at the same calm, steady drip. Calm. Controlled. Consistent. Even if you dump a whole jug in, it still drips at the same speed!
Leaky Bucket — Queue capacity: 5, Output rate: 1 request/second
Many requests flood in (any rate)
↓ ↓ ↓ ↓ ↓ ↓ ↓ ↓
┌──────────────────────────────────────┐
│ QUEUE (holds max 5 requests) │ ← overflow = DROPPED 🚫
│ [ req ][ req ][ req ][ req ][ req ] │
└─────────────────┬────────────────────┘
│
▼ (exactly 1 per second — always steady)
✅ Server processes at fixed rate, always
No matter how many flood in at the top, exactly 1 goes through per second.
Leaky Bucket cannot handle bursts. If your users sometimes need to send 10 requests quickly (like a mobile app sync), those requests queue up and get dropped if the queue fills. Use Token Bucket instead when burst tolerance is needed.
✅ Best For: Video streaming, financial transactions, IoT sensor pipelines — anything needing perfectly smooth throughput.
⭐ Used By: Network routers, CDN providers, payment processors.
📊 Algorithm Comparison — One Quick Look
Algorithm │ Memory │ Accuracy │ Allows Bursts? │ Best For ────────────────────────┼─────────┼───────────┼────────────────┼────────────────────────── Fixed Window Counter │ 🟢 Low │ ⚠️ Low │ ⚠️ Edge only │ Prototypes, simple APIs Sliding Window Log │ 🔴 High │ 🟢 High │ ✅ Yes │ Small, high-accuracy APIs Sliding Window Counter │ 🟢 Low │ 🟢 High │ ✅ Yes │ Production at scale ⭐ Token Bucket │ 🟢 Low │ 🟢 High │ ✅✅ Great │ Burst-friendly APIs ⭐ Leaky Bucket │ 🟢 Low │ 🟢 High │ ❌ None │ Smooth, constant-rate systems
Starting out? → Fixed Window.
Need accuracy at scale? → Sliding Window Counter + Redis.
Users need occasional bursts? → Token Bucket.
Need perfectly smooth output? → Leaky Bucket.
🗺️ Where Does Rate Limiting Live in a Real System?
Rate limiting is not just one thing in one place. Think of it like security checkpoints at an airport — multiple layers, each catching different problems. A real production system has rate limiting at every level.
👤 USER (browser, mobile app, script)
│
▼
┌─────────────────────────────────────────────────────────┐
│ LAYER 1 — CDN / DNS (e.g. Cloudflare, Akamai) │
│ • Blocks known bad IPs before they touch your server │
│ • Handles 100M+ rate limit checks per second globally │
│ • First line of defence against DDoS attacks │
└──────────────────────────┬──────────────────────────────┘
│ ✅ Only clean traffic passes
▼
┌─────────────────────────────────────────────────────────┐
│ LAYER 2 — API Gateway (e.g. Kong, AWS API GW, Nginx) │
│ • Enforces per-API-key and per-user limits │
│ • Different limits per route (/login vs /search) │
│ • Returns 429 + Retry-After headers automatically │
└──────────────────────────┬──────────────────────────────┘
│ ✅ Authenticated, within-limit traffic
▼
┌─────────────────────────────────────────────────────────┐
│ LAYER 3 — Load Balancer (e.g. HAProxy, AWS ALB) │
│ • Limits concurrent connections per IP at network level │
│ • Distributes remaining load evenly across servers │
└──────────────────────────┬──────────────────────────────┘
│ ✅ Balanced traffic
▼
┌─────────────────────────────────────────────────────────┐
│ LAYER 4 — Application Code / Microservice │
│ • Fine-grained business logic limits │
│ • e.g. "User can submit only 3 orders per minute" │
│ • Uses Redis for shared state across all instances │
└──────────────────────────┬──────────────────────────────┘
│ ✅ Final response
▼
📱 Response returned to user
Each layer catches different kinds of abuse. If a bad actor slips past Layer 1, Layer 2 catches them. If they somehow bypass Layer 2, Layer 3 stops them. This multi-layer approach is called "Defence in Depth" — a core security principle.
📡 HTTP Rate Limit Headers — What the Server Says Back
When an API rate limits you, it does not just say "no" and leave you confused. It sends back HTTP headers that explain exactly what happened and when you can try again.
Think of headers like the nutrition label on a food packet. The food (data) is inside — but the label tells you all the important details about it.
── Normal response (within limits) ─────────────────────────
HTTP/1.1 200 OK
X-RateLimit-Limit: 100 ← Max requests allowed per window
X-RateLimit-Remaining: 63 ← How many you have LEFT right now
X-RateLimit-Reset: 1720000060 ← Unix timestamp when limit resets
── When you hit the limit ───────────────────────────────────
HTTP/1.1 429 Too Many Requests
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0 ← Zero left!
Retry-After: 45 ← Wait exactly 45 seconds, then retry
Body:
{
"error": "rate_limit_exceeded",
"message": "Too many requests. Please wait 45 seconds.",
"retry_after": 45
}
Getting a 429 response and immediately retrying in a loop. This makes things worse — you keep getting blocked and you flood the server with wasted requests.
Always read the
Retry-After header and wait that exact number of seconds
before trying again. This is called honouring the rate limit.
🔴 Distributed Rate Limiting with Redis
Here is a tricky problem that catches many engineers off guard. What happens when your service runs on 10 different servers?
If Server A counts a user's requests in its own memory, and the user's next request goes to Server B — Server B has no idea what Server A counted! The user can now make 100 requests per server × 10 servers = 1,000 requests, while only 100 were supposed to be allowed!
❌ BAD — In-Memory Rate Limiting (broken in multi-server setup) Server A: User hit limit → counter = 100. BLOCKED. Server B: User hit limit → counter = 100. BLOCKED. Server C: User hit limit → counter = 100. BLOCKED. Total actual requests: 300! User was only supposed to get 100. 😱 ✅ GOOD — Centralised Redis Counter (correct in any setup) Server A ──┐ Server B ──┼──► REDIS (shared counter) → user:alice = 247 → BLOCKED! Server C ──┘ All servers read and write to the SAME Redis counter. No matter which server gets the request, the total is always accurate.
Redis is an in-memory database that responds in under 1 millisecond. It has an atomic INCR command — it can increment a counter without any race conditions, even if 1,000 servers hit it at the exact same moment. The count is always perfectly accurate. No double-counting. Ever.
💻 Code Examples — Let's Build It!
Example 1 — Simple Fixed Window Rate Limiter (Python)
This builds the simplest possible rate limiter using a Python dictionary as a notebook. It tracks: "How many times has User X made a request in the current minute?" If they exceed the limit, we say NO and ask them to wait until the next window. This is great for learning — but only works on a single server.
import time
from collections import defaultdict
# This class is our "rate limit notebook"
class FixedWindowRateLimiter:
def __init__(self, max_requests: int, window_seconds: int):
# How many requests are allowed per window? e.g. 5
self.max_requests = max_requests
# How long is one window in seconds? e.g. 60
self.window_seconds = window_seconds
# Our notebook: { user_id → { count, window_start } }
self.store = defaultdict(
lambda: {"count": 0, "window_start": time.time()}
)
def is_allowed(self, user_id: str) -> bool:
record = self.store[user_id]
now = time.time()
# Has the current window expired? If yes, start fresh.
if now - record["window_start"] >= self.window_seconds:
record["count"] = 0
record["window_start"] = now
# Is there still room in this window?
if record["count"] < self.max_requests:
record["count"] += 1
return True # ✅ ALLOWED
return False # 🚫 BLOCKED
# ── Test it ──────────────────────────────────────────────────
limiter = FixedWindowRateLimiter(max_requests=5, window_seconds=60)
for i in range(8):
result = limiter.is_allowed("alice")
status = "✅ ALLOWED" if result else "🚫 BLOCKED"
print(f"Request {i+1}: {status}")
🖥️ Expected Output: Request 1: ✅ ALLOWED Request 2: ✅ ALLOWED Request 3: ✅ ALLOWED Request 4: ✅ ALLOWED Request 5: ✅ ALLOWED Request 6: 🚫 BLOCKED Request 7: 🚫 BLOCKED Request 8: 🚫 BLOCKED
Example 2 — Token Bucket Rate Limiter (Python)
This builds a Token Bucket. Think of a piggy bank that automatically gets coins added over time. Every request spends one coin. Empty piggy bank = request blocked. But if you don't spend coins for a while, they build up — letting you burst later. The
threading.Lock() makes it safe when many users hit the same server at once.
import time
import threading
class TokenBucketRateLimiter:
def __init__(self, capacity: int, refill_rate: float):
# Max tokens the bucket can hold (e.g. 5)
self.capacity = capacity
# Tokens added per second (e.g. 1.0 = 1 token per second)
self.refill_rate = refill_rate
self.tokens = float(capacity) # Start with a full bucket
self.last_refill = time.monotonic()
self.lock = threading.Lock() # Safety lock for multiple threads
def _refill(self):
now = time.monotonic()
elapsed = now - self.last_refill # Seconds since last refill
new_tokens = elapsed * self.refill_rate # Tokens earned in that time
self.tokens = min(self.capacity, self.tokens + new_tokens)
self.last_refill = now
def consume(self, tokens_needed: int = 1) -> bool:
with self.lock:
self._refill()
if self.tokens >= tokens_needed:
self.tokens -= tokens_needed
return True # ✅ ALLOWED — token consumed
return False # 🚫 BLOCKED — not enough tokens
# ── Demo: burst, then wait, then more ────────────────────────
limiter = TokenBucketRateLimiter(capacity=5, refill_rate=1.0)
print("=== Burst of 8 requests instantly ===")
for i in range(8):
result = limiter.consume()
print(f" Request {i+1}: {'✅ ALLOWED' if result else '🚫 BLOCKED'}")
time.sleep(3) # Wait 3 seconds → 3 new tokens are added
print("\n=== After waiting 3 seconds ===")
for i in range(4):
result = limiter.consume()
print(f" Request {i+1}: {'✅ ALLOWED' if result else '🚫 BLOCKED'}")
🖥️ Expected Output: === Burst of 8 requests instantly === Request 1: ✅ ALLOWED Request 2: ✅ ALLOWED Request 3: ✅ ALLOWED Request 4: ✅ ALLOWED Request 5: ✅ ALLOWED Request 6: 🚫 BLOCKED Request 7: 🚫 BLOCKED Request 8: 🚫 BLOCKED === After waiting 3 seconds === Request 1: ✅ ALLOWED ← 3 new tokens refilled! Request 2: ✅ ALLOWED Request 3: ✅ ALLOWED Request 4: 🚫 BLOCKED ← only 3 tokens refilled, not 4
Example 3 — Production Redis Rate Limiter (Python)
This is a production-ready distributed rate limiter backed by Redis. Instead of counting requests in server memory (which breaks across multiple servers), we store the counter in Redis — a shared database that every server reads from and writes to. Works correctly even with 100 servers running at the same time. Also shows how to apply different limits to different endpoints.
# Install first: pip install redis
import redis
import time
from dataclasses import dataclass
@dataclass
class RateLimitResult:
allowed: bool # Was the request allowed?
remaining: int # How many requests left this window?
reset_after: int # Seconds until this window resets
class RedisRateLimiter:
def __init__(self, redis_url="redis://localhost:6379"):
self.redis = redis.from_url(redis_url, decode_responses=True)
def check(self, key: str, limit: int, window: int) -> RateLimitResult:
"""
key → unique identifier, e.g. "user:alice" or "ip:1.2.3.4"
limit → max requests allowed per window
window → window duration in seconds
"""
now = int(time.time())
# Key includes the current window number — auto-namespaces by time slice
redis_key = f"rl:{key}:{now // window}"
# Pipeline = send both commands in one trip to Redis (much faster!)
pipe = self.redis.pipeline()
pipe.incr(redis_key) # Atomically add 1 to counter
pipe.expire(redis_key, window) # Auto-delete key when window ends
results = pipe.execute()
current = results[0] # Total requests so far this window
remaining = max(0, limit - current)
reset_after = window - (now % window)
return RateLimitResult(
allowed=current <= limit,
remaining=remaining,
reset_after=reset_after
)
# ── Usage inside a Flask / FastAPI request handler ─────────────
limiter = RedisRateLimiter()
def handle_request(user_id: str, endpoint: str):
# Different endpoints get different limits
limits = {
"/search": (100, 60), # 100 requests per minute
"/checkout": (10, 60), # 10 per minute (very sensitive!)
"/login": (5, 300), # 5 attempts per 5 minutes
}
limit, window = limits.get(endpoint, (1000, 3600))
result = limiter.check(f"user:{user_id}", limit, window)
if not result.allowed:
return {
"status_code": 429,
"error": "rate_limit_exceeded",
"retry_after": result.reset_after
}
return {"status_code": 200, "requests_remaining": result.remaining}
🏆 Advanced Pattern — Tiered Rate Limiting
Real products like Stripe, Twilio, and OpenAI do not give everyone the same rate limit. Different subscription tiers get different quotas. This is how SaaS companies monetise their APIs while keeping the service fair for all users.
Tier │ Rate Limit │ Why ──────────────┼────────────────────┼────────────────────────────────────────── 🆓 Free │ 10 req/min │ Enough to explore, not enough to abuse ⭐ Pro │ 100 req/min │ Real workloads, pay for more 💎 Business │ 1,000 req/min │ Teams and production usage 🏢 Enterprise │ Custom / Unlimited │ Direct contract, dedicated infrastructure
Store the user's tier in your database or inside their JWT token. When a request arrives, look up their tier, get the right limit for that tier, then pass it to your rate limiter. The same rate limiter code serves all tiers — you just feed it different numbers.
Trend — AI-Powered Adaptive Rate Limiting
The biggest shift in rate limiting is moving from fixed numbers (always allow exactly 100 requests/min for everyone) to dynamic, AI-driven limits that adapt in real time based on what is actually happening.
- 📊 Behavioural Analysis — AI models detect bot-like patterns (e.g. requests spaced exactly 100ms apart, or a flood from a single subnet). The limit tightens automatically for that specific user while real users are unaffected.
- 🌡️ Server-Health-Aware Limits — When CPU usage climbs above 80%, the system automatically reduces the global rate limit. When the server recovers, limits automatically relax again. No human intervention needed.
- 🎯 Request Cost Scoring — Not all requests cost the same. A complex database join costs 10× more than a simple cache hit. Modern limiters score each request by its actual resource cost and limit by cost, not just count. OpenAI's Tokens Per Minute model is an early version of this.
- 📈 Predictive Pre-Scaling — ML models trained on historical traffic predict upcoming spikes (Black Friday, a viral product launch) and pre-adjust limits before the spike hits — not after the server is already struggling.
Envoy Proxy (global rate limiting used by Netflix, Lyft) · Kong AI Gateway · Cloudflare Workers Rate Limiting · AWS WAF Adaptive Protection · Upstash Ratelimit (serverless and edge-native)
✅ DOs and ❌ DON'Ts — The Complete Checklist
• Always return a 429 status code and a Retry-After header when blocking. Never leave the client guessing.
• Use Redis the moment your service runs on more than one server.
• Set different limits per endpoint — a login route needs 10× stricter limits than a product search.
• Log every rate limit event — this data helps you detect attacks and tune your limits with real evidence.
• Whitelist your own monitoring IPs — health checks and load tests should never be rate limited.
• Implement exponential backoff with jitter on the client: wait 1s, then 2s, then 4s + random offset.
• Run in dry-run mode first — log who would be blocked without actually blocking them. Tune the numbers. Then enforce.
• Never silently drop requests without a clear error message. Always tell the client why and when to retry.
• Never use in-memory rate limiting across multiple servers. It is broken by design.
• Never use the same limit for all endpoints. A checkout API is not the same risk as a product listing.
• Never forget to handle Redis failure gracefully. If Redis goes down, decide in advance: allow all (fail open) or block all (fail closed)?
• Never rate limit only by IP address in corporate or campus environments — thousands of users can share one IP.
• Never set limits based on guessing. Use real traffic data from your logs and start permissive, then tighten.
• Never let a client retry in a loop immediately after a 429. Always enforce a waiting period.
🗝️ Quick Decision Guide — Which Algorithm Should I Pick?
"I'm building my first API and need something simple fast." → Fixed Window Counter. Takes 10 minutes to write. Good enough for prototypes. "I need accuracy and cannot afford boundary bursts." → Sliding Window Log (small scale) or Sliding Window Counter (large scale). "My users sometimes legitimately need to send many requests quickly." → Token Bucket. Allows short bursts while controlling the average rate. "I need perfectly smooth, constant throughput — no spikes ever." → Leaky Bucket. Great for payments, video streaming, IoT pipelines. "I run on multiple servers and need consistent limits across all of them." → Sliding Window Counter + Redis. The production gold standard. "I want my limits to automatically adjust based on server health and behaviour." → Adaptive Rate Limiting with Envoy, Kong AI Gateway, or Cloudflare Workers.
Happy rate limiting! 🚦✨
Comments
Post a Comment