Idempotency in System Design: Build Reliable APIs and Prevent Duplicate Payments
It is 2 AM. A customer's payment is processing. The network hiccups. The client doesn't know if the charge went through. It retries. The charge goes through twice. The customer wakes up to two charges. Angry emails. Refunds. Lost trust.
This nightmare scenario plays out in poorly designed systems every day. And it can be prevented entirely with one concept: Idempotency.
Idempotency is the guarantee that performing the same operation multiple times produces exactly the same result as performing it once. It is the engineering superpower that makes retries safe, distributed systems reliable, and customers happy. Master it and you will design systems that never accidentally charge someone twice — no matter what the network does.
💳 Stripe: Every payment API call accepts an idempotency key — charges are never duplicated
☁️ AWS: EC2 RunInstances, SQS, SNS — all designed for idempotent retries
🚗 Uber: Trip creation and payment are idempotent — one trip, one charge, always
🛒 Amazon: Order placement uses idempotency to prevent duplicate orders on network errors
📦 Kubernetes: All kubectl apply operations are idempotent — apply the same YAML 100 times, same result
📨 Kafka / SQS: Message delivery is "at-least-once" — consumers must be idempotent
If you are building any distributed system, API, or payment flow — you must understand idempotency. There is no alternative.
🔢 Section 1: The Math Behind Idempotency — Surprisingly Simple
The word "idempotent" comes from Latin: idem (same) + potens (power). In mathematics, a function is idempotent if applying it multiple times gives the same result as applying it once.
🔢 The Idempotency Formula
Applying f once or a thousand times — the result is always the same.
🔥 Section 2: The Core Problem — Networks Are Unreliable, Retries Are Inevitable
In a perfect world, every API call succeeds on the first try. We don't live in a perfect world. Networks drop packets. Servers time out. Load balancers restart. And when that happens, the client must retry.
A light switch is idempotent. Flip it ON → light is on. Flip it ON again → light is still on. No matter how many times you flip it ON, the result is always the same: light is on. ✅
Now imagine a button that adds ₹1,000 to your bill every time you press it. That is NOT idempotent. Each press changes the state. If you press it by accident twice → ₹2,000 charged. 😱
Your payment API should be a light switch, not a bill-adding button.
😱 The Duplicate Request Problem — What Happens Without Idempotency
💀 Scenario A — Request succeeds but response is lost
sends POST
/charge $100
charges $100
✅ SUCCESS
lost! Client
sees TIMEOUT
retries POST
/charge $100
charged!
DUPLICATE!
💀 Scenario B — User double-clicks the "Pay Now" button
"Pay Now"
2 requests sent
2 POST requests
simultaneously
created! 2 charges!
DISASTER!
💀 Scenario C — Load balancer retries on upstream timeout
POST /order
server-1 slow →
retries to server-2
AND server-2
create the order!
In any distributed system, retries are unavoidable. Networks fail. Servers time out. Load balancers retry. You cannot prevent retries — especially in microservices architectures.
The only solution: make your operations idempotent so that retrying them is always safe. An operation that can be safely retried any number of times without side effects is idempotent. Build everything to be idempotent from day one.
🌐 Section 3: HTTP Methods — Which Are Idempotent and Why
The HTTP specification explicitly classifies methods by idempotency. This is foundational knowledge for any API designer.
| 🌐 HTTP Method | 🛡️ Idempotent? | 🔒 Safe? | 📝 Explanation | 💡 Example |
|---|---|---|---|---|
| GET | ✅ YES | ✅ YES | Read-only. Calling 100 times returns same data (assuming no changes). Never modifies state. | GET /users/42 |
| PUT | ✅ YES | ❌ NO | Sets a resource to an exact state. Calling "PUT name=Alice" 100 times → name is still Alice. Same result. | PUT /users/42 with full body |
| DELETE | ✅ YES | ❌ NO | Delete user 42 once → gone. Delete user 42 again → still gone (return 404). End state: user is deleted. | DELETE /users/42 |
| PATCH | ⚠️ DEPENDS | ❌ NO | "Set age=30" is idempotent. But "Increment age by 1" is NOT. Depends entirely on what the patch does. | PATCH /users/42 |
| POST | ❌ NO | ❌ NO | Creates new resources each time. POST /orders twice → two orders. This is the method that NEEDS idempotency keys most urgently. | POST /orders |
Use PUT instead of POST when replacing an entire resource — it's naturally idempotent.
Use POST for creating resources, but add an idempotency key to make it safe to retry.
A good API design uses PUT wherever possible (idempotent by nature) and adds idempotency keys to POST endpoints that create resources.
🔑 Section 4: Idempotency Keys — The Universal Solution
The idempotency key is the most important and most widely used pattern for making non-idempotent operations safe. It is how Stripe, AWS, Uber, and every serious payment and API system works.
When a waiter takes your order, they write it on a numbered ticket: #4729. If they later wonder "did I send order #4729 to the kitchen yet?" they check their records. If yes → don't send again. If no → send now.
The ticket number is the idempotency key. The waiter's record book is the server's idempotency store. The same order number always means the same order — no matter how many times they check.
🔑 How Idempotency Keys Work — Step by Step
The client creates a UUID:
e8f3a2c1-9b4d-4f7e-8a3c-2b1d5e6f9a0b.
This key uniquely identifies THIS specific intent to perform THIS operation.
The key is stored on the client side and sent with every retry attempt.
POST /charges with header:
Idempotency-Key: e8f3a2c1-9b4d-4f7e-8a3c-2b1d5e6f9a0b
The server does an atomic lookup: "Has key
e8f3a2c1... been processed?"If NO: Process the request normally. Store the key + response. Return response.
If YES: Do NOT process again. Return the STORED response from the first attempt. Done!
💳 Section 5: Idempotency in Payment Systems — The Highest Stakes
There is no domain where idempotency matters more than payments. A duplicated charge can lose a customer forever. This is why every serious payment company builds idempotency as a core primitive.
💳 How Stripe Implements Idempotency
Stripe is the gold standard for idempotency implementation. Their approach is elegant, battle-tested, and widely copied.
💳 Stripe's Idempotency System — How It Works
Idempotency-Key: uuid-here.
On ANY retry (timeout, 5xx error, network error), send the SAME key.
SET key value NX EX 86400
(set if not exists, expires in 24 hours). This prevents race conditions
where two simultaneous requests with the same key both try to process.
💳 Payment Idempotency — The Complete State Machine
Key stored in Redis.
Processing started.
Response cached.
Any retry → return cache.
Error cached too!
Retry returns same error.
🗄️ Section 6: Database Idempotency Patterns
At the database level, several patterns naturally implement idempotency. Knowing these makes your data layer crash-safe without any external key store.
🔀 Pattern 1: UPSERT (INSERT ON CONFLICT)
UPSERT = UPDATE + INSERT combined. "If the row already exists → update it. If it doesn't → insert it."
This is inherently idempotent: running the same UPSERT 100 times always results in exactly one row with the specified values. No duplicates. No errors. Pure idempotency at the SQL level.
🔐 Pattern 2: Conditional Updates (Optimistic Locking)
Problem: Two concurrent requests try to update the same record. Last-write-wins causes the first update to be silently overwritten.
Solution — Version Numbers: Every row has a version column.
Any UPDATE must include WHERE version = current_version.
If the version doesn't match → the update fails (someone else already updated it).
The client retries with the fresh version. No silent data loss.
This is called Optimistic Locking — and it's idempotent because the same update with the same version number can safely be retried: first attempt updates and increments version; second attempt finds version doesn't match and fails cleanly.
🗂️ Pattern 3: Deduplication Table
Use when: You need idempotency across multiple tables or complex business operations that can't be expressed as a single UPSERT.
How: Maintain a separate processed_events table.
Before processing any event, try to INSERT the event ID into this table.
If the INSERT succeeds (no conflict) → process the event.
If it fails with a duplicate key error → event was already processed → skip it safely.
Works perfectly with Kafka consumer offsets and SQS message IDs.
📨 Section 7: Message Queues — At-Least-Once vs Exactly-Once
Message queues are where idempotency becomes absolutely non-negotiable. Understanding delivery guarantees is essential for every backend engineer.
Message is sent once. If delivery fails → message is lost. Never retried.
Risk: Message loss. You might miss events permanently.
Use for: Metrics, logs, analytics — where occasional loss is acceptable.
Idempotency needed: No (but you'll lose data).
Message is delivered until the consumer ACKs it. On consumer crash, message is redelivered.
Risk: Duplicate processing — consumer may see the same message multiple times.
Use for: Most systems — Kafka default, SQS standard queues.
Idempotency needed: YES — consumers MUST be idempotent!
Each message is delivered and processed exactly once. No loss. No duplicates.
Reality: True exactly-once is impossible in pure distributed systems
(see: Two Generals Problem). Systems that offer it (Kafka transactions, SQS FIFO)
achieve it by combining at-least-once delivery WITH idempotent consumers.
Idempotency needed: Built into the framework (but understanding it helps greatly).
⚠️ Why At-Least-Once Requires Idempotent Consumers
delivers msg
order_id=789
processes order_id=789
Creates DB row ✅
before ACKing!
Kafka redelivers!
order_id=789 AGAIN
Already processed!
checks DB → already done
Skip! ACK message.
🌐 Section 8: Idempotency in Real-World Systems
Every POST endpoint accepts an Idempotency-Key header.
Keys stored in Redis with a 24-hour TTL. Entire response cached — including error responses.
If a charge fails with "card declined", retry returns the same "card declined" —
it doesn't retry the charge. Client must generate a new key to re-attempt.
The idempotency key is tied to the specific operation, not the user or session.
EC2 RunInstances: ClientToken parameter prevents duplicate instance launches.
SQS FIFO Queues: Message deduplication ID — same ID within 5-minute window → deduplicated.
DynamoDB: Conditional writes with version attributes for optimistic locking.
Lambda: Powertools library provides idempotency decorator for event processing.
kubectl apply -f deployment.yaml can be run 1,000 times.
The cluster always converges to the desired state in the YAML.
If the deployment already matches → no-op. If different → update.
This is called declarative configuration — describing the desired state
rather than the steps to get there. Declarative = inherently idempotent.
Trip creation uses idempotency keys so a network hiccup doesn't create two trips. Payment processing is idempotent — a driver can only be paid once per trip regardless of retries. Uber's Cadence workflow system provides idempotent workflow execution — even if a workflow step crashes and replays, each activity executes at-most-once.
git pull is idempotent — pulling when already up-to-date does nothing.
git checkout branch is idempotent — switching to a branch you're already on does nothing.
Content-addressable storage (SHA hashes) means identical content always has the same hash —
no matter how many times you add it. Git is a masterclass in idempotent design.
💻 Section 9: Code Examples — Building Idempotency from Scratch
This SQL code shows the UPSERT pattern — the simplest and most powerful database-level idempotency technique. It combines INSERT and UPDATE into one atomic operation. If a row with this user_id already exists → update it. If it doesn't → insert it. Running this exact SQL 1,000 times always produces exactly one row with the specified values. No duplicates. No errors on retry. The last three examples show progressively more sophisticated uses: inventory management (ensure minimum stock), payment deduplication, and a counter that only increments if the row is new.
-- ✅ UPSERT PATTERN — Idempotent Write (PostgreSQL) -- Running this 100 times always produces exactly one row with these values. -- Basic UPSERT: create user or update their email INSERT INTO users (user_id, email, updated_at) VALUES ('user-42', 'alice@example.com', NOW()) ON CONFLICT (user_id) -- "if user_id already exists..." DO UPDATE SET email = EXCLUDED.email, -- EXCLUDED = the values we tried to INSERT updated_at = NOW(); -- Result: exactly one row. Same query 100 times = same result. ✅ -- Payment deduplication UPSERT -- Prevents charging twice if the same payment_id is received multiple times INSERT INTO payments (payment_id, user_id, amount, status, created_at) VALUES ('pay-abc123', 'user-42', 9999, 'completed', NOW()) ON CONFLICT (payment_id) DO NOTHING; -- If payment_id exists → do absolutely nothing. Idempotent! ✅ -- First call: inserts the payment. Second call: no-op. Customer charged once. -- Optimistic locking UPSERT — only update if you have the right version -- Prevents "lost update" race condition between concurrent writers UPDATE account_balances SET balance = balance - 100, -- deduct $100 version = version + 1, -- bump version updated_at = NOW() WHERE account_id = 'acct-789' AND version = 5 -- only update if version matches! AND balance >= 100; -- only if sufficient funds -- If version changed (another process updated first) → 0 rows affected → retry! -- The application checks rows_affected == 1. If 0, fetch fresh data and retry. -- This is Optimistic Locking — idempotent and race-condition safe.
This code implements a reusable idempotency middleware
that can be added to any API endpoint with a single decorator.
It works exactly like Stripe's idempotency system.
When a request arrives with an Idempotency-Key header, the middleware:
(1) checks Redis for a previously cached response,
(2) if found, returns it immediately without touching the handler,
(3) if not found, calls the real handler, caches its response with a 24-hour TTL,
and returns it.
The NX flag in the Redis SET ensures atomic check-and-set —
preventing two simultaneous requests with the same key from both processing.
Add this decorator to any POST endpoint and it becomes safely retryable instantly!
# ✅ Idempotency Key Middleware (Python / FastAPI style pseudocode) # Add @idempotent to any POST endpoint to make it safely retryable. import functools, json, uuid from datetime import timedelta KEY_TTL = timedelta(hours=24) # how long we remember processed requests def idempotent(func): """ Decorator that makes any POST endpoint idempotent. Usage: @idempotent on any API handler function. Requires 'Idempotency-Key' header from the client. """ @functools.wraps(func) async def wrapper(request, *args, **kwargs): # Step 1: Extract the idempotency key from the request header idempotency_key = request.headers.get("Idempotency-Key") if not idempotency_key: return error_response( status=400, msg="Idempotency-Key header is required for this endpoint." ) # Step 2: Namespace the key (prevents collisions across endpoints/users) cache_key = f"idem:{request.user.id}:{request.path}:{idempotency_key}" # Step 3: Try to ATOMICALLY claim this key in Redis # NX = "set only if Not eXists" — prevents two concurrent requests # from both thinking they're the first! claimed = redis.set( name = cache_key, value = "PROCESSING", # placeholder while we work nx = True, # only set if key doesn't exist ex = int(KEY_TTL.total_seconds()) ) if not claimed: # Key already exists — check if it's still processing or done cached = redis.get(cache_key) if cached == "PROCESSING": # Another request is processing right now — wait and retry return error_response(status=409, msg="Request is being processed. Retry shortly.") # A previous request completed — return its cached response! stored = json.loads(cached) return response( status = stored["status"], body = stored["body"], headers = {"Idempotent-Replayed": "true"} # tell client it's a replay ) # Step 4: We claimed the key — execute the actual handler try: result = await func(request, *args, **kwargs) status_code = 200 except ClientError as e: result = {"error": str(e)} status_code = e.status_code # cache errors too! (e.g. 400 Bad Request) # Step 5: Cache the response (even errors!) with the full TTL redis.set( name = cache_key, value = json.dumps({"status": status_code, "body": result}), ex = int(KEY_TTL.total_seconds()) ) return response(status=status_code, body=result) return wrapper # Usage — just add @idempotent to any POST endpoint! @idempotent async def charge_customer(request): # This function now safely handles retries automatically. # Even if the client retries 50 times, the charge happens exactly once. charge = payment_gateway.charge( user_id = request.user.id, amount = request.body["amount"] ) return {"charge_id": charge.id, "status": "succeeded"}
This code shows an idempotent Kafka consumer — the correct way to process
messages from any at-least-once delivery system (Kafka, SQS, RabbitMQ).
Since Kafka can deliver the same message multiple times (if the consumer crashes before ACKing),
the consumer MUST check if it already processed this message before doing any work.
We use a deduplication table in the database — before processing any message,
we attempt to INSERT its message ID into processed_events.
If the INSERT fails with a unique constraint error → we've already processed this message → skip it.
If the INSERT succeeds → we haven't seen it → process and ACK.
This pattern works with any database and any message queue!
# ✅ Idempotent Kafka Consumer — Deduplication Table Pattern (Pseudocode) # Works for any at-least-once message queue: Kafka, SQS, RabbitMQ, etc. def process_order_created_event(message): """ Kafka calls this function for every message. Because Kafka is at-least-once, we might see the same message multiple times. We use a deduplication table to ensure we process each order exactly once. """ message_id = message.headers["message_id"] # unique ID per message order_data = message.value # ───────────────────────────────────────────── # Step 1: Try to record this message_id as "being processed" # Uses a DB transaction so this is atomic even under concurrent consumers. # processed_events table has: UNIQUE CONSTRAINT on message_id try: db.execute( """ INSERT INTO processed_events (message_id, topic, processed_at) VALUES (:message_id, :topic, NOW()) """, {"message_id": message_id, "topic": "order_created"} ) except UniqueConstraintViolation: # This message_id already in the table → already processed! # Safely skip — do NOT process again. log(f"[SKIP] Already processed message {message_id}. This was a duplicate delivery.") kafka.commit(message) # acknowledge to Kafka so it stops redelivering return # ───────────────────────────────────────────── # Step 2: INSERT succeeded — this is a NEW message. Process it! try: # Do all the real business work here order_service.create_order(order_data) inventory_service.reserve_items(order_data["items"]) notification_service.send_confirmation(order_data["user_id"]) # Step 3: ACK the message so Kafka knows we're done kafka.commit(message) log(f"[SUCCESS] Processed order {order_data['order_id']} from message {message_id}") except Exception as e: # Processing failed! Remove from deduplication table so we CAN retry. # Next delivery of this message will be processed (not skipped). db.execute( "DELETE FROM processed_events WHERE message_id = :mid", {"mid": message_id} ) # Don't ACK! Kafka will redeliver this message. We want to retry. raise e # let Kafka know to retry
Notice that when processing fails, we DELETE the deduplication row before re-raising the error. This is essential — if we kept the row, future retries would see it and skip the message forever, leaving the order un-created.
Only delete the deduplication row on genuine errors that warrant retry. This pattern ensures: successfully processed messages are skipped on retry. Failed messages are retried until they succeed or hit the dead letter queue. 🎯
📐 Section 10: Implementation Guidelines — Building It Right
The client must generate the key before sending. If the server generated the key,
a client would need a separate round-trip to get it — and THAT round-trip could fail!
Client-side generation (UUID v4) means the key is ready before the first attempt.
import uuid; key = str(uuid.uuid4()) — done.
A key of "abc123" from User A and "abc123" from User B should not collide.
Always namespace your keys:
idem:{user_id}:{endpoint}:{client_key}.
This prevents a malicious user from using another user's idempotency key
to interfere with their requests.
If a request fails with a 400 Bad Request, cache THAT response. A retry with the same key returns "400 Bad Request" without re-executing the handler. This prevents a subtle bug where a bad request accidentally succeeds on retry if the handler is partially side-effectful.
An idempotency key is NOT a security token. It does not replace authentication. Anyone can send any key — the server must still authenticate the user normally. The key only affects deduplication, not access control.
Each unique business intent needs a unique key. If you charge $100 on Monday and want to charge another $100 on Tuesday → generate a NEW key for Tuesday's charge. Reusing Monday's key returns "charged $100" from the cache — without charging again. One key = one operation. Always.
Too short → legitimate retries (hours later) aren't protected → possible duplicate.
Too long → Redis memory bloat → expensive storage.
Typical range: 24 hours to 7 days depending on expected retry window.
Payments: 24 hours. Long-running workflows: up to 30 days.
Match TTL to your SLA for "how long might a client retry?"
🗺️ Section 11: Everything Together — Complete Idempotency Architecture
🛡️ Idempotency — Complete Architecture Map
Idempotency-Key: e8f3a2c1-9b4d-4f7e-8a3c-2b1d5e6f9a0b
idem:user42:POST:/charges:e8f3a2c1
(process charge)
(UPSERT / conditional write)
(publish event)
Idempotency store
24-hour TTL
Deduplication table
processed_events
Idempotent processing
via dedup table
🎓 Section 12: System Design Interview Cheat Sheet
Idempotency appears in system design interviews in many forms. Here is exactly how to handle every variant:
Answer: Client generates UUID before payment attempt. Sends as Idempotency-Key header.
Server atomically checks Redis (SET NX). On first request: process + cache response.
On retry: return cached response. TTL = 24 hours. Also add database-level deduplication
(UNIQUE constraint on payment_id) as a second line of defence.
Answer: Propagate idempotency keys through the entire call chain. The API gateway assigns a request ID. Every downstream service call includes this ID. Each service uses it as an idempotency key. This way, even if the load balancer retries internally, the entire chain is idempotent end-to-end.
Answer: Idempotent consumers via a deduplication table.
Before processing any message, INSERT its message ID into a processed_events table
with a UNIQUE constraint. If INSERT fails (duplicate) → skip. If succeeds → process.
On processing failure → DELETE the row so retries are allowed.
Alternatively: use UPSERT instead of INSERT for the business operation itself.
Answer: GET (read-only, never changes state), PUT (sets resource to exact state), DELETE (resource is gone whether called once or 100 times). POST is NOT idempotent — each call creates a new resource. PATCH depends on the operation — "set field X to Y" is idempotent, "increment by 1" is not.
Answer: On checkout page load, generate a UUID order_token.
Store it in localStorage. Every "Place Order" click sends the same token.
Server uses token as idempotency key in Redis (24-hour TTL).
Database: INSERT INTO orders ... ON CONFLICT (order_token) DO NOTHING.
Even if user clicks 10 times or network retries 5 times → one order created.
Answer: At-least-once → message delivered but might be redelivered on consumer crash. Requires idempotent consumers. Exactly-once → true exactly-once is theoretically impossible in distributed systems (Two Generals Problem). Systems like Kafka Transactions achieve "effectively exactly-once" by combining at-least-once delivery with idempotent producers/consumers. In practice: build idempotent consumers and at-least-once delivery is "good enough exactly-once."
🎉 Final Summary
Idempotency is not an optimisation — it is a correctness requirement for any distributed system.
Networks will fail. Clients will retry. Load balancers will retry. Users will double-click. This is not an edge case — it is the normal operating condition of the internet.
Design every write operation to be idempotent from day one. Not as an afterthought. Not when bugs appear. Building idempotency in later is 10x harder than building it in from the start. The systems that get this right are the ones customers trust with their money — and their lives. 🛡️
Happy Learning! Keep Building! 🔥
Comments
Post a Comment