Unique ID Generation in Distributed Systems: Techniques, Trade-Offs and Scaling
Imagine you run a huge library 📚 with millions of books arriving every single day. Every book needs a unique number on its spine so no two books ever get the same tag — even if 1000 books arrive at the very same second.
That is exactly the problem that giant companies like Twitter, Uber, Discord, and Amazon face every day — billions of records being created every second, and every single one needs its own unique ID.
Twitter generates about 6,000 tweets per second. That is 6,000 unique IDs every single second — all different, all in the correct order, all without any two servers accidentally giving out the same number. 🐦
Why Do We Even Need Unique IDs?
Think about your school. Every student has a roll number. Two students cannot have the same roll number — the teacher would get confused!
In computer systems, every piece of data — every tweet, every order, every user account — must have a unique ID (Identifier). This is how computers find and track things.
┌─────────────────────────────────────────────────────────────────────-┐ │ WITHOUT Unique IDs — 😱 CHAOS! │ │ │ │ User A posts tweet ──► ID = 101 ◄── Same ID! 💥 │ │ User B posts tweet ──► ID = 101 ◄── Same ID! 💥 │ │ │ │ Result: System cannot tell which tweet is which! Everything │ │ breaks. Users lose data. Company loses money. │ ├───────────────────────────────────────────────────────────────────── ┤ │ WITH Unique IDs — ✅ ORDER! │ │ │ │ User A posts tweet ──► ID = 12839102019201 │ │ User B posts tweet ──► ID = 12839102019202 ← Different! ✅ │ │ │ │ Result: Every record is perfectly trackable. System works. 🎉 │ └─────────────────────────────────────────────────────────────────────-┘
🏆 What Makes a "Good" Unique ID?
Not all IDs are created equal! A great unique ID in modern systems must have these qualities:
- 🔢 Globally Unique — No two IDs anywhere on the planet should ever match
- ⚡ Sortable by Time — Newer IDs should be "bigger" so you can sort records by time
- 🚀 Fast to Generate — Generating an ID should take microseconds, not seconds
- 📏 Compact — Shorter IDs save database space and speed up searches
- 🌐 Works Without a Central Boss — Multiple servers should generate IDs independently
- 🛡️ No Security Leak — IDs should not reveal sensitive info like creation time or count
The golden rule of system design: "Your ID system is only as good as its worst failure mode."
An ID system that fails under heavy load, or generates duplicates even once, can bring down an entire company's production system. 🚨
🎲 Solution 1: UUID — The Random Number Giant
UUID stands for Universally Unique Identifier. Think of it as a super long random lottery ticket number — so long that the chance of two people getting the same number is basically impossible.
A UUID looks like this: 550e8400-e29b-41d4-a716-446655440000
It has 128 bits of data (think: 128 coin flips put together). There are 340 undecillion possible UUIDs — that is 340 followed by 36 zeros! You could generate a billion UUIDs every second for 100 years and still not repeat one.
📋 UUID Versions — Which One to Use?
There are 8 versions of UUID. Think of them like different flavours of ice cream 🍦 — each made differently, but all are ice cream (all are 128-bit unique identifiers).
| Version | How it Works | Best For | Status |
|---|---|---|---|
| v1 | Uses current time + MAC address of your computer | When you need sortable IDs | ⚠️ Privacy risk (leaks MAC) |
| v4 | Completely random — like rolling dice 122 times | General purpose, most popular | ✅ Widely used |
| v7 ⭐ NEW | Unix timestamp (ms) + random bits — sortable! | Modern databases, APIs | Recommended |
| v8 ⭐ NEW | Custom-defined by you — flexible format | Specialized systems | 🚀 Emerging standard |
The IETF officially standardised UUID v7 in RFC 9562 (2024). Use UUID v7 for all new projects — it is both random AND time-sortable. UUID v4 is still fine for cases where you don't need sorting.
This code generates a UUID in Python — think of it as pressing a button that spits out a unique lottery ticket number instantly. We show UUID v4 (pure random) and UUID v7 (time-based, sortable — the modern choice). The
import uuid part is like opening your toolbox to get the right tool.
# ── UUID v4: Pure Random UUID ──────────────────────────────────
import uuid
# Generate a v4 UUID (completely random, no time info)
my_id = uuid.uuid4()
print(f"UUID v4: {my_id}")
# Example output: 550e8400-e29b-41d4-a716-446655440000
# ── UUID v7: Time-Sortable UUID (Python 3.13+) ─────────────────
# uuid7 is in Python 3.13+ OR use the 'uuid7' library
# pip install uuid7
from uuid7 import uuid7
id1 = uuid7()
id2 = uuid7()
# id2 will ALWAYS be "greater than" id1 because it was made later.
# This means when you sort your database, newest records come last!
print(f"UUID v7 #1: {id1}")
print(f"UUID v7 #2: {id2}")
print(f"Is id2 newer? {id2 > id1}") # True ✅
- No central server needed — generate on any machine
- Works offline — no network call required
- Practically zero chance of collision (duplicate)
- Supported natively in every language & DB
- UUID v7 is sortable — great for databases
- UUID v4 is NOT sortable (random order in DB = slow)
- 128 bits = bigger than a simple integer (more storage)
- Not human-readable — hard to remember or type
- UUID v1 leaks your network MAC address (privacy risk)
🎟️ Solution 2: Ticket Server — The Queue at a Bakery
Imagine a busy bakery 🥐. When you walk in, you pull a numbered ticket from a dispenser. The machine gives out: 1, 2, 3, 4, 5... in order, one by one. No two people ever get the same number.
A Ticket Server works exactly like that machine. It is one central database that hands out the next number every time someone asks. Flickr (Yahoo's photo site) famously used this approach for billions of photos.
Ticket Server in Action
Order Service
User Service
Payment Svc
SERVER
↑ Each app server asks the Ticket Server for the next ID. The counter ticks up: 1001 → 1002 → 1003 → always unique, always in order.
🔧 How Flickr's Ticket Server Actually Worked
Flickr used a clever SQL trick to make their ticket server fast and safe. Let's understand what this SQL does before reading it:
The SQL below creates a special table called
Tickets64 that acts like the
ticket-dispensing machine at the bakery.
Every time you run
REPLACE INTO + LAST_INSERT_ID(),
the database automatically increments a counter by 1 and returns the new number — safely!
Even if 100 servers ask at the same time, they each get a different number.
It is like 100 people pressing the ticket machine button simultaneously — each gets their own slip.
-- Step 1: Create the "ticket machine" table
-- This table has only ONE row, but that row's id keeps climbing!
CREATE TABLE Tickets64 (
id BIGINT(20) UNSIGNED NOT NULL AUTO_INCREMENT,
stub CHAR(1) NOT NULL DEFAULT '',
PRIMARY KEY (id),
UNIQUE KEY stub (stub)
) ENGINE=InnoDB;
-- ─────────────────────────────────────────────────────────────
-- Step 2: Every time an app server needs a new unique ID,
-- it runs these two lines.
--
-- REPLACE INTO: If the row exists, delete it and re-insert it.
-- MySQL auto-increments the id each time!
-- LAST_INSERT_ID: Returns the freshly incremented id.
-- ─────────────────────────────────────────────────────────────
REPLACE INTO Tickets64 (stub) VALUES ('a');
SELECT LAST_INSERT_ID();
-- Returns: 1001 (then 1002, then 1003... always unique!)
But wait — what if the Ticket Server goes down? That is the big problem. If the one machine breaks, nobody can create any new records. It is a Single Point of Failure (SPOF).
If that one database machine crashes — and it will crash someday — your entire system stops dead. No new IDs = no new orders, no new users, no new anything. This is called a Single Point of Failure and it is every engineer's nightmare. 😱
Flickr's solution? They used two ticket servers and split the work between them. One server generates only even numbers (2, 4, 6, 8…) and the other generates only odd numbers (1, 3, 5, 7…). If one goes down, the other keeps working!
-- ── TICKET SERVER 1 — Generates EVEN numbers ────────────────
-- auto_increment_offset = 1 means: start at 1
-- auto_increment_increment = 2 means: jump by 2 each time
-- So this server gives: 1, 3, 5, 7, 9 ... (odd numbers)
SET @@auto_increment_offset = 1;
SET @@auto_increment_increment = 2;
-- ── TICKET SERVER 2 — Generates ODD numbers ────────────────
-- Starts at 2, jumps by 2 each time.
-- So this server gives: 2, 4, 6, 8, 10 ... (even numbers)
SET @@auto_increment_offset = 2;
SET @@auto_increment_increment = 2;
-- Together, both servers cover ALL numbers with zero overlap!
-- Server 1: 1 3 5 7 9 11 13 ...
-- Server 2: 2 4 6 8 10 12 14 ...
-- Combined: 1 2 3 4 5 6 7 8 9 10 ... ✅
🌐 Solution 3: Multi-Master Replication — Many Bosses, One Rule
What if, instead of one ticket machine (the Ticket Server), you had many machines each generating IDs — but they all follow a simple rule so they never clash?
That is Multi-Master Replication. Think of it like multiple cashiers at a supermarket 🏪. Cashier 1 serves customers 1, 4, 7, 10… Cashier 2 serves customers 2, 5, 8, 11… Cashier 3 serves customers 3, 6, 9, 12… Each cashier has their own lane, and they never collide!
📐 The Math Behind Multi-Master
The formula is beautifully simple. If you have N masters, each master is assigned a unique offset number from 1 to N. Each master only generates IDs in its own lane — like runners on a track.
Formula: next_id = (last_id) + (number_of_masters)
With 3 Masters:
┌──────────────────────────────────────────────────────┐
│ Master 1: starts at 1 → 1, 4, 7, 10, 13, 16 ... │
│ Master 2: starts at 2 → 2, 5, 8, 11, 14, 17 ... │
│ Master 3: starts at 3 → 3, 6, 9, 12, 15, 18 ... │
├──────────────────────────────────────────────────────┤
│ No two masters EVER generate the same number! ✅ │
│ If Master 2 dies, Master 1 & 3 keep going! ✅ │
└──────────────────────────────────────────────────────┘
If you add a 4th master later? PROBLEM!
You have to reconfigure ALL masters. 😬
(This is the main weakness of this approach)
Adding a new master means reconfiguring every existing server — this causes downtime and headaches. It is a fixed-topology solution, not a flexible one.
❄️ Solution 4: Twitter Snowflake — The Engineering Masterpiece
In 2010, Twitter engineers faced a massive problem: MySQL auto-increment IDs could not keep up with millions of tweets per minute. They invented Snowflake — a distributed ID system so elegant and clever that nearly every major company (Discord, Instagram, Sony, Mastodon) copied the idea.
The Snowflake idea is like a clever coded barcode 🔢. Each ID is a 64-bit integer, but inside those 64 bits, it hides three pieces of information at once: when it was created, which machine created it, and a sequence number in case two IDs are born in the same millisecond.
🔍 Breaking Down a Real Snowflake ID
Let's decode a real Snowflake ID step by step. Think of it like cracking open a birthday cracker 🎉 and reading the message inside!
Snowflake ID: 1541815603606036480
Step 1 — Convert to binary (64 bits):
0 | 00010101011000101101111001011101001110 | 0000010101 | 000000000000
Step 2 — Split the sections:
┌─────────────────────────────────────────────────────────────────────┐
│ Sign │ Timestamp (41 bits) │ Machine (10b) │ Sequence (12b)│
│ 0 │ ...milliseconds since 2010 │ 0000010101 │ 000000000000 │
└─────────────────────────────────────────────────────────────────────┘
Step 3 — Decode the timestamp:
Twitter's epoch start = November 4, 2010 (they picked this date!)
Timestamp bits → 1288834974657 ms since then
Convert to date → Created on: Dec 3, 2021, 14:30:22 UTC
Step 4 — Decode the Machine ID:
10 bits = 5 bits datacenter + 5 bits worker
Machine ID = 21 → Datacenter 2, Worker 5
Step 5 — Sequence number = 0 (first ID in that millisecond)
💻 Building Your Own Snowflake Generator in Python
The Python class below is a complete Snowflake ID generator. Imagine you are building your own Twitter! This code creates a machine that stamps a unique number onto every new tweet as fast as lightning — up to 4,096 IDs per millisecond from a single machine.
Key parts explained simply:
__init__= This sets up the machine — gives it a name (machine_id) and a starting positiongenerate()= This is the "stamp button" — call it whenever you need a new unique IDint(time.time() * 1000)= Gets the current time in milliseconds (very precise!)- Bit shifting (
<<) = Like sliding puzzle pieces into specific slots in a 64-bit box threading.Lock()= Like a bathroom lock 🚽 — only ONE person at a time can use it
import time
import threading
class SnowflakeIDGenerator:
"""
A Snowflake ID Generator.
It produces unique 64-bit integer IDs.
Each ID encodes: timestamp + machine_id + sequence number.
Just like Twitter's original design — but yours to own! ❄️
"""
# ── Constants (bit counts for each section) ───────────────
EPOCH = 1288834974657 # Twitter's custom epoch (Nov 4, 2010)
MACHINE_BITS = 10 # 10 bits → up to 1024 machines
SEQUENCE_BITS = 12 # 12 bits → 4096 IDs per millisecond
MAX_MACHINE_ID = (1 << MACHINE_BITS) - 1 # = 1023
MAX_SEQUENCE = (1 << SEQUENCE_BITS) - 1 # = 4095
def __init__(self, machine_id: int):
# Make sure the machine_id is valid (0 to 1023)
if not (0 <= machine_id <= self.MAX_MACHINE_ID):
raise ValueError(f"machine_id must be between 0 and {self.MAX_MACHINE_ID}")
self.machine_id = machine_id
self.sequence = 0 # Starts at 0 each millisecond
self.last_timestamp = -1 # Remembers when the last ID was made
self.lock = threading.Lock() # Safety lock for multi-threading
def _current_millis(self) -> int:
"""Get current time in milliseconds since our custom epoch."""
return int(time.time() * 1000) - self.EPOCH
def generate(self) -> int:
"""
Generate and return the next unique Snowflake ID.
Thread-safe: you can call this from many threads at once safely.
"""
with self.lock: # 🔒 Only one thread can run this block at a time
now = self._current_millis()
# ── If we are in the same millisecond as last time ──────
if now == self.last_timestamp:
self.sequence = (self.sequence + 1) & self.MAX_SEQUENCE
if self.sequence == 0:
# We have used all 4096 IDs in this ms!
# Wait for the next millisecond before continuing.
while now <= self.last_timestamp:
now = self._current_millis()
# ── New millisecond — reset sequence to 0 ──────────────
else:
self.sequence = 0
self.last_timestamp = now
# ── Assemble the final 64-bit ID ────────────────────────
# Think of this like sliding puzzle pieces into exact slots:
# [41-bit time] shifted left by 22 positions
# [10-bit machine_id] shifted left by 12 positions
# [12-bit sequence] stays in place
snowflake_id = (
(now << (self.MACHINE_BITS + self.SEQUENCE_BITS)) |
(self.machine_id << self.SEQUENCE_BITS) |
self.sequence
)
return snowflake_id
# ── DEMO: Let's generate some IDs! ─────────────────────────────
if __name__ == "__main__":
# Create a generator for machine #1
gen = SnowflakeIDGenerator(machine_id=1)
print("🎉 Generating 5 Snowflake IDs:\n")
for i in range(5):
new_id = gen.generate()
print(f" ID #{i+1}: {new_id}")
time.sleep(0.001) # Tiny pause between IDs
# They are always increasing! Sort them and you get them in time order.
# Example output:
# ID #1: 1843921740001234944
# ID #2: 1843921740001234945
# ID #3: 1843921740005429248
# ...
🔄 How Snowflake Handles the Clock Going Backwards
A sneaky problem: what if your server's clock jumps backward? (This can happen after NTP sync.) Suddenly the timestamp is smaller than before — and you could generate a duplicate ID! 😱
# ── Safe Clock-Backward Handler ──────────────────────────────
# Add this check inside your generate() method after getting "now"
if now < self.last_timestamp:
# The clock went backward! This is dangerous.
# How much did it go back?
clock_drift = self.last_timestamp - now
if clock_drift <= 5:
# Tiny drift (≤5 ms): just wait it out.
# It will self-correct in a few milliseconds.
time.sleep(clock_drift / 1000.0)
now = self._current_millis()
else:
# Large drift: something is seriously wrong.
# Raise an error rather than generate a bad ID.
raise RuntimeError(
f"Clock moved backwards by {clock_drift}ms! "
f"Cannot safely generate ID. Check your NTP config."
)
🌍 How Companies Customised Snowflake for Themselves
Snowflake is an idea, not a strict standard. Every company that adopted it tweaked the bit layout to suit their needs. This is like a recipe — the base is the same, but the seasoning differs!
| Company | System Name | Key Difference | Bits |
|---|---|---|---|
| Snowflake | Original design. Epoch = Nov 2010 | 1+41+10+12 | |
| Insta64 | Uses PostgreSQL; shard ID instead of machine | 41+13+10 | |
| 🎮 Discord | Snowflake | Epoch = Jan 1, 2015. Process ID in machine bits | 1+41+10+12 |
| 🍔 Baidu | UidGenerator | Future timestamps pre-cached in a ring buffer | 1+28+22+13 |
| ☁️ Meituan | Leaf | Dual mode: Snowflake + Segment (Ticket Server) | 1+41+10+12 |
| 🦣 Mastodon | Flake | Open source; Ruby impl; per-worker sequence | 1+41+10+12 |
🔄 End-to-End Flow — How a Tweet Gets Its ID (Step by Step)
Let's trace exactly what happens when you hit "Post" on Twitter. Follow each step like a story!
📊 Side-by-Side Comparison — Which One Should You Pick?
Here is a quick cheat-sheet. Think of this as the menu at a restaurant 🍽️ — you choose based on what your appetite (system requirement) is.
| Feature | UUID v4 | UUID v7 | Ticket Server | Multi-Master | ❄️ Snowflake |
|---|---|---|---|---|---|
| Globally Unique | ✅ Yes | ✅ Yes | ✅ Yes | ✅ Yes | ✅ Yes |
| Time-Sortable | ❌ No | ✅ Yes | ✅ Yes | ✅ Yes | ✅ Yes |
| No Single Point of Failure | ✅ Yes | ✅ Yes | ❌ No | ✅ Yes | ✅ Yes |
| Works Without Network | ✅ Yes | ✅ Yes | ❌ No | ❌ No | ✅ Yes* |
| Compact (fits in 64 bits) | ❌ 128 bits | ❌ 128 bits | ✅ 64 bits | ✅ 64 bits | ✅ 64 bits |
| Scales to 1000s of servers | ✅ Yes | ✅ Yes | ⚠️ Limited | ⚠️ Fixed | ✅ Yes (1024) |
| IDs per second (per node) | Millions | Millions | ~10K | ~10K/node | 4M/node |
| Complexity | 🟢 Very Low | 🟢 Very Low | 🟢 Low | 🟡 Medium | 🔴 High |
| Best For | Simple apps | Modern apps | Small systems | Medium systems | Large scale |
* Snowflake works locally once deployed — no network needed to generate IDs after setup.
🌳 Decision Tree — Which ID System is Right for You?
Still not sure which one to use? Follow this decision tree like a flowchart! It is like choosing which ride to take at an amusement park 🎡.
Start Here: Do you need more than 100 IDs per second?
│
├── NO → Is your app a single server or small startup?
│ │
│ ├── YES → Use a simple DB Auto-Increment (INTEGER PRIMARY KEY)
│ │ It is the simplest and most supported option. ✅
│ │
│ └── NO → Use UUID v7 (time-sortable, easy to use everywhere) ✅
│
└── YES → Do you need IDs to be sortable by time (newest = largest)?
│
├── NO → Do you need compact IDs (64-bit, not 128-bit)?
│ │
│ ├── NO → Use UUID v4 (simple, widely supported) ✅
│ └── YES → Use Ticket Server (easy, sequential) ✅
│
└── YES → Will you ever need more than 2 database nodes?
│
├── NO → Use Multi-Master Replication (simple, no extra service) ✅
│
└── YES → Are you handling MILLIONS of events per second?
│
├── NO → Use UUID v7 (easiest sortable option) ✅
└── YES → Use Snowflake or a Snowflake variant ✅ 🏆
🚀 Trend — Modern Alternatives You Should Know
The ID generation world kept evolving. These are the alternatives that have emerged alongside the classic four approaches:
-
🔶 ULID (Universally Unique Lexicographically Sortable Identifier)
Like UUID v7 but encoded in a shorter, more readable base32 string. Example:01ARZ3NDEKTSV4RRFFQ69G5FAV— looks like a password, sorts by time, works offline. -
🔷 KSUID (K-Sortable Globally Unique IDs)
Created by Segment. 20-byte IDs: 4 bytes timestamp + 16 bytes random. Example:0ujtsYcgvSTl8PAuAdqWYSMnLOv. Sorts naturally like timestamps. -
🔹 NanoID
Tiny library generating URL-safe, compact IDs. Example:V1StGXR8_Z5jdHi6B-myT. Great for URLs where shorter = better. -
⭐ CUID2
Security-first IDs resistant to fingerprinting. Designed for web apps and APIs. Replaces CUID which was deprecated in 2024.
| Format | Example ID | Sortable | Size | Best For |
|---|---|---|---|---|
| UUID v7 | 018e–c8f0–7000–… | ✅ | 128 bits | Databases, APIs (standard) |
| ULID | 01ARZ3NDEKTSV4RRFFQ69G5FAV | ✅ | 128 bits | Readable IDs, logs |
| KSUID | 0ujtsYcgvSTl8PAuAdqW | ✅ | 160 bits | Events, analytics streams |
| NanoID | V1StGXR8_Z5jdHi6B | ❌ | ~126 bits | URLs, short codes |
| Snowflake | 1541815603606036480 | ✅ | 64 bits | High-throughput distributed systems |
⚠️ Common Mistakes Engineers Make with ID Generators
UUID v4 is completely random. When you insert rows with random primary keys, the database B-tree index fragments badly — like shuffling a sorted deck of cards with every new card. This causes page splits and slows down inserts drastically at scale. Fix: Use UUID v7 (time-sorted) or Snowflake IDs instead.
If your server's clock goes backward (even by 1 millisecond due to NTP sync), your Snowflake generator can produce duplicate IDs or throw an exception. Always add a clock-backward guard in your implementation! See the code example above. 🛡️
If two processes on the same machine both use machine_id=1 in Snowflake, they will collide when generating IDs at the same millisecond. Always assign a unique machine_id to every process, not just every physical machine.
If your user ID goes 1, 2, 3… anyone can guess that user #100001 exists. Competitors can estimate your user growth by signing up twice and comparing IDs. This is called ID enumeration attack. Use UUIDs or Snowflake IDs for anything user-visible.
- Use UUID v7 as your default choice for new projects — it is now an RFC standard
- Use Snowflake when you need 64-bit IDs and millions per second at scale
- Never share machine IDs across processes in a Snowflake setup
- Always have a clock-skew guard in your Snowflake implementation
- Use two Ticket Servers (odd/even) if you go that route — avoid SPOF
- Store Snowflake IDs as
BIGINT(not VARCHAR) in databases for efficiency - Do NOT use sequential IDs as publicly visible identifiers — security risk!
🧠 Quick Recap — The 4 Solutions at a Glance
┌──────────────────────────────────────────────────────────────────────┐ │ SOLUTION │ ANALOGY │ BEST SCALE │ COMPLEXITY │ ├──────────────────────────────────────────────────────────────────────┤ │ UUID v4 │ Random lottery │ Any scale │ Super easy │ │ UUID v7 │ Dated lottery │ Any scale │ Super easy │ │ Ticket Server │ Bakery queue │ Small/Med │ Easy │ │ Multi-Master │ Many cashiers │ Medium │ Medium │ │ Twitter Snowflake ❄️ │ Coded barcode │ Massive │ Advanced │ └──────────────────────────────────────────────────────────────────────┘ The golden rule of ID design: Start simple → UUID v7 for most apps. Go complex → Snowflake only when you TRULY need billions per day.
❄️ Happy Generating! 🎟️ 🌐 🔢
Comments
Post a Comment