Skip to main content

Unique ID Generation in Distributed Systems: Techniques, Trade-Offs and Scaling

Calculating read time…

Imagine you run a huge library 📚 with millions of books arriving every single day. Every book needs a unique number on its spine so no two books ever get the same tag — even if 1000 books arrive at the very same second.

That is exactly the problem that giant companies like Twitter, Uber, Discord, and Amazon face every day — billions of records being created every second, and every single one needs its own unique ID.

💡 Fun Fact!
Twitter generates about 6,000 tweets per second. That is 6,000 unique IDs every single second — all different, all in the correct order, all without any two servers accidentally giving out the same number. 🐦



 Why Do We Even Need Unique IDs?

Think about your school. Every student has a roll number. Two students cannot have the same roll number — the teacher would get confused!

In computer systems, every piece of data — every tweet, every order, every user account — must have a unique ID (Identifier). This is how computers find and track things.

  ┌─────────────────────────────────────────────────────────────────────-┐
  │  WITHOUT Unique IDs — 😱 CHAOS!                                      │
  │                                                                      │
  │  User A posts tweet  ──►  ID = 101   ◄── Same ID! 💥                 │
  │  User B posts tweet  ──►  ID = 101   ◄── Same ID! 💥                 │
  │                                                                      │
  │  Result: System cannot tell which tweet is which! Everything         │
  │          breaks. Users lose data. Company loses money.               │ 
  ├───────────────────────────────────────────────────────────────────── ┤
  │  WITH Unique IDs — ✅ ORDER!                                         │
  │                                                                      │
  │  User A posts tweet  ──►  ID = 12839102019201                        │
  │  User B posts tweet  ──►  ID = 12839102019202   ← Different! ✅      │
  │                                                                      │
  │  Result: Every record is perfectly trackable. System works. 🎉       │
  └─────────────────────────────────────────────────────────────────────-┘

🏆 What Makes a "Good" Unique ID?

Not all IDs are created equal! A great unique ID in modern systems must have these qualities:

  • 🔢 Globally Unique — No two IDs anywhere on the planet should ever match
  • ⚡ Sortable by Time — Newer IDs should be "bigger" so you can sort records by time
  • 🚀 Fast to Generate — Generating an ID should take microseconds, not seconds
  • 📏 Compact — Shorter IDs save database space and speed up searches
  • 🌐 Works Without a Central Boss — Multiple servers should generate IDs independently
  • 🛡️ No Security Leak — IDs should not reveal sensitive info like creation time or count
✅ Remember This Rule:
The golden rule of system design: "Your ID system is only as good as its worst failure mode."
An ID system that fails under heavy load, or generates duplicates even once, can bring down an entire company's production system. 🚨

🎲 Solution 1: UUID — The Random Number Giant

UUID stands for Universally Unique Identifier. Think of it as a super long random lottery ticket number — so long that the chance of two people getting the same number is basically impossible.

A UUID looks like this: 550e8400-e29b-41d4-a716-446655440000

It has 128 bits of data (think: 128 coin flips put together). There are 340 undecillion possible UUIDs — that is 340 followed by 36 zeros! You could generate a billion UUIDs every second for 100 years and still not repeat one.

 How UUID is Built (128 bits = 32 hex characters)

TIME_LOW
550e8400
32 bits
-
TIME_MID
e29b
16 bits
-
VERSION+TIME
41d4
16 bits
-
CLOCK SEQ
a716
16 bits
-
NODE (MAC)
446655440000
48 bits
✨ FINAL UUID (128 bits total):
550e8400-e29b-41d4-a716-446655440000

📋 UUID Versions — Which One to Use?

There are 8 versions of UUID. Think of them like different flavours of ice cream 🍦 — each made differently, but all are ice cream (all are 128-bit unique identifiers).

Version How it Works Best For Status
v1 Uses current time + MAC address of your computer When you need sortable IDs ⚠️ Privacy risk (leaks MAC)
v4 Completely random — like rolling dice 122 times General purpose, most popular ✅ Widely used
v7 ⭐ NEW Unix timestamp (ms) + random bits — sortable! Modern databases, APIs Recommended
v8 ⭐ NEW Custom-defined by you — flexible format Specialized systems 🚀 Emerging standard
💡 Tip:
The IETF officially standardised UUID v7 in RFC 9562 (2024). Use UUID v7 for all new projects — it is both random AND time-sortable. UUID v4 is still fine for cases where you don't need sorting.
📝 What the code below does (before you read it):

This code generates a UUID in Python — think of it as pressing a button that spits out a unique lottery ticket number instantly. We show UUID v4 (pure random) and UUID v7 (time-based, sortable — the modern choice). The import uuid part is like opening your toolbox to get the right tool.

# ── UUID v4: Pure Random UUID ──────────────────────────────────
import uuid

# Generate a v4 UUID (completely random, no time info)
my_id = uuid.uuid4()
print(f"UUID v4: {my_id}")
# Example output: 550e8400-e29b-41d4-a716-446655440000

# ── UUID v7: Time-Sortable UUID (Python 3.13+) ─────────────────
# uuid7 is in Python 3.13+  OR  use the 'uuid7' library

# pip install uuid7
from uuid7 import uuid7

id1 = uuid7()
id2 = uuid7()

# id2 will ALWAYS be "greater than" id1 because it was made later.
# This means when you sort your database, newest records come last!
print(f"UUID v7 #1: {id1}")
print(f"UUID v7 #2: {id2}")
print(f"Is id2 newer? {id2 > id1}")  # True ✅

✅ UUID Pros (DO use UUID when...)
  • No central server needed — generate on any machine
  • Works offline — no network call required
  • Practically zero chance of collision (duplicate)
  • Supported natively in every language & DB
  • UUID v7 is sortable — great for databases
❌ UUID Cons (DON'T use UUID when...)
  • UUID v4 is NOT sortable (random order in DB = slow)
  • 128 bits = bigger than a simple integer (more storage)
  • Not human-readable — hard to remember or type
  • UUID v1 leaks your network MAC address (privacy risk)

🎟️ Solution 2: Ticket Server — The Queue at a Bakery

Imagine a busy bakery 🥐. When you walk in, you pull a numbered ticket from a dispenser. The machine gives out: 1, 2, 3, 4, 5... in order, one by one. No two people ever get the same number.

A Ticket Server works exactly like that machine. It is one central database that hands out the next number every time someone asks. Flickr (Yahoo's photo site) famously used this approach for billions of photos.

Ticket Server in Action

🖥️ App Server 1
Order Service
🖥️ App Server 2
User Service
🖥️ App Server 3
Payment Svc
⇄
🎟️
TICKET
SERVER
MySQL / Postgres
ID: 1001
ID: 1002
ID: 1003
ID: 1004
ID: 1005

↑ Each app server asks the Ticket Server for the next ID. The counter ticks up: 1001 → 1002 → 1003 → always unique, always in order.

🔧 How Flickr's Ticket Server Actually Worked

Flickr used a clever SQL trick to make their ticket server fast and safe. Let's understand what this SQL does before reading it:

📝 What the code below does (read this first!):

The SQL below creates a special table called Tickets64 that acts like the ticket-dispensing machine at the bakery.

Every time you run REPLACE INTO + LAST_INSERT_ID(), the database automatically increments a counter by 1 and returns the new number — safely! Even if 100 servers ask at the same time, they each get a different number. It is like 100 people pressing the ticket machine button simultaneously — each gets their own slip.

-- Step 1: Create the "ticket machine" table
-- This table has only ONE row, but that row's id keeps climbing!
CREATE TABLE Tickets64 (
  id          BIGINT(20) UNSIGNED NOT NULL  AUTO_INCREMENT,
  stub        CHAR(1)    NOT NULL DEFAULT '',
  PRIMARY KEY (id),
  UNIQUE KEY stub (stub)
) ENGINE=InnoDB;

-- ─────────────────────────────────────────────────────────────
-- Step 2: Every time an app server needs a new unique ID,
--         it runs these two lines.
--
-- REPLACE INTO: If the row exists, delete it and re-insert it.
--               MySQL auto-increments the id each time!
-- LAST_INSERT_ID: Returns the freshly incremented id.
-- ─────────────────────────────────────────────────────────────
REPLACE INTO Tickets64 (stub) VALUES ('a');
SELECT LAST_INSERT_ID();
-- Returns: 1001  (then 1002, then 1003... always unique!)

But wait — what if the Ticket Server goes down?  That is the big problem. If the one machine breaks, nobody can create any new records. It is a Single Point of Failure (SPOF).

❌ DON'T use a single Ticket Server in production without a backup plan!
If that one database machine crashes — and it will crash someday — your entire system stops dead. No new IDs = no new orders, no new users, no new anything. This is called a Single Point of Failure and it is every engineer's nightmare. 😱

Flickr's solution? They used two ticket servers and split the work between them. One server generates only even numbers (2, 4, 6, 8…) and the other generates only odd numbers (1, 3, 5, 7…). If one goes down, the other keeps working!


-- ── TICKET SERVER 1 — Generates EVEN numbers ────────────────
-- auto_increment_offset  = 1 means: start at 1
-- auto_increment_increment = 2 means: jump by 2 each time
-- So this server gives: 1, 3, 5, 7, 9 ...  (odd numbers)
SET @@auto_increment_offset     = 1;
SET @@auto_increment_increment  = 2;

-- ── TICKET SERVER 2 — Generates ODD numbers  ────────────────
-- Starts at 2, jumps by 2 each time.
-- So this server gives: 2, 4, 6, 8, 10 ... (even numbers)
SET @@auto_increment_offset     = 2;
SET @@auto_increment_increment  = 2;

-- Together, both servers cover ALL numbers with zero overlap!
-- Server 1:  1  3  5  7  9  11  13 ...
-- Server 2:  2  4  6  8  10  12  14 ...
-- Combined:  1  2  3  4  5  6  7  8  9  10 ... ✅


🌐 Solution 3: Multi-Master Replication — Many Bosses, One Rule

What if, instead of one ticket machine (the Ticket Server), you had many machines each generating IDs — but they all follow a simple rule so they never clash?

That is Multi-Master Replication. Think of it like multiple cashiers at a supermarket 🏪. Cashier 1 serves customers 1, 4, 7, 10… Cashier 2 serves customers 2, 5, 8, 11… Cashier 3 serves customers 3, 6, 9, 12… Each cashier has their own lane, and they never collide!

3 Masters Generating IDs Without Collision

🖥️
Master 1
offset=1, step=3
→ 1, 4, 7, 10…
🖥️
Master 2
offset=2, step=3
→ 2, 5, 8, 11…
🖥️
Master 3
offset=3, step=3
→ 3, 6, 9, 12…
🔀 Combined Output (all unique, no gaps):
1   2   3   4   5   6   7   8   9   10   11   12 …

📐 The Math Behind Multi-Master

The formula is beautifully simple. If you have N masters, each master is assigned a unique offset number from 1 to N. Each master only generates IDs in its own lane — like runners on a track.


  Formula: next_id = (last_id) + (number_of_masters)

  With 3 Masters:
  ┌──────────────────────────────────────────────────────┐
  │  Master 1: starts at 1  → 1,  4,  7, 10, 13, 16 ... │
  │  Master 2: starts at 2  → 2,  5,  8, 11, 14, 17 ... │
  │  Master 3: starts at 3  → 3,  6,  9, 12, 15, 18 ... │
  ├──────────────────────────────────────────────────────┤
  │  No two masters EVER generate the same number! ✅     │
  │  If Master 2 dies, Master 1 & 3 keep going! ✅       │
  └──────────────────────────────────────────────────────┘

  If you add a 4th master later? PROBLEM!
  You have to reconfigure ALL masters. 😬
  (This is the main weakness of this approach)

❌ DON'T use Multi-Master if you need to scale up by adding new servers dynamically.
Adding a new master means reconfiguring every existing server — this causes downtime and headaches. It is a fixed-topology solution, not a flexible one.
✅ DO use Multi-Master if your number of database servers is fixed and stable. It is simpler than Snowflake, works well for small-to-medium scale, and eliminates the single point of failure problem from the basic Ticket Server.

❄️ Solution 4: Twitter Snowflake — The Engineering Masterpiece

In 2010, Twitter engineers faced a massive problem: MySQL auto-increment IDs could not keep up with millions of tweets per minute. They invented Snowflake — a distributed ID system so elegant and clever that nearly every major company (Discord, Instagram, Sony, Mastodon) copied the idea.

The Snowflake idea is like a clever coded barcode 🔢. Each ID is a 64-bit integer, but inside those 64 bits, it hides three pieces of information at once: when it was created, which machine created it, and a sequence number in case two IDs are born in the same millisecond.

Snowflake ID Being Built (64 bits total)

SIGN BIT
0
1 bit
Always 0
(positive)
TIMESTAMP
41 bits
milliseconds since epoch
~69 years of IDs!
Newest IDs sort last ✅
MACHINE ID
10 bits
datacenter + worker
Up to 1024 machines
at once! ✅
SEQUENCE
12 bits
per-machine counter
4096 IDs/ms
per machine! ✅
📦 Example Snowflake ID (as a 64-bit integer):
1541815603606036480

41 bits of time + 10 bits of machine ID + 12 bits of sequence = 63 bits of pure power ⚡

🔍 Breaking Down a Real Snowflake ID

Let's decode a real Snowflake ID step by step. Think of it like cracking open a birthday cracker 🎉 and reading the message inside!


  Snowflake ID: 1541815603606036480

  Step 1 — Convert to binary (64 bits):
  0 | 00010101011000101101111001011101001110 | 0000010101 | 000000000000

  Step 2 — Split the sections:
  ┌─────────────────────────────────────────────────────────────────────┐
  │ Sign  │ Timestamp (41 bits)         │ Machine (10b) │ Sequence (12b)│
  │  0    │  ...milliseconds since 2010 │   0000010101  │  000000000000 │
  └─────────────────────────────────────────────────────────────────────┘

  Step 3 — Decode the timestamp:
  Twitter's epoch start = November 4, 2010  (they picked this date!)
  Timestamp bits → 1288834974657 ms since then
  Convert to date → Created on: Dec 3, 2021, 14:30:22 UTC

  Step 4 — Decode the Machine ID:
  10 bits = 5 bits datacenter + 5 bits worker
  Machine ID = 21  → Datacenter 2, Worker 5

  Step 5 — Sequence number = 0 (first ID in that millisecond)

💻 Building Your Own Snowflake Generator in Python

📝 What the code below does (read this first — super important!):

The Python class below is a complete Snowflake ID generator. Imagine you are building your own Twitter! This code creates a machine that stamps a unique number onto every new tweet as fast as lightning — up to 4,096 IDs per millisecond from a single machine.

Key parts explained simply:
  • __init__ = This sets up the machine — gives it a name (machine_id) and a starting position
  • generate() = This is the "stamp button" — call it whenever you need a new unique ID
  • int(time.time() * 1000) = Gets the current time in milliseconds (very precise!)
  • Bit shifting (<<) = Like sliding puzzle pieces into specific slots in a 64-bit box
  • threading.Lock() = Like a bathroom lock 🚽 — only ONE person at a time can use it

import time
import threading

class SnowflakeIDGenerator:
    """
    A Snowflake ID Generator.

    It produces unique 64-bit integer IDs.
    Each ID encodes: timestamp + machine_id + sequence number.
    Just like Twitter's original design — but yours to own! ❄️
    """

    # ── Constants (bit counts for each section) ───────────────
    EPOCH           = 1288834974657  # Twitter's custom epoch (Nov 4, 2010)
    MACHINE_BITS    = 10             # 10 bits → up to 1024 machines
    SEQUENCE_BITS   = 12             # 12 bits → 4096 IDs per millisecond
    MAX_MACHINE_ID  = (1 << MACHINE_BITS) - 1   # = 1023
    MAX_SEQUENCE    = (1 << SEQUENCE_BITS) - 1  # = 4095

    def __init__(self, machine_id: int):
        # Make sure the machine_id is valid (0 to 1023)
        if not (0 <= machine_id <= self.MAX_MACHINE_ID):
            raise ValueError(f"machine_id must be between 0 and {self.MAX_MACHINE_ID}")

        self.machine_id     = machine_id
        self.sequence       = 0       # Starts at 0 each millisecond
        self.last_timestamp = -1      # Remembers when the last ID was made
        self.lock           = threading.Lock()  # Safety lock for multi-threading

    def _current_millis(self) -> int:
        """Get current time in milliseconds since our custom epoch."""
        return int(time.time() * 1000) - self.EPOCH

    def generate(self) -> int:
        """
        Generate and return the next unique Snowflake ID.

        Thread-safe: you can call this from many threads at once safely.
        """
        with self.lock:  # 🔒 Only one thread can run this block at a time
            now = self._current_millis()

            # ── If we are in the same millisecond as last time ──────
            if now == self.last_timestamp:
                self.sequence = (self.sequence + 1) & self.MAX_SEQUENCE
                if self.sequence == 0:
                    # We have used all 4096 IDs in this ms!
                    # Wait for the next millisecond before continuing.
                    while now <= self.last_timestamp:
                        now = self._current_millis()

            # ── New millisecond — reset sequence to 0 ──────────────
            else:
                self.sequence = 0

            self.last_timestamp = now

            # ── Assemble the final 64-bit ID ────────────────────────
            # Think of this like sliding puzzle pieces into exact slots:
            #  [41-bit time] shifted left by 22 positions
            #  [10-bit machine_id] shifted left by 12 positions
            #  [12-bit sequence] stays in place
            snowflake_id = (
                (now          << (self.MACHINE_BITS + self.SEQUENCE_BITS)) |
                (self.machine_id << self.SEQUENCE_BITS)                    |
                self.sequence
            )
            return snowflake_id


# ── DEMO: Let's generate some IDs! ─────────────────────────────
if __name__ == "__main__":
    # Create a generator for machine #1
    gen = SnowflakeIDGenerator(machine_id=1)

    print("🎉 Generating 5 Snowflake IDs:\n")
    for i in range(5):
        new_id = gen.generate()
        print(f"  ID #{i+1}: {new_id}")
        time.sleep(0.001)  # Tiny pause between IDs

    # They are always increasing! Sort them and you get them in time order.
    # Example output:
    # ID #1: 1843921740001234944
    # ID #2: 1843921740001234945
    # ID #3: 1843921740005429248
    # ...

🔄 How Snowflake Handles the Clock Going Backwards

A sneaky problem: what if your server's clock jumps backward? (This can happen after NTP sync.) Suddenly the timestamp is smaller than before — and you could generate a duplicate ID! 😱


# ── Safe Clock-Backward Handler ──────────────────────────────
# Add this check inside your generate() method after getting "now"

if now < self.last_timestamp:
    # The clock went backward! This is dangerous.
    # How much did it go back?
    clock_drift = self.last_timestamp - now

    if clock_drift <= 5:
        # Tiny drift (≤5 ms): just wait it out.
        # It will self-correct in a few milliseconds.
        time.sleep(clock_drift / 1000.0)
        now = self._current_millis()
    else:
        # Large drift: something is seriously wrong.
        # Raise an error rather than generate a bad ID.
        raise RuntimeError(
            f"Clock moved backwards by {clock_drift}ms! "
            f"Cannot safely generate ID. Check your NTP config."
        )

🌍 How Companies Customised Snowflake for Themselves

Snowflake is an idea, not a strict standard. Every company that adopted it tweaked the bit layout to suit their needs. This is like a recipe — the base is the same, but the seasoning differs!

Company System Name Key Difference Bits
🐦 Twitter Snowflake Original design. Epoch = Nov 2010 1+41+10+12
📸 Instagram Insta64 Uses PostgreSQL; shard ID instead of machine 41+13+10
🎮 Discord Snowflake Epoch = Jan 1, 2015. Process ID in machine bits 1+41+10+12
🍔 Baidu UidGenerator Future timestamps pre-cached in a ring buffer 1+28+22+13
☁️ Meituan Leaf Dual mode: Snowflake + Segment (Ticket Server) 1+41+10+12
🦣 Mastodon Flake Open source; Ruby impl; per-worker sequence 1+41+10+12

🔄 End-to-End Flow — How a Tweet Gets Its ID (Step by Step)

Let's trace exactly what happens when you hit "Post" on Twitter. Follow each step like a story!

From "Post Button" to Stored Tweet

1
User taps "Post Tweet" on their phone
Your phone sends an HTTP POST /tweet request to Twitter's load balancer. The tweet text travels over the internet to Twitter's data center.
2
Load Balancer routes to an App Server
Twitter has thousands of app servers. The load balancer picks the least-busy one and forwards your request. Think of it like a traffic cop at a busy junction! 🚦
3
App Server asks the Snowflake Service for a new ID
Twitter runs dedicated "Snowflake Servers" — machines whose only job is generating IDs. The app server makes a quick network call: "Give me the next ID!"
4
Snowflake Server generates the ID in microseconds
The Snowflake generator combines: current timestamp + its machine ID + sequence counter. The result: 1843921740001234945 — a unique 64-bit integer, done in <1 microsecond!
5
Tweet is saved to the database with the Snowflake ID as its primary key
The app server stores: {id: 1843921740001234945, text: "Hello World!", user_id: 88} into the database. This record is now permanently retrievable by its unique ID.
6
Response sent back — tweet is live! 🎉
The app server responds with the new tweet data (including its ID). Your phone receives it, shows the tweet on screen, and displays the like/retweet buttons. Total time: usually under 200 milliseconds!

📊 Side-by-Side Comparison — Which One Should You Pick?

Here is a quick cheat-sheet. Think of this as the menu at a restaurant 🍽️ — you choose based on what your appetite (system requirement) is.

Feature UUID v4 UUID v7 Ticket Server Multi-Master ❄️ Snowflake
Globally Unique ✅ Yes ✅ Yes ✅ Yes ✅ Yes ✅ Yes
Time-Sortable ❌ No ✅ Yes ✅ Yes ✅ Yes ✅ Yes
No Single Point of Failure ✅ Yes ✅ Yes ❌ No ✅ Yes ✅ Yes
Works Without Network ✅ Yes ✅ Yes ❌ No ❌ No ✅ Yes*
Compact (fits in 64 bits) ❌ 128 bits ❌ 128 bits ✅ 64 bits ✅ 64 bits ✅ 64 bits
Scales to 1000s of servers ✅ Yes ✅ Yes ⚠️ Limited ⚠️ Fixed ✅ Yes (1024)
IDs per second (per node) Millions Millions ~10K ~10K/node 4M/node
Complexity 🟢 Very Low 🟢 Very Low 🟢 Low 🟡 Medium 🔴 High
Best For Simple apps Modern apps Small systems Medium systems Large scale

* Snowflake works locally once deployed — no network needed to generate IDs after setup.


🌳 Decision Tree — Which ID System is Right for You?

Still not sure which one to use? Follow this decision tree like a flowchart! It is like choosing which ride to take at an amusement park 🎡.

Start Here: Do you need more than 100 IDs per second?
│
├── NO  → Is your app a single server or small startup?
│          │
│          ├── YES → Use a simple DB Auto-Increment (INTEGER PRIMARY KEY)
│          │          It is the simplest and most supported option. ✅
│          │
│          └── NO  → Use UUID v7 (time-sortable, easy to use everywhere) ✅
│
└── YES → Do you need IDs to be sortable by time (newest = largest)?
           │
           ├── NO  → Do you need compact IDs (64-bit, not 128-bit)?
           │          │
           │          ├── NO  → Use UUID v4 (simple, widely supported) ✅
           │          └── YES → Use Ticket Server (easy, sequential) ✅
           │
           └── YES → Will you ever need more than 2 database nodes?
                      │
                      ├── NO  → Use Multi-Master Replication (simple, no extra service) ✅
                      │
                      └── YES → Are you handling MILLIONS of events per second?
                                 │
                                 ├── NO  → Use UUID v7 (easiest sortable option) ✅
                                 └── YES → Use Snowflake or a Snowflake variant ✅ 🏆

🚀 Trend — Modern Alternatives You Should Know

The ID generation world kept evolving. These are the alternatives that have emerged alongside the classic four approaches:

  • 🔶 ULID (Universally Unique Lexicographically Sortable Identifier)
    Like UUID v7 but encoded in a shorter, more readable base32 string. Example: 01ARZ3NDEKTSV4RRFFQ69G5FAV — looks like a password, sorts by time, works offline.
  • 🔷 KSUID (K-Sortable Globally Unique IDs)
    Created by Segment. 20-byte IDs: 4 bytes timestamp + 16 bytes random. Example: 0ujtsYcgvSTl8PAuAdqWYSMnLOv. Sorts naturally like timestamps.
  • 🔹 NanoID
    Tiny library generating URL-safe, compact IDs. Example: V1StGXR8_Z5jdHi6B-myT. Great for URLs where shorter = better.
  • ⭐ CUID2
    Security-first IDs resistant to fingerprinting. Designed for web apps and APIs. Replaces CUID which was deprecated in 2024.
Format Example ID Sortable Size Best For
UUID v7 018e–c8f0–7000–… ✅ 128 bits Databases, APIs (standard)
ULID 01ARZ3NDEKTSV4RRFFQ69G5FAV ✅ 128 bits Readable IDs, logs
KSUID 0ujtsYcgvSTl8PAuAdqW ✅ 160 bits Events, analytics streams
NanoID V1StGXR8_Z5jdHi6B ❌ ~126 bits URLs, short codes
Snowflake 1541815603606036480 ✅ 64 bits High-throughput distributed systems

⚠️ Common Mistakes Engineers Make with ID Generators

❌ Mistake 1 — Using UUID v4 as a database Primary Key without care
UUID v4 is completely random. When you insert rows with random primary keys, the database B-tree index fragments badly — like shuffling a sorted deck of cards with every new card. This causes page splits and slows down inserts drastically at scale. Fix: Use UUID v7 (time-sorted) or Snowflake IDs instead.
❌ Mistake 2 — Forgetting to handle clock skew in Snowflake
If your server's clock goes backward (even by 1 millisecond due to NTP sync), your Snowflake generator can produce duplicate IDs or throw an exception. Always add a clock-backward guard in your implementation! See the code example above. 🛡️
❌ Mistake 3 — Sharing a Machine ID between multiple processes
If two processes on the same machine both use machine_id=1 in Snowflake, they will collide when generating IDs at the same millisecond. Always assign a unique machine_id to every process, not just every physical machine.
❌ Mistake 4 — Using sequential integers as public-facing IDs
If your user ID goes 1, 2, 3… anyone can guess that user #100001 exists. Competitors can estimate your user growth by signing up twice and comparing IDs. This is called ID enumeration attack. Use UUIDs or Snowflake IDs for anything user-visible.
✅ Best Practice Checklist:
  • Use UUID v7 as your default choice for new projects — it is now an RFC standard
  • Use Snowflake when you need 64-bit IDs and millions per second at scale
  • Never share machine IDs across processes in a Snowflake setup
  • Always have a clock-skew guard in your Snowflake implementation
  • Use two Ticket Servers (odd/even) if you go that route — avoid SPOF
  • Store Snowflake IDs as BIGINT (not VARCHAR) in databases for efficiency
  • Do NOT use sequential IDs as publicly visible identifiers — security risk!

🧠 Quick Recap — The 4 Solutions at a Glance

  ┌──────────────────────────────────────────────────────────────────────┐
  │  SOLUTION              │ ANALOGY          │ BEST SCALE  │ COMPLEXITY │
  ├──────────────────────────────────────────────────────────────────────┤
  │  UUID v4               │ Random lottery   │ Any scale   │ Super easy │
  │  UUID v7               │ Dated lottery    │ Any scale   │ Super easy │
  │  Ticket Server         │ Bakery queue     │ Small/Med   │ Easy       │
  │  Multi-Master          │ Many cashiers    │ Medium      │ Medium     │
  │  Twitter Snowflake ❄️  │ Coded barcode    │ Massive     │ Advanced   │
  └──────────────────────────────────────────────────────────────────────┘

  The golden rule of ID design:
  Start simple → UUID v7 for most apps.
  Go complex → Snowflake only when you TRULY need billions per day.

❄️ Happy Generating! 🎟️ 🌐 🔢

Comments