You snap a photo, add a filter, write a caption, and hit Share. In seconds your photo is visible to thousands of followers around the world — perfectly displayed, beautifully compressed, and loading instantly on every device.
Meanwhile, 2 billion other people are scrolling their feeds, watching Reels, checking Stories, and sending DMs — all at the same time, all without a single hiccup.
Instagram is one of the most technically sophisticated systems ever built. It started as a tiny Python app on a single server in 2010. Today it handles 100 million photo uploads per day and serves billions of feed requests per second. The journey from one server to this scale is one of engineering's greatest stories. Let's explore every detail.
📸 100 million+ photos and videos uploaded every day
👥 2 billion+ monthly active users across 170+ countries
🎬 Reels gets over 200 billion plays per day
💬 100 million+ direct messages sent daily
🔍 500 million+ accounts use Instagram Stories every day
❤️ Over 4.2 billion likes happen on Instagram daily
🏷️ Over 25 million businesses use Instagram actively
Instagram was acquired by Facebook (Meta) in 2012 for $1 billion — at the time it had 13 employees and ran on a handful of servers.
🖼️ Section 1: Think of Instagram Like a Smart Global Photo Gallery
Before we touch a single server, let's build the right mental picture using a simple analogy.
Imagine a magical photo gallery that exists in every city in the world simultaneously. You take a photo, hand it to the gallery, and within seconds a copy of it appears in the personal gallery of every person who follows you — no matter where on Earth they are.
But this gallery is incredibly smart. It doesn't show you photos in random order. It studies what you like, what you watch longer, what you comment on — and shows you content it knows you'll love, even from people you don't follow yet. That's the Explore page. That's the Feed algorithm. That's Instagram.
| 🖼️ Photo Gallery Analogy | 📸 Instagram Reality | 🔧 Engineering Concept |
|---|---|---|
| Making multiple print sizes of one photo | Photo encoded in thumbnail, low, high res | Multi-resolution encoding |
| Gallery branches in every neighbourhood | Photo served from CDN closest to viewer | Global CDN (Facebook's PoPs) |
| Curator who knows your taste recommends art | Explore page + Feed ranking algorithm | ML Recommendation Engine |
| Delivering your photo to each follower's mailbox | Pushing post ID to follower feeds | Fan-out (write / read / hybrid) |
| Photos that disappear from walls after 24 hours | Instagram Stories | Ephemeral content + TTL storage |
🗺️ Section 2: The Big Picture — Instagram's Full Architecture
Instagram is a beautifully layered system. Let's see the bird's-eye view then zoom into every component.
🏗️ Instagram High-Level Architecture — Animated
Android
iPhone
Web (Browser)
API Clients
Post
Service
Feed
Service
User
Service
Search
Service
Reels
Service
Story
Service
DM
Service
Notif.
Service
+ Celery Workers (async tasks)
PostgreSQL
Users/Posts (sharded)
Redis
Feed Cache
Memcached
Object Cache
Cassandra
Activity / DMs
Haystack
Photo Storage
CDN
Media Delivery
We will explain every single component in detail below 👇
🆔 Section 3: Instagram's ID Generation — Their Own Snowflake
Every photo, video, comment, user, and story on Instagram needs a globally unique ID. Instagram built their own ID generator — similar in spirit to Twitter's Snowflake, but with a clever twist to work with their sharded PostgreSQL setup.
Simple auto-increment works on one database. But Instagram's PostgreSQL is sharded across many machines. If Database Shard 1 gives ID "5" and Database Shard 2 also gives ID "5" — you have two different posts with the same ID. Chaos! 💥
Instagram's solution: a 64-bit ID composed of:
⏱️ 41 bits — millisecond timestamp (like Snowflake)
🔢 13 bits — logical shard ID (which database generated this)
🔁 10 bits — sequence number (for IDs in the same millisecond)
The result: globally unique, time-sortable IDs generated inside PostgreSQL itself using a custom PL/pgSQL function — no separate ID service needed!
🆔 Inside a 64-bit Instagram Post ID
41 bits
ms since Jan 1 2011
(Instagram epoch)
13 bits
Which PostgreSQL
shard created this
10 bits
Up to 1,024 IDs
per millisecond
📤 Section 4: What Happens When You Post a Photo?
You spend 3 minutes picking the perfect filter, write a thoughtful caption, and hit Share. Let's trace every step that follows — from your phone to your followers' feeds.
📤 The Photo Post Journey — Every Single Step
Before uploading, Instagram's app compresses the image, applies your chosen filter, strips sensitive EXIF metadata (GPS location), and converts to JPEG. This reduces upload size by 60–80%. Smart!
The compressed image uploads via HTTPS to Instagram's upload servers, which store it in Facebook's Haystack photo storage system — a custom-built, distributed blob store optimised for billions of tiny images. The photo gets a unique storage key.
A Celery worker (asynchronous task queue) picks up a job to generate multiple versions of your photo: thumbnail (150×150), low-res feed (612×612), high-res (1080×1080), and story size. Each is stored separately in Haystack with its own key.
A new row is inserted: post ID (Instagram's custom 64-bit ID), author user ID, caption, location, hashtags, image storage keys, timestamp. The shard for this user is determined by
user_id % NUM_SHARDS.
A Kafka event carries
{post_id, user_id, hashtags, location, timestamp}.
Multiple downstream consumers react: Fan-out Service, Search Indexer,
Hashtag Counter, Notification Service, ML Feature Store. All in parallel.
The Fan-out Service reads your follower list and pushes your post ID to the Redis feed cache of each follower. When they scroll their feed, it's already there — instant! (Same hybrid celebrity/regular approach as Twitter.)
🔄 Image Processing Pipeline — One Upload, Many Versions
4MB · 4032×3024
HEIF (iPhone RAW)
Resize · Compress
Convert to JPEG/WebP
150 × 150
~8 KB
320 × 320
~25 KB
640 × 640
~80 KB
1080 × 1080
~200 KB
1080 × 1920
~300 KB
All 5 versions stored in Haystack. When your follower's phone requests the photo, Instagram serves the right size based on their screen resolution and network speed. This is called responsive image delivery.
🏔️ Section 5: Haystack — Instagram's Photo Storage System
Instagram stores hundreds of billions of photos. Regular file systems and databases can't handle this efficiently. After the Facebook acquisition, Instagram moved to Haystack — Facebook's custom-built photo storage system.
Facebook named it Haystack because finding one specific photo among hundreds of billions is like finding a needle in a haystack.
Traditional NAS (Network Attached Storage) systems open a file by name, which requires multiple metadata lookups (directory → file → data). For billions of photos with billions of accesses per day, this becomes unbearably slow.
Haystack's insight: store many photos in one large file and maintain an in-memory index of where each photo lives within that large file. One disk seek → photo found. Dramatically faster!
🏔️ How Haystack Stores Photos Efficiently
Photo request → Look up directory → Find file metadata → Seek to file on disk → Read photo. Each step is a separate disk operation. 3–5 disk seeks per photo! At billions of photos/day, disks become the bottleneck.
Photo request → Look up in-memory index → Find offset in large "haystack" file → One disk seek. The index fits entirely in RAM. All photos are packed into large logical volumes. One disk read. Done. 5x faster than traditional NAS.
Haystack Store: Physical machines holding the actual photo data in large volumes.
Haystack Directory: Maps photo ID → which Store machine holds it. Distributed, replicated.
Haystack Cache: An internal CDN — caches hot (frequently accessed) photos in memory.
🏠 Section 6: The Home Feed — How Instagram Builds What You See
This is Instagram's most visited and most technically complex feature. When you open Instagram and scroll your feed, what happens behind the scenes?
Before 2016, your feed showed posts in reverse-chronological order — newest first. Simple but imperfect. You'd miss posts from people you care about if you didn't check frequently.
Since 2016, Instagram uses a Machine Learning ranking algorithm that predicts which posts you're most likely to engage with — and shows those first. Your feed is now unique to YOU, assembled fresh on every scroll. 🤖
🌊 The Fan-Out Problem — Same as Twitter, Instagram's Solution
📝 Fan-Out on WRITE
When you post, push post ID into every follower's Redis feed cache immediately.
📖 Fan-Out on READ
When you open the app, dynamically fetch recent posts from all you follow.
🏆 Instagram's Hybrid
Regular users → write-time fan-out into Redis.
Celebrities (1M+ followers) → read-time lazy fetch.
Merge at read time.
🏠 Feed Assembly — What Happens When You Open Instagram
From Redis feed cache → get ~300 recent post IDs from people you follow (regular users).
From celebrity caches → fetch last 20 posts each from accounts you follow with 1M+ followers.
From ML Interest Graph → add ~50 posts Instagram thinks you'll love (even from strangers).
Each of the 350 candidates gets a score from Instagram's ranking model. The model predicts: "If I show this post to THIS user, what's the probability they will like it / comment / watch it / share it?" Signals used: your relationship with the author, your past engagement, post age, post engagement rate from others, content type (photo vs Reel), time of day.
Not more than 3 posts in a row from the same account. Content policy violations removed. Sensitive content filtered based on your preferences. Ads injected at defined intervals (positions 3, 7, 13...).
Fetch full post details (caption, like count, author info, photo URLs) from Memcached / PostgreSQL in one batch. Return top 12 posts to your app. The next 12 are pre-fetched as you scroll down.
⭕ Section 7: Instagram Stories — The 24-Hour Ephemeral Feed
500 million people use Instagram Stories every single day. Stories vanish after 24 hours. Engineering ephemeral content at this scale is fascinating.
↑ The coloured ring is purely CSS (gradient border). Each circle represents one person's active story (content within 24 hours). The ring disappears when their story expires!
⭕ Stories Engineering — How Ephemeral Content Works
Photo/video uploaded → Haystack storage → metadata saved in Cassandra (NOT PostgreSQL — Stories have high write volume and Cassandra handles it better). A TTL (Time-To-Live) of 86,400 seconds (24 hours) is set. Cassandra automatically deletes the row when TTL expires!
A Redis sorted set stores "users with active stories" per viewer, scored by story post time. To show the story tray at the top of your feed, Instagram queries Redis for all accounts you follow that have stories. One fast Redis call → complete story tray, ordered by recency.
Cassandra's native TTL handles expiry — no scheduled cleanup jobs required. When the TTL fires: Cassandra deletes the row, the media in Haystack is marked for deletion, and Redis is updated. The story vanishes from all views automatically. The magic of TTL-based architecture! ✨
Every time someone views your story, an event is recorded. You can see exactly who viewed your story. This list is stored in Cassandra, keyed by story ID with viewer user IDs. Sorted by view time. It's optimised for fast writes (views happen constantly) and medium-speed reads (you check your viewer list occasionally).
🎬 Section 8: Reels — The Video Engineering Powerhouse
Reels is Instagram's short-form video product — Instagram's answer to TikTok. With 200 billion Reels plays per day, the video encoding and delivery pipeline needs to be world-class.
🎬 Reels Video Processing Pipeline
Generates: 360p, 480p, 720p, 1080p — each in H.264 and H.265 (HEVC). Also generates: thumbnail, animated preview GIF, and waveform for audio scrubbing.
⏳ Animated: Encoding Progress for a 60-Second Reel
↑ All quality levels encode in parallel — a 60-second Reel is usually ready within 2–5 minutes of upload.
🔍 Section 9: The Explore Page — Discovering What You Don't Know You Love
The Explore page is Instagram's most fascinating product from an ML perspective. It shows you content from accounts you don't follow — yet it always seems to know exactly what you'll love.
Think of the Explore page like a brilliant music shop owner who watches what genres you buy, what you listen to in the store, which artists you ask about — and then recommends an album you've never heard of that you end up absolutely loving.
That's exactly what Instagram's Explore algorithm does. It doesn't just show you more of what you already like — it finds adjacent interests you haven't discovered yet, using signals from millions of similar users.
🔍 Explore Algorithm — 3-Stage Pipeline
From 100 billion posts → narrow to ~500 candidates.
How? Your interest graph: accounts similar to those you follow,
hashtags you use, posts you've liked/saved, accounts your followees interact with.
Also: trending content in your region.
Each candidate scored by a neural network predicting: probability you'll like, comment, share, or save. Signals: your past Explore interactions, post's engagement velocity, creator's relevance to you, content type preference.
Safety filter: remove anything violating Instagram's policies. Diversity filter: not too many posts from one creator. Novelty filter: don't show content you've already seen. Ad injection at defined positions.
🗄️ Section 10: Databases — Instagram's Data Layer
Instagram is famously a polyglot persistence system — they use the right database for each specific job. Let's understand each choice and why it was made.
The primary datastore for structured data. Sharded into thousands of partitions
using Instagram's custom Sharding Framework.
Sharding key for users: user_id % num_shards.
Sharding key for media: media_id % num_shards.
Connection pooling via pgBouncer — Django's ORM opens thousands of connections;
pgBouncer pools them to a manageable number at the database level.
Feed caches: each user has a Redis sorted set of post IDs (scored by timestamp). Session tokens, rate limiting counters, story tray data. Instagram also uses a Redis-based system called Flock for caching the social graph (who follows whom) — essential for fast fan-out. Instagram runs one of the world's largest Redis deployments.
Large objects (full post details, user profiles including profile photo URL, bio, follow counts) are cached in Memcached after first DB load. Instagram uses Facebook's heavily optimised Memcached deployment with custom connection pooling and consistent hashing. Cache hit rate for popular posts: 99%+.
High write-volume, time-series-style data. Stories metadata with TTL expiry. Direct Messages (write-heavy: millions of messages per second). Notification logs. Activity feeds (who liked/commented on what). Cassandra's tunable consistency and time-ordered write performance makes it ideal for these workloads.
Kafka carries all events: new posts, likes, follows, story views. Celery (Python task queue backed by Redis) handles async jobs: image resizing, push notifications, search indexing, ML feature updates. Instagram runs thousands of Celery workers processing millions of tasks per minute.
💻 Section 11: A Peek at the Code — Simplified Examples
Let's look at simplified pseudocode that brings Instagram's core engineering to life. These are teaching examples based on publicly documented patterns — not Instagram's actual code.
This shows the Post Service handling a photo upload. It does three critical things in sequence: (1) uploads the image to Haystack (Instagram's photo storage) and generates a unique ID, (2) saves the post's metadata to the correct PostgreSQL shard using the custom Instagram ID, (3) publishes a Kafka event so all downstream systems — fan-out, notifications, search, ML feature stores — can react asynchronously without slowing down the response to the user. Notice the user gets a response immediately, before fan-out even begins!
# Instagram Post Service — Handle Photo Upload (Pseudocode) # Django-based, runs on thousands of web servers def upload_post(user_id, image_file, caption, hashtags, location): # Step 1: Generate Instagram's custom 64-bit ID # This runs INSIDE PostgreSQL as a PL/pgSQL function # Format: [41-bit timestamp][13-bit shard_id][10-bit sequence] shard_id = user_id % NUM_SHARDS # which shard this post lives on post_id = postgres[shard_id].call("insta_id_generate", shard_id) # Step 2: Upload original image to Haystack storage storage_key = haystack.store( data = image_file.read(), content_type= "image/jpeg", post_id = post_id ) # returns e.g. "hs://photo/123456789/original" # Step 3: Save post metadata to sharded PostgreSQL postgres[shard_id].insert("media", { "id": post_id, "user_id": user_id, "caption": caption, "storage_key": storage_key, "hashtags": hashtags, "location": location, "created_at": now() }) # Step 4: Queue async Celery tasks (non-blocking — happen in background!) celery.send_task("resize_image", args=[post_id, storage_key]) celery.send_task("extract_features", args=[post_id]) # ML signals celery.send_task("index_hashtags", args=[post_id, hashtags]) # Step 5: Publish Kafka event → triggers fan-out, notifications etc kafka.publish(topic="new_post", message={ "post_id": post_id, "user_id": user_id, "hashtags": hashtags, "timestamp": now() }) # Return immediately — fan-out happens asynchronously return {"post_id": post_id, "status": "published"}
This shows the Feed Service building your home feed when you open Instagram. It follows the hybrid fan-out strategy: it reads your pre-built feed from Redis (posts from regular users you follow that were pushed in at write time), fetches recent posts from any celebrity accounts you follow separately (because their fan-out was too expensive at write time), merges both lists sorted by time, scores them with the ML ranking model, and finally hydrates the post IDs into full post objects from Memcached/PostgreSQL. The entire process runs in under 100ms for most users!
# Feed Service — Assemble Home Feed (Pseudocode) # Called every time you pull-to-refresh or open the Instagram app CELEBRITY_THRESHOLD = 1_000_000 # 1M+ followers = celebrity path FEED_CACHE_SIZE = 300 # max post IDs in Redis feed cache def get_home_feed(viewer_id, page=1, count=12): # Step 1: Get pre-built feed from Redis (post IDs from regular followees) # Redis ZREVRANGE returns IDs sorted by score (score = post timestamp) cached_ids = redis.zrevrange( key=f"feed:{viewer_id}", start=0, end=FEED_CACHE_SIZE ) # Step 2: Find celebrities this user follows (from Flock social graph cache) all_following = flock.get_following(viewer_id) celebrities = [uid for uid in all_following if user_service.follower_count(uid) >= CELEBRITY_THRESHOLD] # Step 3: For each celebrity, get their last 20 posts (read-time fan-out) celeb_ids = [] for celeb_id in celebrities: recent = redis.zrevrange(f"user_posts:{celeb_id}", 0, 19) celeb_ids.extend(recent) # Step 4: Merge + de-duplicate + sort newest first # Snowflake-style IDs are time-ordered — larger ID = newer post! all_ids = list(set(cached_ids + celeb_ids)) all_ids.sort(reverse=True) # newest first # Step 5: ML re-ranking — score each post for THIS viewer scored = ranking_model.score(viewer_id=viewer_id, post_ids=all_ids) ranked_ids = [pid for pid, _ in sorted(scored, key=lambda x: x[1], reverse=True)] # Step 6: Paginate start = (page - 1) * count page_ids = ranked_ids[start : start + count] # Step 7: Hydrate — check Memcached first, PostgreSQL for misses # Batch lookup — NOT one query per post! posts = memcached.get_multi([f"post:{pid}" for pid in page_ids]) missing = [pid for pid in page_ids if pid not in posts] if missing: db_posts = postgres.batch_fetch("media", ids=missing) posts.update(db_posts) memcached.set_multi({f"post:{p.id}": p for p in db_posts}) # Step 8: Apply viewer-specific filters (block/mute/sensitive content) return apply_filters([posts[pid] for pid in page_ids if pid in posts], viewer_id=viewer_id)
This shows how Instagram Stories are posted and auto-expired.
The most important concept here is the TTL (Time-To-Live).
When we save the story to Cassandra with TTL = 86400 (seconds in 24 hours),
Cassandra automatically deletes that row after 24 hours — no cleanup job needed!
We also add the story to a Redis sorted set for the viewer's story tray
(scored by timestamp so newer stories show first).
That Redis key also expires after 24 hours — the story tray updates itself automatically.
This is the beauty of TTL-based architecture for ephemeral content.
# Story Service — Post a Story (Pseudocode) # Stories live exactly 24 hours then vanish automatically STORY_TTL_SECONDS = 86_400 # 24 hours in seconds def post_story(user_id, media_file, media_type): # Step 1: Upload media to Haystack (same as regular posts) story_id = generate_story_id() # custom 64-bit ID storage_key = haystack.store(media_file) expires_at = now() + STORY_TTL_SECONDS # Step 2: Save story metadata to Cassandra WITH TTL! # Cassandra will automatically DELETE this row after 86,400 seconds. # No cron job. No cleanup worker. Pure database magic! cassandra.insert( table = "stories", data = { "story_id": story_id, "user_id": user_id, "storage_key": storage_key, "media_type": media_type, "created_at": now() }, ttl = STORY_TTL_SECONDS # ← THIS is the magic line! ) # Step 3: Tell all followers "this user now has an active story" # Add to each follower's story tray sorted set in Redis # Score = current timestamp (so newer stories sort to front) followers = flock.get_followers(user_id) for follower_id in followers: redis.zadd( key=f"story_tray:{follower_id}", score=now_unix_timestamp(), member=user_id ) # This Redis key also auto-expires after 24 hours redis.expire(f"story_tray:{follower_id}", STORY_TTL_SECONDS) return {"story_id": story_id, "expires_at": expires_at} # When user opens the app — get story tray in ONE Redis call! def get_story_tray(viewer_id): # Returns list of user_ids with active stories, sorted by recency return redis.zrevrange(f"story_tray:{viewer_id}", 0, 49)
Using Cassandra's TTL and Redis's EXPIRE for Stories means: no cleanup cron jobs, no manual deletion, no orphaned data. The database manages expiry internally as a first-class feature. At Instagram's scale, eliminating cleanup jobs prevents entire categories of bugs (what if the cleanup job fails? what if it runs late?). Designing expiry into the data model itself is always cleaner than cleanup after the fact. 🎯
📈 Section 12: Scalability — How Instagram Grew from 1 Server to Billions
Instagram launched in October 2010 on a single EC2 instance. Within 18 hours they had 25,000 users. Within a week, 100,000. The scaling story of Instagram is one of the most educational in technology.
📅 Instagram's Scaling Journey — Key Decisions
Instagram's PostgreSQL sharding framework partitions data across thousands of database instances. Each shard is replicated (primary + replicas). Reads go to replicas (scale reads). Writes go to the primary. pgBouncer manages connection pools — preventing thousands of Django web workers from overwhelming the database with direct connections.
Instagram started on Python/Django. Python is developer-friendly but not the fastest. As they scaled, Instagram rewrote the most performance-critical paths (feed ranking, image processing, API response serialisation) in C++ and Go. The business logic stayed in Python; the compute-heavy kernels moved to faster languages. This pragmatic approach balanced developer productivity with raw performance.
Instagram's Memcached cluster absorbs the overwhelming majority of read traffic. A popular post (say, a celebrity post with 5M likes) is loaded once from PostgreSQL and then served from Memcached for every subsequent read. At peak, Memcached serves millions of requests per second — almost none of which touch the database. PostgreSQL is protected from read storms.
🛡️ Section 13: Reliability — Why Instagram Stays Up
Every PostgreSQL shard has a primary and multiple read replicas. Facebook's infrastructure automatically promotes a replica to primary if the primary fails — within seconds. Zero manual intervention. Instagram's data has never been lost due to a single database failure.
Kafka stores all events durably on disk, replicated across brokers. If a Celery worker crashes mid-task (image resize fails), the task is retried. If a fan-out worker crashes, Kafka's consumer group re-assigns the partition to another worker which resumes from where it left off. No posts go un-fanned-out.
If the ML ranking model is slow or unavailable, Instagram falls back to reverse-chronological order — simple but functional. If the celebrity post fetch fails, only the pre-built Redis feed is shown. Partial feeds are always better than blank screens. Every failure mode has a fallback.
Instagram runs across Facebook/Meta's global data centers in the US, Europe, and Asia. Traffic is geo-routed to the nearest region. If an entire region goes offline, traffic fails over to other regions within minutes. The CDN layer (for media) has even more redundancy — photos are served from edge nodes that have no dependency on the origin data center being alive.
🗺️ Section 14: Everything Together — The Complete Instagram Map
📸 Instagram — Complete Architecture Map
Topics: new_post · new_like · new_follow · story_view · reel_play
Users · Posts · Follows (sharded)
Feed Cache · Social Graph · Stories
Object Cache (Posts / Users)
DMs · Stories Meta · Activity
Photo · Video Storage
Global Media Delivery
📐 Section 15: Instagram's Core Design Principles
Instagram launched on one server with a simple Django app. They moved fast. As they hit limits, they fixed them pragmatically — shard the database, add a cache, move hot paths to faster languages. The lesson: don't over-engineer before you need to, but know your scaling breaking points before you reach them.
Instagram's CDN, Memcached, and Redis layers all serve the same purpose: get data as close to the user as possible, as fast as possible. Cold storage (PostgreSQL, Cassandra, Haystack) is the truth. Caches are what you show. Designing around this layered distance-from-truth model keeps systems fast and consistent.
Image resizing, push notifications, search indexing, ML feature extraction — none of these happen in the request-response cycle. They're queued in Kafka and processed by Celery workers in the background. This makes every user-facing action feel instantaneous and decouples services so one slow worker doesn't affect the whole system.
PostgreSQL for structured relational data. Cassandra for high-write time-series. Redis for real-time caching and social graphs. Memcached for large object caching. Haystack for billions of binary blobs. No single database can do all these jobs well. Picking the right tool for each problem is one of the most important skills in distributed systems design.
🎓 Section 16: System Design Interview Cheat Sheet
If you're asked to "Design Instagram" in a system design interview, here is your exact 5-step framework:
- Photo + video upload and display? (yes — photos and Reels/videos)
- Follow system? (yes — asymmetric, like Twitter)
- Home feed? (yes — algorithmic ranking, not purely chronological)
- Stories? (yes — ephemeral, 24-hour expiry)
- Search + Explore? (yes — hashtag search and ML-driven discovery)
- Direct Messages? (yes — like WhatsApp, encrypted)
- Likes, Comments, Saves? (yes — engagement features)
- 2B users, ~500M DAU
- 100M photo/video uploads per day = ~1,160 uploads/second average
- Feed reads vastly outnumber writes (100:1 ratio)
- Each photo: ~200KB (compressed) × 100M/day = 20TB storage/day
- 5 resolutions per photo × 20TB = 100TB media ingested daily
- 4.2B likes/day = ~48,600 likes/second (high-write Cassandra workload)
- Write path: Client → Load Balancer → Post Service → Haystack + PostgreSQL + Kafka
- Async: Kafka → Celery Workers → Image resize, fan-out, notifications, ML
- Fan-out: For regular users → push post ID to Redis feed. For celebrities → read-time lazy fetch
- Read path: Feed Service → Redis cache + celebrity posts → ML rank → Memcached hydrate
- Stories: Cassandra with TTL + Redis sorted sets for tray
- Media delivery: Haystack → Meta CDN → User's device
- Explain Instagram's custom 64-bit ID generation (sharding-aware Snowflake variant)
- Explain hybrid fan-out strategy for regular users vs celebrities
- Explain Haystack photo storage — why it beats traditional NAS for billions of tiny files
- Explain multi-resolution image encoding pipeline with Celery workers
- Explain Stories TTL design using Cassandra's native TTL feature
- Explain Explore page 3-stage pipeline: candidate gen → ML rank → policy filter
- Celebrity posts (Cristiano Ronaldo, 600M followers) → read-time fan-out prevents write storm
- Viral post suddenly gets 10M likes/hour → Memcached handles read storm; Cassandra handles like write storm
- Story doesn't disappear after 24hrs → Cassandra TTL + Redis EXPIRE both fire; belt-and-suspenders approach
- User deletes a post → remove from PostgreSQL, mark Haystack objects for deletion, invalidate Memcached, remove from Redis feed caches
- Copyright video uploaded as Reel → content fingerprinting at upload (like YouTube's Content ID)
🎉 Final Summary
Instagram's greatest engineering insight is this: the read path must be blindingly fast, because reads outnumber writes 100 to 1.
Every architectural decision — pre-computing feeds at write time, Memcached for post objects, multi-resolution photos for network conditions, CDN for media — exists to make the read experience as close to instant as possible.
Optimise ruthlessly for what happens most frequently. Everything else is secondary. 📸
Happy Learning! Keep Building! 🔥
Comments
Post a Comment