Skip to main content

Netflix System Design: Video Streaming, CDN, Encoding and Scaling Explained

Calculating read time…

You open Netflix. You see a beautiful thumbnail of your favourite show. You click play. Within 2 seconds, crisp HD video is streaming on your screen — perfectly smooth, perfectly synced.

Simultaneously, 300 million other people around the world are doing exactly the same thing. Different shows. Different languages. Different internet speeds. All working perfectly.

How does Netflix pull this off without breaking a sweat? That's exactly what we're going to unpack today — completely, from scratch.


💡 Netflix by the Numbers — 2026

🎬 300 million+ paying subscribers across 190 countries
▶️ 500 million+ hours of video streamed every single day
📦 Netflix uses about 15% of all global internet bandwidth during peak hours
🎞️ Netflix's content library: 36,000+ hours of original and licensed content
☁️ Netflix runs almost entirely on Amazon Web Services (AWS)
🌐 Netflix built its own private CDN called Open Connect — embedded in ISP data centers
🐒 Netflix intentionally breaks its own servers every day to test resilience (Chaos Engineering!)

This is one of the most sophisticated engineering systems ever built. Let's explore it all!

📚 Section 1: Think of Netflix Like a Magical Library + Pizza Delivery System

Before touching a single server, let's build the right mental picture.

Imagine a magical library that has every movie and TV show ever made. You're a member. You can request any book (video) at any time. The moment you ask, a copy of that book is delivered to your doorstep in under 2 seconds — no matter where you live on Earth.

But here's the twist: the library doesn't send the entire book at once. It sends it page by page, just slightly ahead of where you're reading. If you're reading slowly (slow internet), it gives you a shorter, easier-to-carry version. If you're reading fast (fast internet), it gives you the beautiful full-colour hardcover.

📚 Library Analogy 🎬 Netflix Reality 🔧 Engineering Concept
Local branch library near your home CDN server in your city Open Connect CDN
Book available in large print OR normal Video in 4K / 1080p / 720p / 480p Adaptive Bitrate Streaming
Librarian recommends books you'll love The "Because you watched..." row Recommendation Engine
Multiple copies of popular books Popular show cached on many servers Content Replication & Caching
Counter staff check your membership card Login + subscription verification Authentication & Authorization

🗺️ Section 2: The Big Picture — Three Planes of Netflix

Here's the most important insight about Netflix's architecture: Netflix is actually three separate systems working together. Engineers call them the three "planes":

✈️ Control Plane — The Brain

Runs on AWS (Amazon Web Services). Handles everything that is NOT the actual video: user accounts, login, search, recommendations, the homepage you see, billing, device management. This is what runs in AWS data centers around the world.

📡 Data Plane — The Muscle (Open Connect CDN)

Netflix's own private CDN called Open Connect. Thousands of custom-built servers placed inside ISP (internet provider) data centers worldwide. The actual video bytes you watch come from here — NOT from AWS. This is Netflix's biggest competitive advantage.

📱 Client Plane — The Eyes & Ears

Your Netflix app (on TV, phone, laptop, tablet, game console). The client is remarkably smart — it monitors your internet speed every second, picks the best quality, prefetches the next episode, manages DRM, and reports playback quality back to Netflix for monitoring.

🏗️ Netflix Complete Architecture — Animated Overview

📺
Smart TV
💻
Laptop
📱
Mobile
🎮
Console
📟
Fire Stick
⬇️
API calls (login, browse, search)
⬇️
Video bytes (actual streaming)
☁️ AWS Cloud (Control Plane)
🚪 Zuul API Gateway
👤 User Service (Auth)
🤖 Recommendation Engine
🔍 Search Service (ElasticSearch)
🎞️ Video Processing Pipeline
📊 Analytics (Apache Spark)
📨 Kafka Event Streaming
🗄️ Cassandra / MySQL / EVCache
🌐 Open Connect CDN (Data Plane)
📦 OCA Server — India (Mumbai)
📦 OCA Server — USA (NY, LA, Chicago)
📦 OCA Server — Europe (London, Paris)
📦 OCA Server — Asia (Singapore, Japan)
📦 1000s more cities worldwide...
✅ Pre-loaded with popular content nightly!

↑ Your API calls go to AWS. Your video bytes come from the nearest Open Connect server. Two completely separate paths!


🌐 Section 3: Open Connect — Netflix's Secret Weapon

Most companies use third-party CDNs like Akamai or Cloudflare. Netflix said "No thanks" and built their own private CDN. This is called Open Connect and it is Netflix's single biggest engineering advantage.

💡 The Brilliant ISP Deal

Netflix went to every major internet provider (Airtel, Jio, AT&T, Comcast, BT...) and said: "Let us put our own servers inside your data center — for free. We'll fill them with popular Netflix content. Your customers get faster Netflix. You save bandwidth costs. Win-win!"

Today, Open Connect Appliances (OCAs) — Netflix's custom servers — sit inside thousands of ISP data centers worldwide. When you watch Netflix, the video comes from a server that might be just a few kilometres from your home!

🌐 How Open Connect Delivers Video to You

❌ Without Open Connect

📱 You in Jaipur, India
⬇️ Data travels 15,000km
🖥️ AWS Server in Oregon, USA

😩 300ms+ latency · Expensive · Chokes ISP backbone

✅ With Open Connect

📱 You in Jaipur, India
⬇️ Data travels 5–50km!
📦 OCA Server inside Jio/Airtel data center (Mumbai)

🚀 5–15ms latency · Fast start · Smooth 4K!

🌙 The Nightly Fill — How OCAs Stay Up to Date

During off-peak hours (typically late night), Netflix's central AWS system pushes fresh content to all the OCA servers worldwide.

This is called proactive caching. Netflix predicts which shows will be popular tomorrow in each region, and makes sure copies are already on the local OCA servers before anyone requests them.

🌙 Nightly OCA Fill Cycle

🌙 11 PM
Off-peak starts
➡️
📊 AWS predicts
tomorrow's top shows
➡️
📡 Pushes video files
to regional OCAs
➡️
☀️ 6 AM
OCAs are ready!

By the time users wake up and start watching, the popular content is already sitting in their local OCA. Zero wait time. Zero long-distance travel.


🎞️ Section 4: Netflix's Video Encoding — Better Than Everyone Else

When a studio hands Netflix a finished movie or show, Netflix doesn't just put it online as-is. They process it through one of the most sophisticated video encoding pipelines on Earth.

💡 What Is Video Encoding?

A raw, uncompressed 2-hour movie at 4K can be 500+ GB in size. You can't stream that over the internet! Encoding compresses it while trying to keep the picture quality as high as possible.

It's like packing your suitcase. A good packer (encoder) fits everything neatly — perfect quality, minimum space. A bad packer either wastes space or crumples your clothes!

🎞️ Netflix's Video Processing Pipeline — From Raw Film to Your Screen

1
Studio uploads raw master file
Netflix receives the original mezzanine (uncompressed) file from the studio. For a movie, this can be hundreds of gigabytes — pristine, highest quality.
⬇️
2
Media Quality Validation
Automated systems check for artefacts, audio sync issues, black frames, colour problems, and missing subtitles. Any issues are flagged for human review. Only perfect files move to the next stage.
⬇️
3
🔥 Per-Title & Per-Scene Encoding (Netflix's Secret Sauce!)
Regular streaming: one set of quality levels for all shows. Netflix: unique quality settings for every title, even every scene. An action-packed car chase needs more data per second than a simple dialogue scene. Netflix analyses every scene and allocates exactly the right amount of data. Result: same file size, dramatically better picture quality.
⬇️
4
Encode into hundreds of versions simultaneously
On thousands of AWS cloud machines running in parallel, the video is encoded into multiple resolutions, multiple codecs (H.264, H.265/HEVC, AV1), multiple audio tracks, multiple subtitle languages — all at the same time. A 2-hour movie might generate 1,200+ different encoded files!
⬇️
5
Quality Verification using VMAF Score
Netflix created a quality measurement tool called VMAF (Video Multi-Method Assessment Fusion). It predicts how good a video looks to human eyes, not just mathematically. Every encoded version is scored. Only files above a quality threshold are accepted. VMAF is now used industry-wide!
⬇️
6
All versions uploaded to Amazon S3 + Open Connect OCAs
The finished encoded files are stored in Amazon S3 (master copies) and pushed to Open Connect servers worldwide. The title is now live!

🆕 AV1 — Netflix's Latest Codec Breakthrough

A codec is the compression method used to shrink a video. Think of it as the packing technique for your suitcase.

🎥 Codec 📅 Year 📦 File Size ⚡ Quality 📱 Support
H.264 (AVC) 2003 Baseline (1x) Good Every device
H.265 (HEVC) 2013 50% smaller Better Most modern devices
AV1 🏆 2018 70% smaller! Best Growing fast (2024+)
✅ Why AV1 Changes Everything:

With AV1, Netflix can deliver the same video quality at 70% less data. That means a show that needed 15 Mbps now only needs 4.5 Mbps! Users on mobile data see their data bill drop dramatically. Netflix saves enormous bandwidth costs. And the quality is actually better. It's a pure win on every dimension. 🎉

📶 Section 5: Adaptive Bitrate Streaming — Why Netflix Never Buffers

This is the technology behind the smoothest streaming experience on Earth. The idea is deceptively simple: always play the best quality your internet can handle, and switch automatically when your speed changes.

💡 The Water Tank Analogy

Imagine your Netflix app is a water tank. Video data is water flowing in from the internet. Video playback is water draining out the bottom at a constant rate.

If water flows in faster than it drains → tank fills up → use higher quality!
If water flows in slower than it drains → tank empties → drop quality to avoid buffer!

Netflix's ABR algorithm monitors this "tank level" every few seconds and adjusts quality silently in the background. You never notice. You never buffer!

📶 Adaptive Bitrate — Quality Switches in Real Time

360p
Slow start
480p
Getting better
720p
Good WiFi
1080p HD
Stable fast internet
4K HDR
Strong fibre
480p
Train tunnel!
1080p HD
Signal back!
← Time → (All these quality switches happen silently. You never see a loading spinner!)

🔪 How Video is Chunked (MPEG-DASH / HLS)

Netflix doesn't send your video as one giant file. It breaks it into tiny 2–4 second chunks. Each chunk is available in every quality level. At any moment, the app can switch quality at the next chunk — completely seamlessly.

📦 A 2-Hour Movie Broken Into Chunks

Chunk 1
4s · 4K
Chunk 2
4s · 4K
Chunk 3
4s · 1080p
Chunk 4
4s · 1080p
Chunk 5
4s · 720p
Chunk 6
4s · 1080p
Chunk 7
4s · 4K
... × 1,800 chunks for a 2-hour movie

Each chunk is independent. Quality can change at every chunk boundary. The player always downloads 2–3 chunks ahead — your buffer — so you never see a spinner.


▶️ Section 6: What Happens in the First 2 Seconds After You Press Play?

Let's trace the exact journey — every step, every millisecond.

▶️ The Complete Play Flow (0 → 2 seconds)

1
0ms — You click Play on "Stranger Things S4 E1"
Your Netflix app sends an API request to the Zuul API Gateway on AWS: "I want to play title ID: 12345678. Here is my auth token. My device is a 4K TV."
⬇️
2
10ms — Zuul validates your session + entitlement
The API Gateway checks: are you logged in? Does your subscription allow 4K? Is this show available in your country (licensing)? These checks hit the EVCache (Netflix's Redis-based cache) for speed.
⬇️
3
30ms — Steering Service finds the BEST OCA server for you
A smart routing service called the Playback Steering Service checks: which OCA servers near you have this specific video? Which one is least loaded right now? Which one is geographically closest? It picks the top 3 candidates and returns them to your client.
⬇️
4
50ms — Your app contacts the chosen OCA directly
Your app connects to the nearest OCA server and starts downloading the first few video chunks. Since the OCA is nearby (same city or ISP), data arrives fast.
⬇️
5
⏱️ Under 2 Seconds — 🎉 Video Starts Playing!
The first chunk is rendered on screen. Meanwhile the app silently downloads the next 20–30 seconds of video in the background (the "buffer"). The ABR algorithm immediately assesses your internet speed and picks the right quality for all subsequent chunks.

🔧 Section 7: Netflix's Microservices — 700+ Services on AWS

Netflix runs over 700 microservices on AWS — each one responsible for exactly one thing. They were among the first companies to go "all in" on microservices architecture, and they've open-sourced many of the tools they built to manage it.

🚪
Zuul — API Gateway

The front door to Netflix's entire backend. Every API request from every device hits Zuul first. It handles authentication, rate limiting, routing, and security. Handles millions of requests per second.

📡
Eureka — Service Discovery

With 700+ services, how does one service find another? Eureka is a registry — like a phone book for services. Every service registers itself: "Hi, I'm the Recommendation Service, I'm at this address." Others look it up before making requests.

⚡
Hystrix — Circuit Breaker

If one service (e.g., Recommendations) is slow or broken, Hystrix "trips" — like a fuse — and returns a fallback (e.g., generic popular shows) instead of letting the failure cascade and crash the whole site.

⚖️
Ribbon — Client Load Balancer

Instead of a central load balancer, Netflix uses client-side load balancing. Each service knows about all instances of services it depends on, and spreads its own requests across them. No single bottleneck!

📊
Atlas — Metrics & Monitoring

Netflix generates billions of metrics per day — CPU, memory, response times, error rates. Atlas ingests all of it and lets engineers query and visualize system health in real time. If something is wrong, engineers know within seconds.

💡 How Hystrix Circuit Breaker Works (The Fuse Box Analogy)

In your home, if a short circuit happens in one room, the fuse for that room blows — but the rest of the house still has power. You don't lose ALL electricity!

Hystrix does the same for software: if the Recommendation Service is having trouble, Hystrix "blows its fuse" and Netflix shows you a generic list of popular shows instead. The rest of Netflix (login, search, playback) continues working perfectly. A partially degraded Netflix is infinitely better than a completely dead Netflix! 🔌

🐒 Section 8: Chaos Engineering — Netflix Breaks Things on Purpose!

This is the most unusual and brilliant thing Netflix does. They deliberately destroy parts of their own system — while it's serving real users — to make sure it can handle failures gracefully.

💡 The Fire Drill Analogy

Schools do fire drills regularly. Why? Because if you wait for a real fire to test your evacuation plan, it might be too late.

Netflix does the same thing with their servers. They run fake "disasters" regularly so they can be sure that when a REAL disaster happens (and it always eventually does), their system handles it perfectly.

🐒 The Simian Army — Netflix's Chaos Engineering Toolbox

🐒
Chaos Monkey

The original. Randomly kills (terminates) individual server instances during BUSINESS HOURS to make sure the system can survive server failures without any user impact. If it can survive Chaos Monkey, it's robust!

🦍
Chaos Gorilla

Takes down an entire Availability Zone (one of AWS's isolated data centers). Simulates a full data center failure. Does Netflix survive? It must.

🌍
Chaos Kong

The most extreme test. Takes down an entire AWS Region (e.g. all of US-East-1). Netflix must automatically shift all traffic to other regions without any downtime. Used only in planned drills.

🔒
Security Monkey

Monitors all AWS configurations for security vulnerabilities. Checks that S3 buckets aren't publicly exposed, security groups are correct, and IAM permissions are properly restricted.

💸
Janitor Monkey

Finds and cleans up unused cloud resources — old servers nobody turned off, orphaned storage, abandoned test environments. Saves Netflix millions of dollars monthly in AWS costs!

🐒 Chaos Monkey in Action — What Actually Happens

🖥️ Server A
Healthy ✅
🖥️ Server B
KILLED 💥
by Chaos Monkey!
🖥️ Server C
Healthy ✅
🖥️ Server D
Healthy ✅
→ Load Balancer instantly routes Server B's traffic to A, C, D
→ Netflix users notice NOTHING. Streaming continues perfectly.

This test happens EVERY DAY in Netflix's production environment. If the system can't handle it — engineers find out and fix the weakness before a real outage occurs!


🤖 Section 9: The Recommendation Engine — How Netflix Reads Your Mind

Netflix claims that 80% of content watched on Netflix comes from recommendations — not from search. That little "Because you watched..." row is worth billions of dollars to them.

💡 The Netflix Prize — $1 Million for Better Recommendations

In 2006, Netflix ran a public competition: "Improve our recommendation algorithm by 10% and we'll pay you $1,000,000." It took 3 years and thousands of the world's best data scientists to crack it! A team called BellKor's Pragmatic Chaos won in 2009. This competition pushed the entire field of machine learning forward and gave us the collaborative filtering techniques used everywhere today.
👥 Collaborative Filtering

"People who watched what you watched also liked THIS."
If 10,000 people who loved Breaking Bad also binge-watched Ozark, Netflix recommends Ozark to you. You become part of a "taste cluster" with millions of others who share your viewing preferences.

🎬 Content-Based Filtering

"This show has properties similar to shows you've loved."
Netflix tags every piece of content with hundreds of attributes: dark tone, crime drama, anti-hero protagonist, non-linear storytelling, plot twists. If you love shows with these tags → find more shows with these tags.

⏱️ Contextual Signals

"What are you in the mood for RIGHT NOW?"
Netflix tracks time of day, day of week, device type, location, weather (via data partnerships). Friday night on a big TV → suggest long movies or episodes. Tuesday morning on a phone → suggest short episodes or documentaries.

🖼️ Thumbnail Personalisation

The thumbnail you see is not the same as what your friend sees!
Netflix has multiple thumbnails for each show (different characters, scenes, moods). If you watch a lot of movies with a specific actor, Netflix shows you a thumbnail featuring that actor for other shows they appear in — even if it's a small role. This dramatically increases click-through rates.

📡 Signals Netflix Uses to Understand You:
⏱️ Watch duration & completion rate ⭐ Explicit ratings (thumbs up/down) 🔁 Rewatches & re-episodes 🔍 Search history 📅 Day + time of viewing 📺 Device type ⏸️ Pause / skip behaviour 🌍 Region and language 👥 Household profiles

🧪 Section 10: A/B Testing — Netflix Tests EVERYTHING

Netflix makes decisions based on data, not opinions. They run hundreds of A/B tests simultaneously — testing different thumbnails, different layouts, different recommendation algorithms, different loading times — all on live users.

🧪 How Netflix A/B Tests Thumbnails

Group A — 50% of users
🎭 Thumbnail showing
the villain's face
Click rate: 3.2%
Group B — 50% of users
💥 Thumbnail showing
an action scene
Click rate: 5.8% ← WINNER!
After 1 week:
Group B is statistically significantly better. Netflix rolls out the action-scene thumbnail to 100% of users.
📈 +81% increase in clicks
for this title!
✅ Netflix Tests More Than Just Thumbnails:

🔹 Homepage layout — how many rows? which row first?
🔹 Preview videos — should the show auto-preview on hover?
🔹 Loading spinners — does showing a spinner cause users to cancel?
🔹 Skip Intro button — tested before rolling out to everyone
🔹 Audio quality — do users notice the difference between codecs?
🔹 Cancellation flow — which screen designs reduce cancellations?

🗄️ Section 11: Databases — Netflix's Data Layer

🗄️
Apache Cassandra — User Activity Data

Stores viewing history, play positions ("resume from 43:21"), interaction logs. Cassandra handles billions of writes per day effortlessly with horizontal scaling. Netflix shards by user ID — all data for one user lives on the same Cassandra node for fast reads.

⚡
EVCache — Netflix's In-Memory Cache

Built on Memcached, EVCache stores session data, user preferences, homepage data. It's replicated across multiple AWS regions — if one cache goes down, the others serve the data. Netflix open-sourced EVCache and it's used by many big companies.

🐬
MySQL — Billing & Subscriptions

Financial data (billing, subscriptions, payment methods) requires ACID transactions — absolutely no data loss or inconsistency. MySQL on AWS RDS handles this with automated backups, read replicas, and multi-AZ deployment.

🔍
ElasticSearch — Content Search

Powers Netflix's search bar. When you type "sci fi space opera with robots", ElasticSearch finds matches across 36,000+ titles in milliseconds using full-text search and semantic understanding. Netflix also uses Vespa (Yahoo's search engine) for recommendations.

📨
Apache Kafka — Event Streaming

Every user action (play, pause, rate, search, scroll) generates an event. Kafka handles trillions of these events daily. The Recommendation Engine, Analytics, A/B Testing — all consume from Kafka to update their models in near-real-time.


💻 Section 12: A Peek at the Code — Simplified Examples

Let's look at simplified code that illustrates key Netflix concepts. These are teaching examples — not Netflix's actual code!

📌 What This Code Does (Read Before The Code!)

This shows how the Zuul API Gateway handles an incoming "play video" request. It acts like a security guard + receptionist combo at the entrance of a building. Before your request reaches any backend service, Zuul: (1) verifies your identity, (2) checks you're allowed to watch this content, and (3) forwards you to the right service. Without this gateway, every microservice would need to handle security individually — messy!

# Simplified Zuul API Gateway — Play Request Handler (Pseudocode)

def handle_play_request(request):

    # Step 1: Authenticate — who is this user?
    # Token is checked against EVCache (fast!) before hitting DB
    user = auth_service.validate_token(request.auth_token)
    if not user:
        return response(status=401, message="Not logged in")

    # Step 2: Check subscription (does their plan include this content?)
    # e.g. 4K is only on Premium plan
    entitlement = subscription_service.check(
        user_id     = user.id,
        content_id  = request.content_id,
        quality     = request.requested_quality  # e.g. "4K"
    )
    if not entitlement.allowed:
        return response(status=403, message="Upgrade plan for 4K")

    # Step 3: Rate limiting — prevent abuse
    # Max 5 simultaneous streams per account
    active_streams = stream_limiter.count(user.account_id)
    if active_streams >= 5:
        return response(status=429, message="Too many streams")

    # Step 4: Route to the correct backend service
    # (Zuul doesn't do the work itself — it delegates!)
    playback_url = playback_service.get_manifest(
        content_id = request.content_id,
        user_id    = user.id,
        device     = request.device_type,
        quality    = entitlement.max_quality
    )

    # Return a manifest URL that contains the list of all video chunks
    return response(status=200, data={
        "manifest_url": playback_url,
        "oca_servers":   get_best_oca_servers(user.location)
    })
📌 What This Code Does (Read Before The Code!)

This is a dramatically simplified version of what Chaos Monkey does. It randomly selects a running server from your cloud infrastructure and shuts it down — on purpose! The idea is: if you do this regularly and your service survives, you've proven it's resilient. If it crashes, you've found a weakness before a real outage does it for you.

# Simplified Chaos Monkey — Intentional Server Terminator (Pseudocode)
# Real Chaos Monkey is open-source: github.com/Netflix/chaosmonkey

def run_chaos_monkey():

    # Only run during business hours — engineers must be awake!
    # Don't destroy servers at 3am when no one can respond
    if not is_business_hours():
        return

    # Get a list of all running server instances on AWS
    all_instances = aws.ec2.list_running_instances()

    # Randomly pick ONE instance to terminate
    victim = random.choice(all_instances)

    # Safety check: don't kill if there's only ONE instance left!
    # We always need at least 2 for high availability
    instances_in_group = aws.ec2.count_in_group(victim.auto_scaling_group)
    if instances_in_group <= 1:
        log("Skipping — only 1 instance in group. Fix your architecture!")
        return

    # BOOM! Kill the server.
    log(f"🐒 Chaos Monkey terminating: {victim.instance_id} ({victim.name})")
    aws.ec2.terminate(victim.instance_id)

    # Monitor what happens next
    # Does the load balancer re-route traffic automatically? ✅
    # Does the auto-scaler spin up a replacement? ✅
    # Do any alerts fire showing user impact? If yes → BUG FOUND! 🐛
    monitor_impact(duration=minutes(10), alert_on_user_impact=True)
📌 What This Code Does (Read Before The Code!)

This shows the simplified logic of Adaptive Bitrate Streaming running inside your Netflix app. Every few seconds, the app measures your internet speed, checks how full your video buffer is, and picks the best quality for the next chunk. This decision happens silently, dozens of times per minute, completely automatically. This is why Netflix almost never shows a loading spinner!

# Adaptive Bitrate (ABR) Algorithm — Inside the Netflix Player (Pseudocode)
# This runs on YOUR device, not on Netflix's servers!

# Available quality levels for this show (pre-defined by Netflix)
QUALITY_LEVELS = [
    {"name": "4K HDR",  "bitrate_mbps": 16.0},
    {"name": "1080p HD", "bitrate_mbps": 5.0},
    {"name": "720p",    "bitrate_mbps": 3.0},
    {"name": "480p",    "bitrate_mbps": 1.5},
    {"name": "360p",    "bitrate_mbps": 0.5}
]

def pick_quality_for_next_chunk():

    # Measure how fast data is arriving from the OCA server
    current_speed_mbps = measure_download_speed()

    # Check how many seconds of video are buffered ahead of playhead
    buffer_seconds = player.get_buffer_ahead()

    # Safety margin: use only 85% of measured speed
    # (Internet speed fluctuates — be conservative!)
    safe_speed = current_speed_mbps * 0.85

    # If buffer is critically low, drop to lowest quality immediately!
    if buffer_seconds < 5:
        return QUALITY_LEVELS[-1]  # 360p — emergency mode

    # Find the HIGHEST quality that fits within our safe speed
    for quality in QUALITY_LEVELS:
        if quality["bitrate_mbps"] <= safe_speed:
            log(f"📶 Selecting {quality['name']} | Speed: {current_speed_mbps:.1f} Mbps | Buffer: {buffer_seconds}s")
            return quality

    # Fallback to lowest quality
    return QUALITY_LEVELS[-1]
✅ Why the 0.85 Safety Margin Matters:

Internet speed is never perfectly stable — it fluctuates constantly. If you pick a quality that needs exactly 100% of your current speed, any tiny drop will cause buffering. By using only 85% of measured speed, Netflix creates a cushion. The video keeps playing smoothly even through small speed dips. 🎯

📈 Section 13: Scalability — How Netflix Handles 500M Hours Daily

🌩️ 1. Cloud-Native on AWS

Netflix runs on AWS across 3 AWS regions simultaneously. During peak demand (Friday night), AWS auto-scales — Netflix can go from 10,000 to 100,000 servers in minutes, automatically. They pay for what they use. No idle capacity during off-peak hours.

🌐 2. Open Connect Offloads 95% of Traffic

By having OCAs inside ISP data centers, approximately 95% of Netflix traffic never even touches AWS. The video goes directly from the ISP's data center to your home. This means Netflix's AWS infrastructure only handles API calls — a tiny fraction of total traffic.

⚡ 3. Client-Side Intelligence

Netflix's app is extremely smart. It measures bandwidth, picks quality, manages the buffer, decides which OCA to use, reports telemetry. By moving this intelligence to the client, Netflix's servers do far less work per stream. 300 million smart clients = massively distributed computation.

📊 4. Stateless Microservices

Every Netflix microservice is stateless — it doesn't remember anything between requests. All state lives in databases (Cassandra, EVCache). This means you can scale any service by simply adding more server instances. No complex session stickiness, no shared memory between instances.


🗺️ Section 14: Everything Together — The Complete Netflix Map

🎬 Netflix Complete Architecture Map

── YOUR DEVICES ──
📺 TV
💻 Laptop
📱 Mobile
🎮 Console
⬇️
API Calls
⬇️
Video Bytes
☁️ AWS Cloud
🚪 Zuul API Gateway
👤 Auth + Subscription Service
🤖 Recommendation Engine
🔍 Search (ElasticSearch)
📡 Playback Steering Service
🎞️ Video Encoding Pipeline
📊 Analytics (Spark + Kafka)
🗄️ Cassandra · MySQL · EVCache · S3
🐒 Chaos Monkey (always running!)
🌐 Open Connect CDN
📦 OCAs in 1000s of cities
📦 OCA — India (Jio, Airtel)
📦 OCA — USA (Comcast, AT&T)
📦 OCA — Europe (BT, Deutsche T.)
📦 OCA — Asia (NTT, SingTel)
+ 1000s more ISP partnerships...
✅ Serves 95% of all traffic!
Video arrives in <15ms

📐 Section 15: Netflix's Core Engineering Principles

🔥 Design for Failure (Resilience over Perfection)

Netflix assumes every component WILL fail. Instead of preventing all failures (impossible), they design the system to survive and self-heal when failures occur. Chaos Monkey is the embodiment of this principle.

📊 Data over Opinion

Every decision — thumbnail design, recommendation algorithm, UI layout — is validated with A/B tests on real users. If data says a change hurts engagement, it gets reverted no matter how "good" the idea seemed. Data is the boss.

🌐 Local Delivery at Global Scale

Open Connect is proof that sometimes the best way to scale globally is to go hyperlocal. By placing servers inside ISPs, Netflix removed the biggest bottleneck (long-distance internet) entirely.

♻️ Graceful Degradation

When a service fails, Netflix degrades gracefully — shows less-perfect results rather than showing an error. Recommendations fail? Show popular content. Personalized artwork fails? Show a generic thumbnail. A partially working Netflix is infinitely better than a broken Netflix.

🔓 Freedom and Responsibility

Netflix gives its engineering teams almost total freedom to choose their own tech stacks, databases, and architectures. But with that freedom comes responsibility: your service must be reliable. Own-it, run-it, fix-it. This culture of ownership drives engineering quality and accountability.


🎓 Section 16: Cheat Sheet

If you're asked to "Design Netflix" in a system design interview, here's your structured 5-step game plan:

Step 1: Clarify Requirements (5 mins)
  • VOD only? Or live streaming too? (say: VOD focus, live secondary)
  • How many users? (say: 300 million, 100M+ concurrent peak)
  • Video quality support? (say: 480p to 4K HDR, Dolby Atmos audio)
  • Offline downloads? (say: yes, DRM-protected)
  • Global? (say: 190 countries, multiple languages)
Step 2: Estimate Scale (5 mins)
  • 500M+ hours of video streamed daily
  • Peak: ~100M concurrent streams (Friday 9pm)
  • 1080p stream = ~5 Mbps → 100M streams = 500 Tbps total bandwidth!
  • New content: ~5,000 hours/week encoded into 1,200 versions each
  • 15% of all global internet traffic during peak hours
Step 3: High-Level Design (10 mins)
  • Control Plane (AWS): Zuul → Auth → Playback Steering → OCA list
  • Data Plane: Client contacts nearest OCA for actual video bytes
  • Encoding Pipeline: Raw file → validate → per-title encode → VMAF check → S3 → OCAs
  • Nightly fill: AWS pushes popular content to OCAs proactively
Step 4: Deep Dive (15 mins)
  • Explain Open Connect CDN architecture (OCA inside ISP) in detail
  • Explain Adaptive Bitrate Streaming with chunking and buffer management
  • Explain per-title and per-scene encoding + VMAF quality metric
  • Explain Chaos Engineering philosophy (Chaos Monkey, Gorilla, Kong)
  • Explain Hystrix circuit breaker pattern for graceful degradation
  • Explain recommendation pipeline (collaborative + content-based filtering)
Step 5: Edge Cases (5 mins)
  • New show launches (Squid Game 2) → huge spike → pre-warm all OCAs globally
  • ISP outage → Playback Steering automatically redirects to next-best OCA
  • DRM for offline downloads → content decryptable only on licensed device
  • 4K TV → HDR metadata embedded in manifest, client picks compatible stream
  • Bad actor sharing accounts → device fingerprinting + concurrent stream limits

🎉 Final Summary 

🎬 Three Planes — Control Plane (AWS), Data Plane (Open Connect), Client Plane (your app)
🌐 Open Connect CDN — Netflix's own CDN inside ISP data centers, serving 95% of traffic from your city
🎞️ Per-Title / Per-Scene Encoding — Every show gets custom encoding settings. Action scenes get more data. Quiet scenes get less
📶 Adaptive Bitrate Streaming — 4-second chunks, quality switches silently based on your internet speed every few seconds
🚪 Zuul API Gateway — Every API call enters through one door, gets authenticated, rate-limited, and routed
⚡ Hystrix Circuit Breaker — When one service breaks, the fuse trips and graceful fallback kicks in — rest of Netflix stays alive
🐒 Chaos Engineering — Netflix deliberately kills servers daily to prove resilience BEFORE real outages happen
🤖 Recommendation Engine — 80% of views come from recommendations; personalised thumbnails; collaborative + content-based filtering
🧪 A/B Testing Culture — Every feature, thumbnail, and UI change is tested on real users before full rollout
🆕 AV1 Codec — 70% smaller file size, same quality — saves bandwidth for both Netflix and its users
✅ The Biggest Lesson From Netflix's Architecture:

Netflix's greatest insight was realising that the hardest problem wasn't storing videos — it was delivering them. Their entire system is engineered around one goal: get the first byte of video to your screen as fast as possible, then keep it flowing smoothly no matter what.

Open Connect, ABR, per-title encoding, the client-side ABR algorithm — every single decision is in service of that one goal. 🎯


Happy Learning! Keep Building! 🔥

Comments