Skip to main content

Master and Worker Nodes in Docker Orchestration: Scaling & Fault Tolerance for ML Workloads

Calculating read time…

Imagine you run the world's most popular pizza restaurant. On a normal Tuesday, 2 chefs can handle all the orders easily.

But then a football match ends. Suddenly 500 people want pizza at the same time. Your 2 chefs can't keep up. Orders pile up. Pizzas get cold. Customers leave angry.

Now imagine you have a Restaurant Manager who watches the order queue in real time. The moment it gets busy, they call in extra chefs from a waiting list. If one chef gets sick, the manager immediately reassigns their orders to someone else. Customers never notice a thing.




💡 That Restaurant Manager = the Master Node.
The Chefs = Worker Nodes.
The Pizzas = your ML model containers (fraud detection, image recognition, etc.)

When your ML model API receives 10,000 requests per second at 3 PM, the Master Node decides how many containers to spin up, which Worker Nodes handle them, and what happens when a Worker Node crashes.

📚 What We'll Cover

  • 🔹 Why a Single Container Is Never Enough in Production
  • 🔹 What Is Container Orchestration? (The Big Idea)
  • 🔹 The Master Node — The Brain of the Cluster
  • 🔹 The Worker Node — The Muscle of the Cluster
  • 🔹 What Lives Inside a Master Node (All 4 Components)
  • 🔹 What Lives Inside a Worker Node (All 3 Components)
  • 🔹 How Load is Distributed Across Worker Nodes
  • 🔹 What Happens When a Worker Node Fails
  • 🔹 What Happens When the Master Node Fails (High Availability)
  • 🔹 Scaling Up and Down — Auto-Scaling for ML Traffic Spikes
  • 🔹 Kubernetes on Oracle OCI (OKE) — Real Deployment
  • 🔹 Docker Swarm vs Kubernetes — Which One and When
  • 🔹 Full Working Example — ML Model Cluster on OCI OKE
  • 🔹 Best Practices for ML Clusters
  • 🔹 High-Level Summary

⚠️ Section 1: Why a Single Container Is Never Enough in Production

You've trained a fraud detection model. You've Dockerized it. You run one container on one OCI Compute instance. It works!

But then your company launches a marketing campaign. Traffic jumps from 100 requests/second to 50,000 requests/second. Your single container handles maybe 200 requests/second. The rest time out. Customers are blocked. The fraud model stops working. Your company loses money every second.

🔥 The 3 Production Problems a Single Container Can't Solve

  • Problem 1: Traffic Spikes
    One container can only handle so many simultaneous ML predictions. When 10,000 people hit your API at once, one container collapses. You need many containers running in parallel.
  • Problem 2: Container Crashes
    Containers crash. Memory leaks, buggy code, corrupted input data — all can bring a container down. If you only have one, your ML service goes offline. You need automatic restarts and failover.
  • Problem 3: Machine Failures
    The OCI Compute instance itself can fail — hardware fault, network issue, OCI maintenance. If all your containers are on one machine, they all die together. You need containers spread across multiple machines.
✅ Container Orchestration solves all three problems automatically:
  • ✅ Traffic spikes → automatically spawn more containers across multiple machines
  • ✅ Container crashes → automatically detect and restart failed containers
  • ✅ Machine failures → automatically move containers to healthy machines
All without any human intervention. The system heals itself. 🔄

🏗️ Section 2: What Is Container Orchestration? (The Big Idea)

Container Orchestration is an automated system that manages a group of computers (called a cluster) and decides:

  • Which machine runs which container
  • How many copies of each container to run
  • What to do when traffic increases or decreases
  • What to do when a container or machine crashes
  • How containers communicate with each other
  • How traffic from outside reaches the right container

The most popular orchestration tools:

Tool Full Name Best For Complexity
Kubernetes ⭐ K8s Production ML systems, OCI OKE High — industry standard
Docker Swarm Swarm Simple setups, small teams Low — built into Docker
Oracle OKE OCI Kubernetes Engine Managed K8s on OCI — no master setup needed Medium — OCI handles master
A KUBERNETES / DOCKER SWARM CLUSTER:

┌─────────────────────────────────────────────────────────┐
│ CLUSTER │
│ │
│ ┌──────────────────────────────┐ │
│ │ MASTER NODE 🧠 │ │
│ │ (The Restaurant Manager) │ │
│ │ - Watches everything │ │
│ │ - Makes all decisions │ │
│ │ - Assigns work to workers │ │
│ └──────────────┬───────────────┘ │
│ │ commands │
│ ┌────────────┼────────────┐ │
│ ▼ ▼ ▼ │
│ ┌───────┐ ┌───────┐ ┌───────┐ │
│ │WORKER │ │WORKER │ │WORKER │ │
│ │ 1 💪 │ │ 2 💪 │ │ 3 💪 │ │
│ │[ML] │ │[ML] │ │[ML] │ │
│ │[ML] │ │[ML] │ │[ML] │ │
│ └───────┘ └───────┘ └───────┘ │
│ │
└─────────────────────────────────────────────────────────┘

Each [ML] box = one running Docker container (your ML model API)

🧠 Section 3: The Master Node — The Brain of the Cluster

The Master Node (also called the Control Plane in Kubernetes) is the decision-making center of the entire cluster. It never runs your ML model containers directly. Instead, it focuses entirely on managing the cluster.

🎬 The Air Traffic Controller Analogy

Think of a busy airport. Dozens of planes are landing and taking off every minute. The air traffic controller in the tower doesn't fly any planes. But they watch every plane on radar, give each pilot instructions, and make sure no two planes collide or sit on the wrong runway.

The Master Node is your air traffic controller. Worker Nodes are the planes. ML containers are the passengers. ✈️

MASTER NODE INTERNALS:

┌─────────────────────────────────────────────────────────┐
│ MASTER NODE 🧠 │
│ │
│ ┌──────────────┐ ┌──────────────┐ │
│ │ API Server │ │ etcd │ │
│ │ 📡 Gateway │ │ 📚 Database │ │
│ │ (front door) │ │ (memory of │ │
│ │ │ │ everything) │ │
│ └──────────────┘ └──────────────┘ │
│ │
│ ┌──────────────┐ ┌──────────────┐ │
│ │ Scheduler │ │ Controller │ │
│ │ 📋 Planner │ │ Manager │ │
│ │ (assigns │ │ 🔄 Watchdog │ │
│ │ work) │ │ (fixes │ │
│ │ │ │ problems) │ │
│ └──────────────┘ └──────────────┘ │
└─────────────────────────────────────────────────────────┘

🔬 The 4 Components Inside a Master Node

Component 1: API Server 📡

The API Server is the front door of the entire cluster. Every command goes through it — whether you're deploying a new ML model, scaling from 3 to 10 containers, or checking the cluster health.

When you type kubectl apply -f fraud-model-deployment.yaml, that request hits the API Server first. It validates the request, authenticates you, and forwards it to the right component.

💡 Think of the API Server as the reception desk at a hotel. Every request — check in, room service, housekeeping — goes through reception first. Nobody talks directly to the chef or the cleaner. Everything is coordinated through one central point.

Component 2: etcd 📚

etcd is a distributed key-value database — the long-term memory of the cluster. It stores the desired state of everything: how many copies of your ML model should run, what resources each needs, which nodes exist, what's healthy and what's not.

When the master node restarts after a crash, it reads etcd to remember exactly what the cluster was doing before. Nothing is forgotten. 🧠

💡 etcd is like the restaurant's order book. Even if the manager goes home for the night, the order book remembers every active order. When the manager comes back, they pick up exactly where they left off.

Component 3: Scheduler 📋

The Scheduler decides which Worker Node runs which container. When you ask for 5 copies of your fraud model to run, the Scheduler looks at all available Worker Nodes and asks:

  • Which nodes have enough free RAM for this container?
  • Which nodes have the required CPU capacity?
  • Are any nodes already overloaded?
  • Are there any placement rules (e.g., "keep containers in the same OCI region")?
  • Are there GPU requirements (e.g., deep learning models need NVIDIA GPU nodes)?

It then assigns each container to the most suitable Worker Node — like a smart waiter assigning tables to the chef with the lowest workload and the right skills.

Component 4: Controller Manager 🔄

The Controller Manager is the watchdog of the cluster. It constantly compares the actual state of the cluster against the desired state stored in etcd. When they don't match, it takes action to fix things.

Examples of what it fixes automatically:

  • You asked for 5 ML containers. Only 4 are running. → Start a 5th container.
  • Worker Node 2 crashed. It had 2 containers. → Reschedule those on healthy nodes.
  • A container has been stuck in "starting" state for 10 minutes. → Kill it, try again.
  • You deleted a deployment. Containers are still running. → Terminate them.
✅ The Controller Manager is why Kubernetes is "self-healing." It runs in an infinite loop checking: "Is what's running what I want?" The moment the answer is "no", it starts fixing automatically. This is the core superpower of container orchestration for ML systems.

💪 Section 4: The Worker Node — The Muscle of the Cluster

Worker Nodes are the actual machines that run your ML model containers. Each Worker Node is an OCI Compute instance (VM or bare metal) with Docker installed and connected to the Master Node.

A Worker Node doesn't make any decisions on its own. It simply receives instructions from the Master and executes them: "Start this container", "Stop that container", "Report your resource usage."

WORKER NODE INTERNALS:

┌──────────────────────────────────────────────────────┐
│ WORKER NODE 💪 │
│ (OCI Compute Instance) │
│ │
│ ┌────────────┐ ┌────────────┐ ┌────────────┐ │
│ │ kubelet │ │ kube-proxy │ │ Container │ │
│ │ 👂 Ears │ │ 🌐 Network │ │ Runtime │ │
│ │ (listens │ │ Routing │ │ ⚙️ Docker │ │
│ │ to Master)│ │ │ │ │ │
│ └────────────┘ └────────────┘ └────────────┘ │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ ML Pod │ │ ML Pod │ │ ML Pod │ │
│ │[Container│ │[Container│ │[Container│ │
│ │ fraud │ │ fraud │ │ fraud │ │
│ │ model] │ │ model] │ │ model] │ │
│ └──────────┘ └──────────┘ └──────────┘ │
└──────────────────────────────────────────────────────┘

🔬 The 3 Components Inside Every Worker Node

Component 1: kubelet 👂

The kubelet is the agent on every Worker Node that maintains the communication channel with the Master. It receives "Pod specifications" from the Master (instructions describing which containers to run and with what settings), starts them, and reports back their health status regularly.

If a container crashes on the worker, kubelet notices immediately and reports it to the Master so the Controller Manager can react.

💡 kubelet is like a chef's sous chef. The head chef (Master) sends down the order: "Make 3 margherita pizzas with these ingredients." The sous chef (kubelet) executes the order, checks them, and reports back: "3 pizzas ready, all looking good." 🍕

Component 2: kube-proxy 🌐

kube-proxy handles network routing on the Worker Node. When a request arrives for your ML model API, kube-proxy decides which running container should receive that request. It also enables containers on different Worker Nodes to communicate with each other.

Component 3: Container Runtime ⚙️

The Container Runtime is the engine that actually runs containers. The most common runtime is containerd (which Docker also uses under the hood). It pulls images from OCIR, starts containers, allocates CPU and RAM, and stops/removes containers when instructed.

🫛 What is a Pod? (The Kubernetes Unit)

In Kubernetes, containers don't run alone — they run inside Pods. A Pod is the smallest deployable unit in Kubernetes. It's a wrapper around one or more containers that share the same network and storage.

💡 A Pod is like a bento box. The bento box (Pod) contains one main dish (your ML model container) and optionally a side dish (a logging sidecar container). The box has one address. Both dishes share the same dining tray (network namespace).

⚖️ Section 5: How Load Is Distributed Across Worker Nodes

When 10,000 users hit your ML model API simultaneously, the cluster distributes that load across all Worker Nodes automatically. Here's exactly how it works, step by step.

📬 The Load Distribution Flow

STEP 1: Request arrives
User → HTTP Request → OCI Load Balancer (external IP)

STEP 2: Load Balancer routes to a Node
OCI Load Balancer → picks a healthy Worker Node (round-robin)

STEP 3: kube-proxy routes to a Pod
Worker Node receives request → kube-proxy checks iptables rules
→ randomly picks one healthy ML Pod from the pool

STEP 4: Pod processes the request
ML Pod runs inference → returns prediction

STEP 5: Response travels back
Prediction → kube-proxy → OCI Load Balancer → User

TOTAL: sub-50ms for the entire round trip on a healthy cluster 🚀

🔄 Round-Robin Load Balancing

By default, Kubernetes distributes requests in a round-robin pattern — like dealing cards at a poker table. Request 1 goes to Pod A. Request 2 goes to Pod B. Request 3 goes to Pod C. Request 4 back to Pod A. And so on.

This ensures no single ML container gets overloaded while others sit idle.

🏋️ Resource-Based Scheduling

When the Scheduler places a new Pod on a Worker Node, it looks at resource requests and limits defined in your deployment YAML.

# Each ML model Pod gets these guaranteed resources:
resources:
  requests:
    memory: "512Mi"   # Minimum RAM guaranteed to this Pod
    cpu: "500m"       # Minimum CPU (500 millicores = half a CPU core)
  limits:
    memory: "2Gi"     # Maximum RAM this Pod can use
    cpu: "2000m"      # Maximum CPU (2 full cores)

The Scheduler will only place a Pod on a Worker Node that has enough free capacity to meet the requests. If no Worker Node has enough resources, the Pod stays in "Pending" state until capacity becomes available or a new Worker Node is added.

💡 For ML models, always set resource requests and limits!
Without them, a single ML model can consume all RAM on a Worker Node, starving other Pods of memory and causing cascading failures across the cluster. A greedy ML model is the most common cause of cluster instability.

💥 Section 6: What Happens When a Worker Node Fails

This is the scenario every ML engineer dreads: a machine dies in production. Here is exactly what Kubernetes does — automatically, without any human action.

🔴 The Node Failure Timeline

T+0:00    Worker Node 2 has a hardware fault and goes offline.
          It was running 3 ML model Pods.

T+0:05    kubelet on Worker Node 2 stops sending heartbeats to the Master.

T+0:40    Master Node notices: "Haven't heard from Worker 2 in 40 seconds."
          Node status changes from "Ready" → "NotReady".

T+5:00    After 5 minutes of NotReady status, Controller Manager acts.
          "Worker 2 is gone. Its 3 Pods must be rescheduled elsewhere."

T+5:05    Scheduler finds 2 healthy Workers (Node 1 and Node 3).
          It assigns 2 new Pods to Node 1 and 1 new Pod to Node 3.

T+5:15    New Pods pull the ML model image from OCIR and start up.

T+5:45    Health checks pass. New Pods start receiving traffic.

T+5:45    The 3 lost Pods are now fully replaced. Service restored. ✅
          Total downtime: ~45 seconds (during Pod startup).

Meanwhile: OCI detects the failed VM and provisions a replacement.
New Node 4 joins the cluster. Cluster is healthy again.

📊 Visual: Before and After Node Failure

BEFORE FAILURE (9 Pods across 3 Nodes):

Node 1 ✅: [ML][ML][ML]
Node 2 ✅: [ML][ML][ML] ← THIS NODE FAILS
Node 3 ✅: [ML][ML][ML]

DURING FAILURE (6 Pods, Node 2 offline):

Node 1 ✅: [ML][ML][ML]
Node 2 ❌: [ ][ ][ ] ← OFFLINE
Node 3 ✅: [ML][ML][ML]

AFTER RECOVERY (9 Pods, 3 rescheduled):

Node 1 ✅: [ML][ML][ML][ML][ML] ← gets 2 extra Pods
Node 2 ❌: offline (OCI replacing)
Node 3 ✅: [ML][ML][ML][ML] ← gets 1 extra Pod
(New Node 4 joins later and Pods redistribute again)
✅ Key Insight: Why You Should Run at Least 3 Replicas

With only 1 Pod: node failure = 100% downtime until restart.
With 2 Pods on 2 nodes: node failure = 50% capacity loss (1 Pod still running).
With 3+ Pods on 3+ nodes: node failure = 33% capacity loss (2/3 Pods still running).

For ML APIs serving real users: always run minimum 3 replicas spread across 3 nodes.

🧠💥 Section 7: What Happens When the Master Node Fails (High Availability)

Here's a question beginners always ask: "If the Master controls everything, what happens when the Master itself fails?"

This is a great question. In a single-master setup (development/test environments), if the Master fails, the existing Worker Nodes keep running their containers. But no new scheduling happens, no failures are recovered, no scaling occurs. The cluster is "frozen" until the master comes back.

For production ML systems, we use High Availability (HA) Master Setup — multiple Master Nodes running simultaneously.

HIGH AVAILABILITY MASTER SETUP (3 Masters):

┌────────────────────────────────────────────────────────────┐
│ CONTROL PLANE (HA) │
│ │
│ Master 1 (LEADER) ──── Master 2 (STANDBY) ──── Master 3 │
│ 🟢 Active 🟡 Ready 🟡 Ready │
│ etcd ◄─────────────── etcd ◄──────────────────── etcd │
│ (sync) (sync) (sync) │
│ │
│ If Master 1 dies → etcd election → Master 2 becomes │
│ new LEADER in ~2–3 seconds. Master 3 remains STANDBY. │
└────────────────────────────────────────────────────────────┘

OKE (Oracle Kubernetes Engine) manages this for you automatically!
You never touch the master nodes when using OKE.

🗳️ Leader Election — How Masters Vote

When a Master fails, the remaining Masters vote to elect a new Leader. This uses a consensus algorithm called Raft. You need a majority to elect a leader — that's why you always use odd numbers of Master Nodes (1, 3, 5 — never 2 or 4).

Masters Can Tolerate Failures Use Case
1 Master 0 failures (single point of failure) Development/testing only
3 Masters 1 failure (2 of 3 still form majority) Standard production setup
5 Masters 2 failures (3 of 5 still form majority) Critical ML systems (financial, healthcare)
✅ OCI OKE handles all of this for you automatically. When you create an OKE cluster, Oracle runs and manages 3 Master Nodes (the control plane) in a highly available configuration. You only pay for and manage the Worker Nodes. Never worry about Master Node failures on OCI OKE!

📈 Section 8: Scaling Up and Down — Auto-Scaling for ML Traffic Spikes

One of the most powerful features of container orchestration is automatic scaling. The cluster grows when traffic is high and shrinks when traffic is low — saving compute costs without any human intervention.

Three Types of Scaling in Kubernetes

1️⃣ Horizontal Pod Autoscaler (HPA) — More Copies

HPA automatically adds or removes Pod replicas based on CPU, RAM, or custom metrics (like requests-per-second or ML inference queue depth).

NORMAL TRAFFIC (2 PM Tuesday):
  Requests/sec: 200     → CPU usage: 15%
  HPA target: 70% CPU → 3 Pods running (plenty of headroom)

TRAFFIC SPIKE (5 PM Friday — end of trading day):
  Requests/sec: 8,000 → CPU usage: 85% (above 70% target!)
  HPA detects: "Need more capacity"
  HPA scales: 3 Pods → 8 Pods (in ~30 seconds)
  CPU drops to: 32% → Stable ✅

TRAFFIC DROPS (11 PM):
  Requests/sec: 50 → CPU usage: 4%
  HPA scales down: 8 Pods → 2 Pods (after 5 min cooldown)
  Cost savings: running 6 fewer containers overnight 💰
# fraud-model-hpa.yaml — Horizontal Pod Autoscaler
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: fraud-model-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: fraud-model-deployment

  minReplicas: 2    # Never go below 2 Pods (always have spare capacity)
  maxReplicas: 20   # Never go above 20 Pods (cost control)

  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70   # Scale up when CPU > 70%

    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: 80   # Scale up when RAM > 80%

2️⃣ Vertical Pod Autoscaler (VPA) — Bigger Containers

VPA adjusts the CPU and RAM allocated to each Pod. If your ML model is consistently using 3.5 GB RAM but you only allocated 2 GB, VPA automatically increases the allocation to 4 GB on the next restart.

3️⃣ Cluster Autoscaler — More Machines

The Cluster Autoscaler adds or removes entire Worker Nodes (OCI Compute instances) from the cluster.

SITUATION: HPA wants to add 5 more ML Pods.
PROBLEM: No Worker Node has enough free RAM for 5 new Pods.
Pods stuck in "Pending" state...

CLUSTER AUTOSCALER DETECTS: "Pods can't be scheduled — nodes are full."
CLUSTER AUTOSCALER ACTS: "Provision 2 new OCI Compute instances."

OCI provisions VM.Standard.E4.Flex with 16 OCPUs, 128 GB RAM.
New Worker Nodes join the cluster.
Pending Pods get scheduled onto new Nodes. ✅

LATER (traffic drops):
HPA scales down Pods. Cluster Autoscaler sees empty Nodes.
It terminates the empty OCI instances. Cost drops immediately. 💰

☁️ Section 9: Kubernetes on Oracle OCI (OKE) — Real Deployment

Oracle Container Engine for Kubernetes (OKE) is Oracle's fully managed Kubernetes service. Oracle runs and maintains the Master Node (Control Plane) for you. You provision and pay only for the Worker Nodes.

📌 Code Purpose — Deploy ML Model to OCI OKE

What this code does: Creates a complete Kubernetes deployment for a fraud detection ML model on OCI OKE. Defines 3 replicas of the ML API Pod, connects to OCIR for the image, sets resource limits, configures a LoadBalancer Service to expose the API externally, and adds a HorizontalPodAutoscaler for automatic scaling.

Why OKE over self-managed K8s: Oracle manages all 3 Master Nodes, upgrades, backups, and HA. You focus only on deploying your ML model — not managing infrastructure.
# ─────────────────────────────────────────────────────────────────────
# fraud-model-deployment.yaml
# Deploy fraud detection ML model to OCI OKE
# kubectl apply -f fraud-model-deployment.yaml
# ─────────────────────────────────────────────────────────────────────

# ── PART 1: Deployment (manages Pod replicas) ─────────────────────────
apiVersion: apps/v1
kind: Deployment
metadata:
  name: fraud-model-deployment
  namespace: ml-production
  labels:
    app: fraud-model
    version: "2.1.0"
    team: ml-engineering
spec:
  replicas: 3          # Start with 3 Pods (1 per Worker Node ideally)

  selector:
    matchLabels:
      app: fraud-model

  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1    # At most 1 Pod offline during updates
      maxSurge: 1          # At most 1 extra Pod during updates
                           # → zero-downtime deployments! ✅

  template:
    metadata:
      labels:
        app: fraud-model
    spec:

      # Pull image from Oracle Container Registry (OCIR)
      imagePullSecrets:
        - name: ocir-secret    # K8s Secret with OCIR credentials

      containers:
        - name: fraud-api
          image: ap-mumbai-1.ocir.io/mytenancy/fraud-model:2.1.0

          ports:
            - containerPort: 8083

          # Resource limits prevent greedy ML models from starving cluster
          resources:
            requests:
              memory: "512Mi"
              cpu: "500m"
            limits:
              memory: "2Gi"
              cpu: "2000m"

          # Environment variables
          env:
            - name: ENVIRONMENT
              value: "production"
            - name: OCI_REGION
              value: "ap-mumbai-1"
            - name: LOG_LEVEL
              value: "INFO"

          # Health check — Master uses this to know if Pod is ready for traffic
          readinessProbe:
            httpGet:
              path: /health
              port: 8083
            initialDelaySeconds: 15   # Wait 15s for model to load first
            periodSeconds: 10
            failureThreshold: 3       # 3 failed checks → remove from load balancer

          # Liveness check — Master uses this to know if Pod should be restarted
          livenessProbe:
            httpGet:
              path: /health
              port: 8083
            initialDelaySeconds: 30
            periodSeconds: 30
            failureThreshold: 3       # 3 failed checks → kill and restart Pod

      # Spread Pods across different Worker Nodes (fault tolerance!)
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: kubernetes.io/hostname
          whenUnsatisfiable: DoNotSchedule
          labelSelector:
            matchLabels:
              app: fraud-model

---
# ── PART 2: Service (exposes Pods via OCI Load Balancer) ──────────────
apiVersion: v1
kind: Service
metadata:
  name: fraud-model-service
  namespace: ml-production
  annotations:
    # OCI-specific: create an OCI Load Balancer with 100 Mbps bandwidth
    service.beta.kubernetes.io/oci-load-balancer-shape: "flexible"
    service.beta.kubernetes.io/oci-load-balancer-shape-flex-min: "10"
    service.beta.kubernetes.io/oci-load-balancer-shape-flex-max: "100"
spec:
  type: LoadBalancer    # OCI automatically provisions an OCI Load Balancer
  selector:
    app: fraud-model    # Routes traffic to all Pods with this label
  ports:
    - protocol: TCP
      port: 80           # External port (users call http://x.x.x.x/predict)
      targetPort: 8083   # Internal Pod port

---
# ── PART 3: HorizontalPodAutoscaler ───────────────────────────────────
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: fraud-model-hpa
  namespace: ml-production
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: fraud-model-deployment
  minReplicas: 3
  maxReplicas: 30
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70
# ── DEPLOY TO OKE ────────────────────────────────────────────────────

# Set up kubectl to connect to your OKE cluster
oci ce cluster create-kubeconfig \
    --cluster-id ocid1.cluster.oc1.ap-mumbai-1.aaaaa... \
    --file $HOME/.kube/config \
    --region ap-mumbai-1 \
    --token-version 2.0.0

# Create the namespace
kubectl create namespace ml-production

# Create OCIR pull secret (so Kubernetes can pull from your private OCIR)
kubectl create secret docker-registry ocir-secret \
    --docker-server=ap-mumbai-1.ocir.io \
    --docker-username="mytenancy/alice@company.com" \
    --docker-password="$OCI_AUTH_TOKEN" \
    --namespace ml-production

# Apply all resources at once
kubectl apply -f fraud-model-deployment.yaml

# ── WATCH THE DEPLOYMENT HAPPEN IN REAL TIME ─────────────────────────

# Watch Pods being created and scheduled to Worker Nodes
kubectl get pods -n ml-production -w

# NAME                                    READY   STATUS              NODE
# fraud-model-deployment-7d8f9b-x2k4p     0/1     ContainerCreating   worker-1
# fraud-model-deployment-7d8f9b-m9p3q     0/1     ContainerCreating   worker-2
# fraud-model-deployment-7d8f9b-n7r8s     0/1     ContainerCreating   worker-3
# fraud-model-deployment-7d8f9b-x2k4p     1/1     Running             worker-1 ✅
# fraud-model-deployment-7d8f9b-m9p3q     1/1     Running             worker-2 ✅
# fraud-model-deployment-7d8f9b-n7r8s     1/1     Running             worker-3 ✅

# Get the external IP of the OCI Load Balancer
kubectl get service fraud-model-service -n ml-production

# NAME                    TYPE           EXTERNAL-IP       PORT(S)
# fraud-model-service     LoadBalancer   140.238.12.34     80:32456/TCP ✅

# Test the deployed ML API
curl http://140.238.12.34/health
# {"status":"healthy","environment":"production","replicas":3}

🔄 Section 10: Zero-Downtime Updates — Rolling Updates Explained

When you train a new, better fraud model and want to deploy it, you don't want your API to go offline during the update. Kubernetes uses Rolling Updates to replace old Pods with new ones one at a time — so the service keeps running throughout.

ROLLING UPDATE TIMELINE (updating fraud model v2.1.0 → v2.2.0):

START: 3 Pods running v2.1.0 — all healthy, handling traffic
[Pod A: v2.1.0 ✅] [Pod B: v2.1.0 ✅] [Pod C: v2.1.0 ✅]

STEP 1: Start new Pod D with v2.2.0. Wait for it to pass health checks.
[Pod A: v2.1.0 ✅] [Pod B: v2.1.0 ✅] [Pod C: v2.1.0 ✅] [Pod D: v2.2.0 🆕]

STEP 2: Terminate old Pod A. Traffic now served by B, C, D.
[Pod B: v2.1.0 ✅] [Pod C: v2.1.0 ✅] [Pod D: v2.2.0 ✅]

STEP 3: Start new Pod E with v2.2.0. Wait for health checks.
[Pod B: v2.1.0 ✅] [Pod C: v2.1.0 ✅] [Pod D: v2.2.0 ✅] [Pod E: v2.2.0 🆕]

STEP 4: Terminate Pod B.
[Pod C: v2.1.0 ✅] [Pod D: v2.2.0 ✅] [Pod E: v2.2.0 ✅]

STEP 5: Start Pod F with v2.2.0. Terminate Pod C.
[Pod D: v2.2.0 ✅] [Pod E: v2.2.0 ✅] [Pod F: v2.2.0 ✅]

RESULT: All 3 Pods now on v2.2.0. Zero downtime. Zero dropped requests. ✅
# Update ML model image version — triggers rolling update automatically!
kubectl set image deployment/fraud-model-deployment \
    fraud-api=ap-mumbai-1.ocir.io/mytenancy/fraud-model:2.2.0 \
    -n ml-production

# Watch the rolling update happen
kubectl rollout status deployment/fraud-model-deployment -n ml-production
# Waiting for deployment "fraud-model-deployment" rollout to finish...
# 1 out of 3 new replicas have been updated...
# 2 out of 3 new replicas have been updated...
# 3 out of 3 new replicas have been updated...
# deployment "fraud-model-deployment" successfully rolled out ✅

# Something wrong with the new model? INSTANT ROLLBACK:
kubectl rollout undo deployment/fraud-model-deployment -n ml-production
# Rolled back to v2.1.0 in seconds!

🥊 Section 11: Docker Swarm vs Kubernetes — Which One and When

Both Docker Swarm and Kubernetes solve the Master/Worker orchestration problem. But they're built for different scenarios. Here's how to choose.

Feature Docker Swarm Kubernetes (OKE)
Setup complexity ⭐ Very easy — built into Docker ⭐⭐⭐ Complex (easy on OKE)
Learning curve Low — 1–2 days High — weeks to months
Auto-scaling Manual or basic ✅ Full HPA + VPA + Cluster Autoscaler
Ecosystem Limited ✅ Huge (Helm, Istio, Prometheus, etc.)
GPU scheduling Limited ✅ Full NVIDIA GPU plugin support
ML platform support Basic ✅ Kubeflow, Seldon, KServe, Ray
Industry adoption Declining ✅ Dominant (95% of orgs)
Best for Hobby projects, learning basics All production ML systems
✅ Recommendation:

Learning / hobby project: Use Docker Swarm first — understand master/worker concepts with zero setup.

Production ML system: Use OCI OKE (managed Kubernetes) — Oracle handles the master plane, you focus on deploying models.

Deep learning with GPUs: OKE with GPU node pools (OCI VM.GPU.A10.1 worker nodes) — Kubernetes is the only serious option here.

🐝 Docker Swarm — Quick Setup for Learning

# ── DOCKER SWARM SETUP (3 OCI instances) ─────────────────────────────

# On the MANAGER node (Master) — initialize the swarm
docker swarm init --advertise-addr 10.0.0.1
# Output: docker swarm join --token SWMTKN-1-abc123... 10.0.0.1:2377

# On each WORKER node — join the swarm
# (run the command from the output above on Worker 1 and Worker 2)
docker swarm join --token SWMTKN-1-abc123... 10.0.0.1:2377
# "This node joined a swarm as a worker." ✅

# From MANAGER: verify all nodes joined
docker node ls

# ID              HOSTNAME    STATUS    AVAILABILITY    MANAGER STATUS
# abc123 *        manager-1   Ready     Active          Leader
# def456          worker-1    Ready     Active
# ghi789          worker-2    Ready     Active

# Deploy ML model as a Swarm service (3 replicas across workers)
docker service create \
    --name fraud-api \
    --replicas 3 \
    --publish published=80,target=8083 \
    --update-parallelism 1 \
    --update-delay 10s \
    ap-mumbai-1.ocir.io/mytenancy/fraud-model:2.1.0

# Watch pods (called "tasks" in Swarm) distribute across nodes
docker service ps fraud-api

# ID           NAME           NODE        DESIRED STATE   CURRENT STATE
# xyz111       fraud-api.1    worker-1    Running         Running 12s
# xyz222       fraud-api.2    worker-2    Running         Running 11s
# xyz333       fraud-api.3    manager-1   Running         Running 10s

# Scale up to 6 replicas (simulate traffic spike)
docker service scale fraud-api=6

# Update image version (rolling update, zero downtime)
docker service update --image fraud-model:2.2.0 fraud-api

🧭 Section 12: Complete Cluster Health Monitoring

Once your ML cluster is running, you need to monitor the health of both Master and Worker nodes continuously.

# ════════════════════════════════════════════════════════════════
#  CLUSTER-LEVEL HEALTH COMMANDS
# ════════════════════════════════════════════════════════════════

# Check all nodes in the cluster
kubectl get nodes -o wide

# NAME         STATUS   ROLES    AGE   VERSION   INTERNAL-IP   OS-IMAGE
# master-1     Ready    master   5d    v1.29.1   10.0.1.10     Oracle Linux 8
# worker-1     Ready    worker   5d    v1.29.1   10.0.1.11     Oracle Linux 8
# worker-2     Ready    worker   5d    v1.29.1   10.0.1.12     Oracle Linux 8
# worker-3     Ready    worker   5d    v1.29.1   10.0.1.13     Oracle Linux 8

# Check detailed node health (CPU, RAM usage)
kubectl top nodes

# NAME         CPU(cores)   CPU%   MEMORY(bytes)   MEMORY%
# worker-1     1240m        31%    4.2Gi           53%
# worker-2     890m         22%    3.8Gi           48%
# worker-3     1560m        39%    5.1Gi           64%

# Describe a specific node (find out WHY it might be unhealthy)
kubectl describe node worker-2

# Check all Pods across all namespaces
kubectl get pods --all-namespaces

# Check Pods in our ML namespace with resource usage
kubectl top pods -n ml-production

# NAME                                    CPU(cores)   MEMORY(bytes)
# fraud-model-deployment-7d8f9b-x2k4p     245m         812Mi
# fraud-model-deployment-7d8f9b-m9p3q     189m         778Mi
# fraud-model-deployment-7d8f9b-n7r8s     312m         891Mi

# Check HPA status (scaling decisions)
kubectl get hpa -n ml-production

# NAME                REFERENCE               TARGETS      MINPODS   MAXPODS   REPLICAS
# fraud-model-hpa     Deployment/fraud-model   45%/70%      3         30        3

# ════════════════════════════════════════════════════════════════
#  SIMULATING AND RECOVERING FROM A NODE FAILURE
# ════════════════════════════════════════════════════════════════

# Simulate draining a node (like before OCI maintenance)
# Safely moves all Pods off the node first
kubectl drain worker-2 --ignore-daemonsets --delete-emptydir-data

# Node is now empty — safe for maintenance, reboot, OCI shape change
# Pods have been rescheduled to worker-1 and worker-3

# After maintenance — bring node back into service
kubectl uncordon worker-2
# Scheduler will start placing Pods on worker-2 again

# Check events to understand what happened on a node
kubectl get events -n ml-production --sort-by='.lastTimestamp' | tail -20

✅ Section 13: Best Practices for ML Clusters

✅ DOs — What Every MLOps Engineer Should Do:
  • ✅ Always run minimum 3 replicas of ML model Pods for fault tolerance
  • ✅ Set resource requests AND limits — prevents one greedy model from crashing the cluster
  • ✅ Use topologySpreadConstraints to spread Pods across different Worker Nodes
  • ✅ Add readinessProbe to every ML Pod — prevents traffic routing to Pods still loading the model
  • ✅ Use Rolling Update strategy — never take your ML API offline for updates
  • ✅ Set up HPA — let the cluster auto-scale during traffic spikes
  • ✅ Use OCI OKE — let Oracle manage the Master Nodes so you focus on ML, not infrastructure
  • ✅ Store images in OCIR with specific version tags — never use :latest in production deployments
  • ✅ Always have a rollback plan — kubectl rollout undo is your safety net
  • ✅ Use namespaces to isolate environments (ml-dev, ml-staging, ml-production)
❌ DON'Ts — Mistakes That Take Down ML Clusters:
  • ❌ Don't run ML models without resource limits — one leaky model tanks the whole cluster
  • ❌ Don't put all replicas on the same Worker Node — defeats the purpose of clustering
  • ❌ Don't skip readinessProbes — users get errors during model warm-up without them
  • ❌ Don't use a single Master Node for production — one hardware fault = entire cluster frozen
  • ❌ Don't load large ML models (5 GB+) inside the container image — use OCI Object Storage and download at startup
  • ❌ Don't ignore OOMKilled events — if your ML Pod keeps getting killed, you need higher memory limits
  • ❌ Don't forget to set HPA maxReplicas — an infinite scale-up can bankrupt your OCI bill

🏆 High-Level Summary !

  • 🔹 Single containers fail under traffic spikes, crashes, and machine failures — you need a cluster.
  • 🔹 Master Node = the brain. Never runs ML models. Only makes decisions (API Server + etcd + Scheduler + Controller Manager).
  • 🔹 Worker Node = the muscle. Runs your ML containers as Pods (kubelet + kube-proxy + container runtime).
  • 🔹 Load distribution: OCI LB → round-robin across Nodes → kube-proxy → Pod. Automatic, fast, fair.
  • 🔹 Worker failure: Master detects in 40s → reschedules Pods to healthy nodes in 5 min → full recovery. Automatic.
  • 🔹 Master failure: HA setup with 3 Masters + etcd election selects new leader in ~3 seconds. OKE handles this for you.
  • 🔹 HPA = auto-add/remove Pods. VPA = auto-resize Pods. Cluster Autoscaler = auto-add/remove OCI Compute instances.
  • 🔹 Rolling Updates = deploy new model versions with zero downtime. One Pod at a time. Health checks gate each step.
  • 🔹 Docker Swarm = simple, built-in, great for learning. Kubernetes / OKE = production standard for all ML systems.
  • 🔹 Oracle OKE = Oracle manages the Master. You manage Worker Node pools. Best of both worlds for ML on OCI.
You started knowing nothing about Masters and Workers. Now you understand why single containers fail, how traffic gets distributed, what happens during node failures, how the cluster self-heals, and how to deploy a real ML model to OCI OKE with zero downtime.

Every major ML product in the world — recommendation systems, fraud detection, image recognition, LLM inference — runs on exactly this architecture.

Your ML models are now production-grade, fault-tolerant, and infinitely scalable. 🐳☁️✨

Keep clustering, keep scaling, keep shipping! 🐼✨

Comments