Skip to main content

Kubernetes Node Services Explained: API Server, etcd, Scheduler, Controller & Kubelet

Calculating read time…

Let's imagine you manage a giant school. The school has a Principal's Office and many Classrooms.

The Principal's Office keeps records, makes decisions, assigns teachers to rooms, and monitors whether every class is running correctly. The Classrooms are where the actual teaching happens.

Now swap "teaching" for "running your ML model container" — and you have exactly how Kubernetes works.




💡 The School Analogy:

🏫 Principal's Office = Master Node
   — API Server = the reception desk
   — etcd = the filing cabinet
   — Scheduler = the timetable planner
   — Controller = the school inspector

🏫 Classrooms = Worker Nodes
   — Container Runtime = the whiteboard & chalk
   — Kubelet = the classroom teacher

Together they run your ML models reliably, at scale, on Oracle OCI — 24 hours a day, 7 days a week. 🚀

📚 What We'll Cover

  • 🔹 The Big Picture — How All 6 Services Fit Together
  • 🔹 Master Node Service 1: API Server — The Front Door
  • 🔹 Master Node Service 2: etcd — The Brain's Memory
  • 🔹 Master Node Service 3: Scheduler — The Smart Planner
  • 🔹 Master Node Service 4: Controller Manager — The Self-Healing Watchdog
  • 🔹 Worker Node Service 1: Container Runtime — The Engine Room
  • 🔹 Worker Node Service 2: Kubelet — The Loyal Agent
  • 🔹 How All 6 Services Work Together — Full ML Deployment Walkthrough
  • 🔹 What Happens During a Failure — Service-by-Service Recovery
  • 🔹 Practical Commands to Inspect Each Service
  • 🔹 Best Practices & Common Mistakes
  • 🔹 High-Level Summary

🗺️ Section 1: The Big Picture — All 6 Services in One View

Before diving into each service, let's see them all together. This diagram is the mental model you'll use for everything else in this post.

┌──────────────────────────────────────────────────────────────────┐
│               MASTER NODE (Control Plane)                │
│                                                         │
│  ┌────────────────┐  ┌────────────────┐                │
│  │ API Server     │  │ etcd Store     │                │
│  │ 📡 The Gateway │  │ 📚 The Memory  │                │
│  └────────────────┘  └────────────────┘                │
│                                                         │
│  ┌────────────────┐  ┌────────────────┐                │
│  │ Scheduler      │  │ Controller Mgr │                │
│  │ 📋 The Planner │  │ 🔄 The Fixer   │                │
│  └────────────────┘  └────────────────┘                │
└────────────────────────────┬─────────────────────────────────────┘
                            │ commands via API Server
               ┌───────────┴───────────┐
               ▼                       ▼
┌──────────────────────────┐   ┌──────────────────────────┐
│  WORKER NODE 1            │   │  WORKER NODE 2            │
│  ┌──────────────────────┐   │   │  ┌──────────────────────┐   │
│  │ Container Runtime   │   │   │  │ Container Runtime   │   │
│  │ ⚙️ Actually runs    │   │   │  │ ⚙️ Actually runs    │   │
│  │   containers        │   │   │  │   containers        │   │
│  └──────────────────────┘   │   │  └──────────────────────┘   │
│  ┌──────────────────────┐   │   │  ┌──────────────────────┐   │
│  │ Kubelet              │   │   │  │ Kubelet              │   │
│  │ 👂 Reports to Master │   │   │  │ 👂 Reports to Master │   │
│  └──────────────────────┘   │   │  └──────────────────────┘   │
└──────────────────────────┘   └──────────────────────────┘
Quick Cheat Sheet — Who Does What:
  • 🟣 API Server — Every request comes through here first. Nothing bypasses it.
  • 🟣 etcd — Remembers the state of everything. The cluster's notebook.
  • 🟣 Scheduler — Decides which Worker Node gets which container. The matchmaker.
  • 🟣 Controller Manager — Watches everything. Fixes problems automatically.
  • 🟠 Container Runtime — Actually starts and stops containers on the Worker.
  • 🟠 Kubelet — The Worker's agent. Listens to Master, reports back health.

🧠 MASTER NODE SERVICES

The Master Node runs 4 core services. None of them run your ML model containers directly. Their only job is to manage the entire cluster.


📡 Master Service 1: API Server — The Front Door of Everything

🏨 The Hotel Reception Analogy

Imagine a grand hotel. You want a room, room service, housekeeping, maintenance — everything goes through the front desk receptionist. You never walk directly into the kitchen to order food. You never go to the maintenance room yourself. Everything is coordinated through one central point.

The API Server is that receptionist. Every single interaction with the Kubernetes cluster — deploying an ML model, scaling up containers, checking cluster health — must go through the API Server first. Nothing talks to anything else directly.

EVERYTHING FLOWS THROUGH THE API SERVER:

kubectl (your laptop) ──────────────────▶ API Server
OCI OKE Dashboard ──────────────────────▶ API Server
CI/CD Pipeline (OCI DevOps) ────────────▶ API Server
Scheduler ──────────────────────────────▶ API Server
Controller Manager ─────────────────────▶ API Server
Kubelet on Worker Nodes ────────────────▶ API Server

NO direct connections between components.
Everything passes through the API Server. Always. 📡

🔒 What the API Server Actually Does

  • Authentication — "Who are you?" Checks your OCI IAM credentials or Kubernetes service account token. Rejects unknown callers immediately.
  • Authorization — "Are you allowed to do this?" Checks RBAC (Role-Based Access Control) rules. A data scientist can deploy models but cannot delete Worker Nodes.
  • Validation — "Is your request well-formed?" Checks that your YAML is correctly structured before accepting it. Rejects requests that reference non-existent images or invalid CPU values.
  • Routing — "Who needs to act on this?" Forwards validated requests to the right component: scheduling requests to the Scheduler, storage to etcd, worker commands to Kubelet.
✅ Why centralizing through the API Server is genius:

If every component talked to every other component directly, you'd have a chaotic web of connections impossible to secure or debug. With one gateway: one place to add security, one place to log requests, one place to rate-limit, one place to audit. Clean. Scalable. Secure. 🔐

💻 API Server in Action

📌 Code Purpose — Interacting with the API Server

What this shows: Every kubectl command is actually an HTTP request to the API Server. These examples show what's happening under the hood and how to interact with the API Server directly, which is useful for debugging and automation in OCI OKE deployments.
# Every kubectl command hits the API Server first
# kubectl is just a user-friendly wrapper around API Server HTTP calls

# Check API Server is reachable and healthy
kubectl cluster-info

# Kubernetes control plane is running at https://xxx.xxx.oraclecloud.com:6443
# CoreDNS is running at https://xxx.xxx.oraclecloud.com:6443/api/v1/...

# See your current connection context (which cluster's API Server you're talking to)
kubectl config current-context

# List all available API endpoints the API Server exposes
kubectl api-resources

# NAME                    SHORTNAMES   APIVERSION   NAMESPACED   KIND
# pods                    po           v1           true         Pod
# deployments             deploy       apps/v1      true         Deployment
# services                svc          v1           true         Service

# Call API Server DIRECTLY (no kubectl wrapper) to see raw HTTP
kubectl get pods -n ml-production -v=8
# -v=8 shows full HTTP request/response between kubectl and API Server
# You'll see: GET https://api-server-url/api/v1/namespaces/ml-production/pods
#             Response: 200 OK + JSON list of all pods

# Check API Server component health
kubectl get componentstatus
# NAME                 STATUS    MESSAGE
# scheduler            Healthy   ok
# controller-manager   Healthy   ok
# etcd-0               Healthy   {"health":"true"}
💡 On OCI OKE: Oracle manages the API Server for you. It runs on Oracle-managed infrastructure with automatic HA. The API Server endpoint is an HTTPS URL that OCI provides when you create your OKE cluster. You never SSH into the master node or restart the API Server manually.

📚 Master Service 2: etcd — The Cluster's Memory

🗄️ The Filing Cabinet Analogy

Every school has a central records office — a place that stores every student's file: which class they're in, what grades they have, who their teacher is, their attendance records.

If the records office burns down and everything is lost, the school is in chaos. Nobody knows who belongs where. etcd is that records office for Kubernetes. It stores the entire state of the cluster — what should be running, what is running, how many copies of each thing, all configurations. Lose etcd and the cluster has amnesia. 🧠💾

🔬 What etcd Actually Stores

ETCD KEY-VALUE STORE — SAMPLE CONTENTS:

/registry/pods/ml-production/fraud-model-pod-1
  → {name, image, resources, status, node: "worker-1"}

/registry/pods/ml-production/fraud-model-pod-2
  → {name, image, resources, status, node: "worker-2"}

/registry/deployments/ml-production/fraud-model
  → {replicas: 5, image: "ocir.../fraud-model:2.1", strategy: RollingUpdate}

/registry/nodes/worker-1
  → {status: Ready, CPU: 4, RAM: 16GB, pods: 3}

/registry/nodes/worker-2
  → {status: NotReady, last-heartbeat: 5 mins ago}

/registry/services/ml-production/fraud-api-service
  → {type: LoadBalancer, port: 80, targetPort: 8083}

⚡ Desired State vs Actual State — The Core Concept

etcd stores two critical things for every resource:

  • Desired State — What you want to be running. "I want 5 copies of my fraud detection model running." This is what you wrote in your YAML and applied with kubectl apply.
  • Actual State — What is currently running. "Right now, only 4 copies are running because Pod 3 crashed." This is reported by Kubelets on Worker Nodes.

The gap between desired and actual state is what triggers action. The Controller Manager watches for this gap and closes it automatically. etcd is the source of truth for both values.

💡 etcd is a distributed database. In a production HA cluster (like OCI OKE), etcd runs on 3 or 5 nodes simultaneously. All copies stay in sync using the Raft consensus algorithm. If one etcd node fails, the others still have all the data. This is why etcd is never a single point of failure in OKE.

OCI OKE backs up etcd automatically. You can restore an entire cluster from an etcd backup if disaster strikes.

💻 Exploring etcd (What's Inside)

📌 Code Purpose — Reading and Understanding etcd State

What this shows: You normally never interact with etcd directly — you use kubectl which reads/writes etcd through the API Server. These commands show the etcd-equivalent state through kubectl, helping you understand what's stored in etcd at any given moment.
# etcd stores "desired state" — you set this with kubectl apply
# kubectl get reads the current desired + actual state FROM etcd via API Server

# See the DESIRED state of your ML deployment (stored in etcd)
kubectl get deployment fraud-model-deployment -n ml-production -o yaml

# apiVersion: apps/v1
# kind: Deployment
# spec:
#   replicas: 5          ← DESIRED: 5 Pods
#   template:
#     spec:
#       containers:
#         image: ap-mumbai-1.ocir.io/mytenancy/fraud-model:2.1.0

# See the ACTUAL state (what's really running vs desired)
kubectl get pods -n ml-production

# NAME                               READY   STATUS    RESTARTS
# fraud-model-7d8f9b-x2k4p           1/1     Running   0         ← Pod 1: OK
# fraud-model-7d8f9b-m9p3q           1/1     Running   0         ← Pod 2: OK
# fraud-model-7d8f9b-n7r8s           0/1     Pending   0         ← Pod 3: PENDING!
# fraud-model-7d8f9b-p4q5r           1/1     Running   0         ← Pod 4: OK
# (Only 4 running, 1 pending — desired is 5, actual is 4 — etcd sees this gap!)

# Watch etcd-tracked events in real time
kubectl get events -n ml-production --sort-by='.lastTimestamp'

# REASON              MESSAGE
# Scheduled           Successfully assigned fraud-model-pod to worker-1
# Pulled              Container image pulled from OCIR
# Started             Container started successfully
# BackOff             Container keeps crashing — back-off restarting

📋 Master Service 3: Scheduler — The Smart Placement Engine

🏫 The School Timetable Analogy

At school, there's a timetable planner who must assign every class to a room for every time slot. They can't put 50 students in a room built for 20. They can't schedule the chemistry lab and cooking class in the same room. They need to balance load across all available rooms.

The Scheduler is that timetable planner for Kubernetes. When a new ML model Pod needs to run, the Scheduler asks: "Which Worker Node is the best fit for this Pod right now?" Then it makes a decision and records it in etcd via the API Server.

🔍 How the Scheduler Makes Its Decision — Step by Step

NEW POD REQUEST: "Run fraud-model container, needs 1 CPU, 2GB RAM"

STEP 1 — FILTERING (eliminate unsuitable nodes)
  Worker 1: 4 CPUs, 16 GB RAM, currently using 3.5 CPU → only 0.5 free → ❌ SKIP
  Worker 2: 4 CPUs, 16 GB RAM, currently using 2 CPU → 2 free → ✅ KEEP
  Worker 3: 8 CPUs, 32 GB RAM, currently using 1 CPU → 7 free → ✅ KEEP
  Worker 4: NotReady status → ❌ SKIP

STEP 2 — SCORING (rank the suitable nodes)
  Worker 2: score 62/100 (has enough, not much spare)
  Worker 3: score 88/100 (has plenty of spare capacity)

STEP 3 — BINDING (assign the Pod)
  Winner: Worker 3 🏆
  Scheduler writes to etcd: "fraud-model-pod-5 → Worker 3"
  Kubelet on Worker 3 picks this up and starts the container.

🎯 Scheduler Scoring Factors

  • Available CPU & RAM — prefer nodes with more free resources
  • Node affinity rules — "only place on GPU nodes" (critical for deep learning models that need OCI GPU shapes)
  • Pod anti-affinity — "spread replicas across different nodes" (never put all fraud model copies on the same machine)
  • Taints and tolerations — special labels that say "this node is reserved for ML workloads only — no other Pods allowed"
  • Data locality — prefer the node closest to the data (e.g., same OCI Availability Domain as the object storage)

💻 Influencing the Scheduler's Decisions

📌 Code Purpose — Scheduler Directives in Pod YAML

What this shows: How to give the Scheduler hints and rules about where to place your ML model Pods. These YAML snippets tell the Scheduler "put this ML Pod only on GPU nodes" and "never put two copies of this Pod on the same physical machine" — ensuring both performance and fault tolerance for ML deployments on OCI OKE.
# fraud-model-deployment.yaml — Scheduler directives

spec:
  template:
    spec:

      # RULE 1: Only schedule on GPU Worker Nodes (for deep learning models)
      nodeSelector:
        accelerator: "nvidia-a10"    # OCI VM.GPU.A10.1 nodes have this label

      # RULE 2: Spread replicas across different Worker Nodes (fault tolerance)
      # Never put two fraud-model Pods on the same node
      affinity:
        podAntiAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            - labelSelector:
                matchLabels:
                  app: fraud-model     # match our own pods
              topologyKey: kubernetes.io/hostname
              # topologyKey = "different hostnames" = different physical nodes

      # RULE 3: Resource requirements (Scheduler filters nodes without this)
      containers:
        - name: fraud-api
          resources:
            requests:
              memory: "2Gi"          # Need minimum 2GB RAM
              cpu: "1000m"           # Need minimum 1 full CPU core
            limits:
              memory: "4Gi"          # Max 4GB RAM
              cpu: "2000m"           # Max 2 CPU cores
# See where the Scheduler placed each Pod (which Worker Node)
kubectl get pods -n ml-production -o wide

# NAME                           NODE        STATUS
# fraud-model-7d8f9b-x2k4p       worker-1    Running   ← assigned to worker-1
# fraud-model-7d8f9b-m9p3q       worker-2    Running   ← assigned to worker-2
# fraud-model-7d8f9b-n7r8s       worker-3    Running   ← assigned to worker-3
# (all on DIFFERENT nodes due to podAntiAffinity rule above ✅)

# Check why a Pod is stuck in Pending (Scheduler can't find suitable node)
kubectl describe pod fraud-model-7d8f9b-pending -n ml-production | grep -A5 "Events"

# Events:
#   Warning  FailedScheduling  "0/3 nodes are available:
#                               3 Insufficient memory"
#   → All Worker Nodes are full! Need to add more nodes or reduce resource requests.

🔄 Master Service 4: Controller Manager — The Self-Healing Watchdog

🏥 The Hospital Monitor Analogy

Imagine a hospital patient connected to monitors. Every monitor beeps when a reading goes outside the normal range: heart rate too low, blood pressure too high, oxygen dropping. Nurses and doctors respond immediately to bring everything back to normal.

The Controller Manager is the hospital monitoring system for your Kubernetes cluster. It continuously watches the actual state of every resource and compares it to the desired state stored in etcd. When they don't match — it acts. Immediately. Automatically.

🔬 The Controller Manager is Actually Many Controllers

The "Controller Manager" is a single process that runs many specialized mini-controllers inside it. Each one watches a different type of Kubernetes resource.

CONTROLLERS INSIDE THE CONTROLLER MANAGER:

Deployment Controller
  "Desired: 5 fraud-model pods. Actual: 3 running."
  → Creates 2 new Pods. Sends request to API Server → Scheduler picks nodes.

ReplicaSet Controller
  "A fraud-model Pod crashed on worker-2."
  → Immediately creates a replacement Pod. No human needed.

Node Controller
  "Worker-3 hasn't sent a heartbeat in 40 seconds."
  → Marks Worker-3 as 'NotReady'. After 5 min: evicts all Pods from that node.

Service Account Controller
  "New namespace ml-staging was created but has no default service account."
  → Creates default service account automatically.

Job Controller
  "ML training batch job has been running for 6 hours — should be done in 4."
  → Marks job as Failed. Optionally retries based on backoffLimit setting.

♾️ The Control Loop — How Controllers Work

Every controller runs an infinite loop called a reconciliation loop:

WHILE cluster is running:

  1. READ desired state from etcd
     "fraud-model deployment should have 5 Pods"

  2. READ actual state from cluster
     "Only 4 Pods are currently running"

  3. COMPARE: Is actual == desired?
     NO — there's a gap (5 desired, 4 actual)

  4. ACT to close the gap
     "Create 1 new Pod via API Server"

  5. WAIT (typically a few seconds)

  6. GO BACK TO STEP 1

This loop runs FOREVER. This is why Kubernetes is "self-healing". 🔄

💻 Watching the Controller Manager in Action

📌 Code Purpose — Observing Controller Manager Reconciliation

What this shows: How to observe the Controller Manager doing its job in real time. We deliberately break something (delete a Pod manually) and watch the Controller Manager notice and fix it automatically — demonstrating the reconciliation loop.
# ── DEMONSTRATE CONTROLLER MANAGER SELF-HEALING ──────────────────────

# First, check current state (5 Pods running as desired)
kubectl get pods -n ml-production
# fraud-model-pod-1   Running
# fraud-model-pod-2   Running
# fraud-model-pod-3   Running
# fraud-model-pod-4   Running
# fraud-model-pod-5   Running

# DELIBERATELY delete a Pod (simulating a crash!)
kubectl delete pod fraud-model-pod-3 -n ml-production
# pod "fraud-model-pod-3" deleted

# Watch what happens IMMEDIATELY (within 5 seconds!)
kubectl get pods -n ml-production -w

# NAME                READY   STATUS              REASON
# fraud-model-pod-1   1/1     Running
# fraud-model-pod-2   1/1     Running
# fraud-model-pod-3   0/1     Terminating         ← being deleted
# fraud-model-pod-4   1/1     Running
# fraud-model-pod-5   1/1     Running
# fraud-model-pod-6   0/1     ContainerCreating   ← NEW! Controller created it!
# fraud-model-pod-6   1/1     Running             ← RESTORED in ~20 seconds ✅

# The Controller Manager noticed: "Desired=5, Actual=4 → CREATE 1 MORE!"
# You never typed a single recovery command. It just happened.

# See Controller Manager decisions in the events log
kubectl get events -n ml-production | grep "fraud-model" | tail -5
# SuccessfulCreate  Created pod: fraud-model-pod-6  ← Controller Manager created this

💪 WORKER NODE SERVICES

Worker Nodes run exactly 2 core services (plus optional extras like kube-proxy). These 2 services receive instructions from the Master and execute them by actually starting, stopping, and monitoring your ML model containers.


⚙️ Worker Service 1: Container Runtime — The Engine Room

🚗 The Car Engine Analogy

A car has a dashboard full of controls — steering wheel, gear stick, pedals. But none of that moves the car. The engine does. You use the controls to tell the engine what to do, but the engine is the actual power source.

The Container Runtime is the engine of the Worker Node. The Kubelet (controls) tells it what to do. The Container Runtime (engine) actually does it: pulling images, creating containers, allocating memory, starting processes, and stopping them.

🔬 The 3 Most Common Container Runtimes

Runtime Full Name Used By Status
containerd ⭐ Container Daemon OCI OKE default, most K8s clusters Industry standard ✅
CRI-O Container Runtime Interface - OCI Red Hat OpenShift Lightweight alternative
Docker Engine Docker Standalone Docker, Docker Desktop Dev/testing (K8s deprecated it)
💡 Important Note: Kubernetes removed direct Docker Engine support (called "dockershim") in 2022. Modern clusters use containerd directly. But here's the good news: Docker still builds your images. containerd just runs them differently. Your docker build commands still work fine — the change is invisible to developers. 🔄

📦 What the Container Runtime Does Step by Step

WHEN KUBELET SAYS: "Start fraud-model container"

Container Runtime does this:

Step 1: Check if image exists locally
  "ap-mumbai-1.ocir.io/mytenancy/fraud-model:2.1.0 — not cached locally"

Step 2: Pull image from OCIR
  Authenticate with OCIR using imagePullSecrets
  Download each image layer (Python base + pip packages + your code)
  Cache layers for future use

Step 3: Create container namespace
  Create isolated process space (PID namespace)
  Create isolated network namespace (IP address)
  Create isolated filesystem (container root)

Step 4: Mount volumes
  Attach /app/models → OCI Block Volume (persistent model storage)

Step 5: Start the container process
  Run: /app/entrypoint.sh
  Container PID 1 starts. Python/FastAPI starts. Model loads.

Step 6: Report to Kubelet
  "Container is running, PID=1234, IP=10.244.1.15"

💻 Inspecting the Container Runtime

📌 Code Purpose — Inspect What Container Runtime is Running

What this shows: How to verify which container runtime is installed on your OCI OKE Worker Nodes, inspect running containers at the container runtime level (bypassing Kubernetes), and check image caching status. Useful when debugging containers that Kubernetes can't start.
# Check which container runtime is running on each Worker Node
kubectl get nodes -o wide

# NAME       STATUS   VERSION   CONTAINER-RUNTIME
# worker-1   Ready    v1.29.1   containerd://1.7.11    ← containerd on OKE
# worker-2   Ready    v1.29.1   containerd://1.7.11

# SSH into a Worker Node to inspect containerd directly
# (In OCI OKE, Worker Nodes are YOUR OCI Compute instances)
ssh -i ~/.ssh/oci_key opc@

# On the Worker Node: list ALL containers via containerd
# (crictl = CLI for container runtimes, works with containerd and CRI-O)
sudo crictl ps

# CONTAINER   IMAGE                                STATUS   NAME
# a9f3b2c1    fraud-model:2.1.0                   Running  fraud-api
# b8d4e7f2    pause:3.9                            Running  POD (sandbox)

# List images cached on this Worker Node
sudo crictl images

# IMAGE                                        TAG     SIZE
# ap-mumbai-1.ocir.io/.../fraud-model          2.1.0   612MB   ← cached ✅
# python                                       3.10    125MB   ← base layer cached

# See container logs at the runtime level (bypasses Kubernetes)
sudo crictl logs a9f3b2c1

# ── Check container runtime health from Kubernetes side ──────────────
kubectl describe node worker-1 | grep -A3 "Container Runtime"

# Container Runtime Version: containerd://1.7.11
# Operating System:          Oracle Linux Server 8.9
# Architecture:              amd64

👂 Worker Service 2: Kubelet — The Worker's Loyal Agent

📞 The Branch Manager Analogy

Imagine a bank with a head office (Master Node) and many branches (Worker Nodes). Each branch has a branch manager who:

  • Receives instructions from head office ("open 3 new accounts today")
  • Carries out those instructions locally ("I've opened them")
  • Sends regular status reports back ("all accounts active, no issues")
  • Alerts head office immediately if something goes wrong ("ATM is down!")

The Kubelet is exactly that branch manager for each Worker Node. It is the only Kubernetes component that runs on Worker Nodes. Everything else on the worker is managed by the Kubelet.

🔬 Everything Kubelet Does

  • Registers itself with the Master
    When a Worker Node starts up, its Kubelet introduces itself to the API Server: "Hi, I'm Worker-3. I have 8 CPUs and 32 GB RAM. I'm available for Pods."
  • Watches for new Pod assignments
    Kubelet watches the API Server continuously for any Pod that has been scheduled onto its node. The moment the Scheduler assigns a Pod to Worker-3, Kubelet on Worker-3 sees it and begins starting the container.
  • Instructs the Container Runtime
    Kubelet reads the Pod specification (image name, resources, env vars, volumes) and tells the Container Runtime: "Start this container with these settings."
  • Runs health checks
    After starting a container, Kubelet continuously runs the readinessProbe and livenessProbe defined in your YAML. "Is the ML API responding on /health? Is it still alive?"
  • Restarts failed containers
    If a container crashes and the Pod's restart policy allows it, Kubelet restarts it locally — without even asking the Master first. Fast local recovery.
  • Sends heartbeats to the Master
    Every 10 seconds (by default), Kubelet reports its health and resource usage back to the API Server → saved in etcd. If heartbeats stop, the Node Controller marks it as NotReady.

💻 Kubelet in Action — Health Probes for ML Models

📌 Code Purpose — Kubelet Health Probe Configuration

What this shows: How to configure the two health probes that Kubelet runs on your ML containers — readinessProbe (is it ready for traffic?) and livenessProbe (is it still alive?). These are the most critical settings for ML models in production because large models take time to load at startup and sometimes hang silently. Kubelet uses these probes to make automatic decisions.
# Kubelet health probe configuration for your ML model Pod
# These probes are run by the Kubelet on the Worker Node

containers:
  - name: fraud-api
    image: ap-mumbai-1.ocir.io/mytenancy/fraud-model:2.1.0
    ports:
      - containerPort: 8083

    # ── READINESS PROBE ───────────────────────────────────────────────
    # Kubelet asks: "Is this container READY to serve traffic?"
    # Before readiness passes: Pod gets NO requests from load balancer
    # After readiness passes: Pod starts receiving traffic
    readinessProbe:
      httpGet:
        path: /health        # Kubelet calls GET http://pod-ip:8083/health
        port: 8083
      initialDelaySeconds: 20  # Wait 20s before first check (model loads in ~15s)
      periodSeconds: 10        # Check every 10 seconds
      successThreshold: 1      # 1 successful response = READY
      failureThreshold: 3      # 3 consecutive failures = NOT READY (removed from LB)

    # ── LIVENESS PROBE ────────────────────────────────────────────────
    # Kubelet asks: "Is this container still ALIVE and healthy?"
    # If liveness fails: Kubelet KILLS and RESTARTS the container
    # Use case: catches frozen ML models that stopped responding but didn't crash
    livenessProbe:
      httpGet:
        path: /health
        port: 8083
      initialDelaySeconds: 45  # Wait longer — don't kill during model loading!
      periodSeconds: 30        # Check every 30 seconds
      failureThreshold: 3      # 3 failures in a row = KILL AND RESTART

    # ── STARTUP PROBE (for slow-starting ML models) ───────────────────
    # New in K8s 1.18 — protects slow-starting containers
    # Disables readiness and liveness until startup is complete
    # Perfect for: models that take 60+ seconds to load (large LLMs!)
    startupProbe:
      httpGet:
        path: /health
        port: 8083
      initialDelaySeconds: 10
      periodSeconds: 10
      failureThreshold: 30     # Give 30 * 10s = 5 minutes to start up
      # Once startupProbe passes, liveness and readiness take over
# ── CHECK KUBELET STATUS ON A WORKER NODE ────────────────────────────
# SSH into an OCI Worker Node
ssh -i ~/.ssh/oci_key opc@

# Check if Kubelet is running
systemctl status kubelet

# ● kubelet.service - Kubernetes Kubelet
#    Loaded: loaded (/etc/systemd/system/kubelet.service)
#    Active: active (running) since Mon 2025-03-10 08:23:11 UTC  ← running ✅
#    Main PID: 1234 (kubelet)

# View Kubelet logs (to debug container startup issues)
journalctl -u kubelet -f --since "10 minutes ago"

# Mar 10 08:45:01 worker-1 kubelet[1234]: Successfully pulled image from OCIR
# Mar 10 08:45:03 worker-1 kubelet[1234]: Created container fraud-api
# Mar 10 08:45:03 worker-1 kubelet[1234]: Started container fraud-api
# Mar 10 08:45:25 worker-1 kubelet[1234]: Container is ready ← readinessProbe passed!

# ── CHECK HEALTH PROBE STATUS FROM KUBERNETES SIDE ───────────────────
# Back on your laptop (not the Worker Node):
kubectl describe pod fraud-model-pod-1 -n ml-production

# Liveness:   http-get http://:8083/health delay=45s timeout=1s period=30s #success=1 #failure=3
# Readiness:  http-get http://:8083/health delay=20s timeout=1s period=10s #success=1 #failure=3
# Conditions:
#   Ready         True   ← readinessProbe passing — Pod is in the load balancer ✅
# Events:
#   Normal   Pulled     Successfully pulled image "fraud-model:2.1.0" in 8.3s
#   Normal   Started    Started container fraud-api
❌ The most common ML model mistake with Kubelet probes:

Setting initialDelaySeconds too short on the livenessProbe. Your fraud model takes 25 seconds to load. If initialDelaySeconds=10, Kubelet checks health at 10 seconds, the model hasn't loaded yet, health returns 503, Kubelet counts as a failure, kills the container, starts again — and gets stuck in an infinite restart loop (CrashLoopBackOff)! 🔁

✅ Always set initialDelaySeconds on livenessProbe to at least 2x your model's loading time. Use startupProbe for models that take over 60 seconds to load.

🔄 Section 3: All 6 Services Working Together — Full ML Deployment Walkthrough

Let's trace exactly what happens when you run kubectl apply -f fraud-model-deployment.yaml — following the request through every single service.

YOU TYPE: kubectl apply -f fraud-model-deployment.yaml

━━━ MASTER NODE SERVICES ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

STEP 1 — API Server receives the request
  • Authenticates: "Is this a valid OCI/K8s credential?" ✅
  • Authorizes: "Is this user allowed to create Deployments?" ✅
  • Validates: "Is the YAML well-formed? Does the image exist?" ✅
  • Saves to etcd: "Desired state: 5 fraud-model Pods."

STEP 2 — etcd stores the desired state
  • Writes to disk: Deployment spec with replicas=5
  • Available to all other components via API Server

STEP 3 — Controller Manager detects work to do
  • Deployment Controller: "Desired=5 Pods, Actual=0 Pods"
  • Creates 5 Pod objects in etcd (status: "Pending")

STEP 4 — Scheduler assigns Pods to Nodes
  • Sees 5 unscheduled Pods in etcd
  • Filters nodes: all 3 Worker Nodes have capacity
  • Scores: Worker-3 has most free RAM → gets 2 Pods
  • Updates etcd: Pod-1→Worker-1, Pod-2→Worker-2, Pod-3→Worker-3,
                   Pod-4→Worker-1, Pod-5→Worker-3

━━━ WORKER NODE SERVICES ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

STEP 5 — Kubelet on each Worker detects its assigned Pods
  • Worker-1 Kubelet sees: "I have 2 new Pods to start"
  • Reads Pod spec: image, resources, env vars, probes, volumes

STEP 6 — Container Runtime pulls and starts containers
  • containerd checks: "Is fraud-model:2.1.0 cached?" → No
  • Pulls from OCIR (authenticated via imagePullSecret)
  • Creates container with isolated filesystem + network
  • Starts /app/entrypoint.sh → FastAPI + ML model loading

STEP 7 — Kubelet runs health checks
  • After 20s: calls GET /health → 200 OK → readinessProbe PASSES
  • Reports to API Server: "Pod-1 is Ready on Worker-1" ✅

STEP 8 — etcd updated with actual state
  • etcd now shows: Desired=5, Actual=5, all Running ✅
  • Controller Manager reconciliation loop confirms: "gap = zero, nothing to do"

STEP 9 — Traffic flows to your ML model 🎉
  • OCI Load Balancer → Worker Nodes → Pods → ML predictions!
  • Total time from kubectl apply to first request served: ~45 seconds

💥 Section 4: Service-by-Service Failure Scenarios

Service Fails Immediate Effect What Still Works Recovery
API Server No new deployments, no scaling, no kubectl commands Existing containers keep running normally OKE auto-restarts it; HA has 3 replicas
etcd Cluster "freezes" — no state changes possible Existing containers keep running OKE runs 3-node etcd cluster; OCI backs up daily
Scheduler New Pods stay in "Pending" — not assigned to nodes Already-running Pods continue unaffected Kubernetes restarts it automatically
Controller Manager No self-healing — crashes not auto-recovered Running containers keep serving traffic Kubernetes restarts it automatically
Container Runtime No new containers can start on that Worker Node Already-running containers may continue briefly systemd restarts containerd; or drain the node
Kubelet Node marked NotReady after 40s; Pods evicted after 5min Containers already running stay up temporarily systemd restarts Kubelet; Controller reschedules Pods

✅ Best Practices & Common Mistakes

✅ DOs:
  • ✅ Always set readinessProbe — prevents traffic to Pods still loading your ML model
  • ✅ Set initialDelaySeconds on livenessProbe longer than your model loading time
  • ✅ Use startupProbe for large ML models (LLMs, heavy CNNs) that take 60+ seconds to load
  • ✅ Always define resource requests and limits — the Scheduler needs them to make smart decisions
  • ✅ Use podAntiAffinity — force the Scheduler to spread your ML replicas across different Worker Nodes
  • ✅ On OCI: use OKE — Oracle manages the 4 Master services so you focus entirely on your ML deployment
  • ✅ Monitor etcd backup status on OKE — it's your cluster's insurance policy
❌ DON'Ts:
  • ❌ Don't skip health probes — Kubelet will send traffic to containers still loading the model
  • ❌ Don't set livenessProbe too aggressively — CrashLoopBackOff will restart your ML model repeatedly
  • ❌ Don't ignore "Pending" Pod status — it means the Scheduler can't find a suitable node (resource shortage!)
  • ❌ Don't set resource limits without requests — the Scheduler can't plan correctly
  • ❌ Don't run ML workloads directly on Master Nodes — they're reserved for control plane services only
  • ❌ Don't ignore etcd warnings in OKE logs — a corrupted etcd = total cluster state loss

🏆 High-Level Summary: All 6 Services Mastered!

🧠 MASTER NODE — 4 Services (The Brain):

  • 📡 API Server — The front door. Every request enters here. Authenticates, validates, routes. Nothing bypasses it.
  • 📚 etcd — The cluster's memory. Stores desired + actual state. The source of truth for everything. Distributed, durable, backed up.
  • 📋 Scheduler — The smart matchmaker. Decides which Worker Node runs which Pod. Filters unsuitable nodes, scores the rest, picks the winner.
  • 🔄 Controller Manager — The self-healing watchdog. Watches for desired vs actual state gaps. Creates, deletes, and restarts resources automatically. Infinite reconciliation loop.

💪 WORKER NODE — 2 Services (The Muscle):

  • ⚙️ Container Runtime (containerd) — The engine. Actually pulls images from OCIR, creates containers, allocates CPU/RAM, starts processes. The Kubelet's hands.
  • 👂 Kubelet — The loyal agent. Registers with Master, watches for Pod assignments, instructs Container Runtime, runs health probes, reports heartbeats. The Worker Node's only voice to the Master.
ChatGPT, Google's recommendation engine, Netflix streaming predictions, every fraud detection system, every real-time ML API in production — they all run on clusters powered by exactly these 6 services.

You've gone from "what is a Master Node?" to understanding every component, what it does, why it exists, and how it fails.

Keep learning, keep deploying 🐼✨

Comments