Kubernetes Node Services Explained: API Server, etcd, Scheduler, Controller & Kubelet
Let's imagine you manage a giant school. The school has a Principal's Office and many Classrooms.
The Principal's Office keeps records, makes decisions, assigns teachers to rooms, and monitors whether every class is running correctly. The Classrooms are where the actual teaching happens.
Now swap "teaching" for "running your ML model container" — and you have exactly how Kubernetes works.
🏫 Principal's Office = Master Node
— API Server = the reception desk
— etcd = the filing cabinet
— Scheduler = the timetable planner
— Controller = the school inspector
🏫 Classrooms = Worker Nodes
— Container Runtime = the whiteboard & chalk
— Kubelet = the classroom teacher
Together they run your ML models reliably, at scale, on Oracle OCI — 24 hours a day, 7 days a week. 🚀
📚 What We'll Cover
- 🔹 The Big Picture — How All 6 Services Fit Together
- 🔹 Master Node Service 1: API Server — The Front Door
- 🔹 Master Node Service 2: etcd — The Brain's Memory
- 🔹 Master Node Service 3: Scheduler — The Smart Planner
- 🔹 Master Node Service 4: Controller Manager — The Self-Healing Watchdog
- 🔹 Worker Node Service 1: Container Runtime — The Engine Room
- 🔹 Worker Node Service 2: Kubelet — The Loyal Agent
- 🔹 How All 6 Services Work Together — Full ML Deployment Walkthrough
- 🔹 What Happens During a Failure — Service-by-Service Recovery
- 🔹 Practical Commands to Inspect Each Service
- 🔹 Best Practices & Common Mistakes
- 🔹 High-Level Summary
🗺️ Section 1: The Big Picture — All 6 Services in One View
Before diving into each service, let's see them all together. This diagram is the mental model you'll use for everything else in this post.
│ MASTER NODE (Control Plane) │
│ │
│ ┌────────────────┐ ┌────────────────┐ │
│ │ API Server │ │ etcd Store │ │
│ │ 📡 The Gateway │ │ 📚 The Memory │ │
│ └────────────────┘ └────────────────┘ │
│ │
│ ┌────────────────┐ ┌────────────────┐ │
│ │ Scheduler │ │ Controller Mgr │ │
│ │ 📋 The Planner │ │ 🔄 The Fixer │ │
│ └────────────────┘ └────────────────┘ │
└────────────────────────────┬─────────────────────────────────────┘
│ commands via API Server
┌───────────┴───────────┐
▼ ▼
┌──────────────────────────┐ ┌──────────────────────────┐
│ WORKER NODE 1 │ │ WORKER NODE 2 │
│ ┌──────────────────────┐ │ │ ┌──────────────────────┐ │
│ │ Container Runtime │ │ │ │ Container Runtime │ │
│ │ ⚙️ Actually runs │ │ │ │ ⚙️ Actually runs │ │
│ │ containers │ │ │ │ containers │ │
│ └──────────────────────┘ │ │ └──────────────────────┘ │
│ ┌──────────────────────┐ │ │ ┌──────────────────────┐ │
│ │ Kubelet │ │ │ │ Kubelet │ │
│ │ 👂 Reports to Master │ │ │ │ 👂 Reports to Master │ │
│ └──────────────────────┘ │ │ └──────────────────────┘ │
└──────────────────────────┘ └──────────────────────────┘
- 🟣 API Server — Every request comes through here first. Nothing bypasses it.
- 🟣 etcd — Remembers the state of everything. The cluster's notebook.
- 🟣 Scheduler — Decides which Worker Node gets which container. The matchmaker.
- 🟣 Controller Manager — Watches everything. Fixes problems automatically.
- 🟠 Container Runtime — Actually starts and stops containers on the Worker.
- 🟠 Kubelet — The Worker's agent. Listens to Master, reports back health.
🧠 MASTER NODE SERVICES
The Master Node runs 4 core services. None of them run your ML model containers directly. Their only job is to manage the entire cluster.
📡 Master Service 1: API Server — The Front Door of Everything
🏨 The Hotel Reception Analogy
Imagine a grand hotel. You want a room, room service, housekeeping, maintenance — everything goes through the front desk receptionist. You never walk directly into the kitchen to order food. You never go to the maintenance room yourself. Everything is coordinated through one central point.
The API Server is that receptionist. Every single interaction with the Kubernetes cluster — deploying an ML model, scaling up containers, checking cluster health — must go through the API Server first. Nothing talks to anything else directly.
kubectl (your laptop) ──────────────────▶ API Server
OCI OKE Dashboard ──────────────────────▶ API Server
CI/CD Pipeline (OCI DevOps) ────────────▶ API Server
Scheduler ──────────────────────────────▶ API Server
Controller Manager ─────────────────────▶ API Server
Kubelet on Worker Nodes ────────────────▶ API Server
NO direct connections between components.
Everything passes through the API Server. Always. 📡
🔒 What the API Server Actually Does
- Authentication — "Who are you?" Checks your OCI IAM credentials or Kubernetes service account token. Rejects unknown callers immediately.
- Authorization — "Are you allowed to do this?" Checks RBAC (Role-Based Access Control) rules. A data scientist can deploy models but cannot delete Worker Nodes.
- Validation — "Is your request well-formed?" Checks that your YAML is correctly structured before accepting it. Rejects requests that reference non-existent images or invalid CPU values.
- Routing — "Who needs to act on this?" Forwards validated requests to the right component: scheduling requests to the Scheduler, storage to etcd, worker commands to Kubelet.
If every component talked to every other component directly, you'd have a chaotic web of connections impossible to secure or debug. With one gateway: one place to add security, one place to log requests, one place to rate-limit, one place to audit. Clean. Scalable. Secure. 🔐
💻 API Server in Action
What this shows: Every
kubectl command is actually an
HTTP request to the API Server. These examples show what's happening
under the hood and how to interact with the API Server directly,
which is useful for debugging and automation in OCI OKE deployments.
# Every kubectl command hits the API Server first
# kubectl is just a user-friendly wrapper around API Server HTTP calls
# Check API Server is reachable and healthy
kubectl cluster-info
# Kubernetes control plane is running at https://xxx.xxx.oraclecloud.com:6443
# CoreDNS is running at https://xxx.xxx.oraclecloud.com:6443/api/v1/...
# See your current connection context (which cluster's API Server you're talking to)
kubectl config current-context
# List all available API endpoints the API Server exposes
kubectl api-resources
# NAME SHORTNAMES APIVERSION NAMESPACED KIND
# pods po v1 true Pod
# deployments deploy apps/v1 true Deployment
# services svc v1 true Service
# Call API Server DIRECTLY (no kubectl wrapper) to see raw HTTP
kubectl get pods -n ml-production -v=8
# -v=8 shows full HTTP request/response between kubectl and API Server
# You'll see: GET https://api-server-url/api/v1/namespaces/ml-production/pods
# Response: 200 OK + JSON list of all pods
# Check API Server component health
kubectl get componentstatus
# NAME STATUS MESSAGE
# scheduler Healthy ok
# controller-manager Healthy ok
# etcd-0 Healthy {"health":"true"}
📚 Master Service 2: etcd — The Cluster's Memory
🗄️ The Filing Cabinet Analogy
Every school has a central records office — a place that stores every student's file: which class they're in, what grades they have, who their teacher is, their attendance records.
If the records office burns down and everything is lost, the school is in chaos. Nobody knows who belongs where. etcd is that records office for Kubernetes. It stores the entire state of the cluster — what should be running, what is running, how many copies of each thing, all configurations. Lose etcd and the cluster has amnesia. 🧠💾
🔬 What etcd Actually Stores
/registry/pods/ml-production/fraud-model-pod-1
→ {name, image, resources, status, node: "worker-1"}
/registry/pods/ml-production/fraud-model-pod-2
→ {name, image, resources, status, node: "worker-2"}
/registry/deployments/ml-production/fraud-model
→ {replicas: 5, image: "ocir.../fraud-model:2.1", strategy: RollingUpdate}
/registry/nodes/worker-1
→ {status: Ready, CPU: 4, RAM: 16GB, pods: 3}
/registry/nodes/worker-2
→ {status: NotReady, last-heartbeat: 5 mins ago}
/registry/services/ml-production/fraud-api-service
→ {type: LoadBalancer, port: 80, targetPort: 8083}
⚡ Desired State vs Actual State — The Core Concept
etcd stores two critical things for every resource:
-
Desired State — What you want to be running.
"I want 5 copies of my fraud detection model running."
This is what you wrote in your YAML and applied with
kubectl apply. - Actual State — What is currently running. "Right now, only 4 copies are running because Pod 3 crashed." This is reported by Kubelets on Worker Nodes.
The gap between desired and actual state is what triggers action. The Controller Manager watches for this gap and closes it automatically. etcd is the source of truth for both values.
OCI OKE backs up etcd automatically. You can restore an entire cluster from an etcd backup if disaster strikes.
💻 Exploring etcd (What's Inside)
What this shows: You normally never interact with etcd directly — you use kubectl which reads/writes etcd through the API Server. These commands show the etcd-equivalent state through kubectl, helping you understand what's stored in etcd at any given moment.
# etcd stores "desired state" — you set this with kubectl apply
# kubectl get reads the current desired + actual state FROM etcd via API Server
# See the DESIRED state of your ML deployment (stored in etcd)
kubectl get deployment fraud-model-deployment -n ml-production -o yaml
# apiVersion: apps/v1
# kind: Deployment
# spec:
# replicas: 5 ← DESIRED: 5 Pods
# template:
# spec:
# containers:
# image: ap-mumbai-1.ocir.io/mytenancy/fraud-model:2.1.0
# See the ACTUAL state (what's really running vs desired)
kubectl get pods -n ml-production
# NAME READY STATUS RESTARTS
# fraud-model-7d8f9b-x2k4p 1/1 Running 0 ← Pod 1: OK
# fraud-model-7d8f9b-m9p3q 1/1 Running 0 ← Pod 2: OK
# fraud-model-7d8f9b-n7r8s 0/1 Pending 0 ← Pod 3: PENDING!
# fraud-model-7d8f9b-p4q5r 1/1 Running 0 ← Pod 4: OK
# (Only 4 running, 1 pending — desired is 5, actual is 4 — etcd sees this gap!)
# Watch etcd-tracked events in real time
kubectl get events -n ml-production --sort-by='.lastTimestamp'
# REASON MESSAGE
# Scheduled Successfully assigned fraud-model-pod to worker-1
# Pulled Container image pulled from OCIR
# Started Container started successfully
# BackOff Container keeps crashing — back-off restarting
📋 Master Service 3: Scheduler — The Smart Placement Engine
🏫 The School Timetable Analogy
At school, there's a timetable planner who must assign every class to a room for every time slot. They can't put 50 students in a room built for 20. They can't schedule the chemistry lab and cooking class in the same room. They need to balance load across all available rooms.
The Scheduler is that timetable planner for Kubernetes. When a new ML model Pod needs to run, the Scheduler asks: "Which Worker Node is the best fit for this Pod right now?" Then it makes a decision and records it in etcd via the API Server.
🔍 How the Scheduler Makes Its Decision — Step by Step
STEP 1 — FILTERING (eliminate unsuitable nodes)
Worker 1: 4 CPUs, 16 GB RAM, currently using 3.5 CPU → only 0.5 free → ❌ SKIP
Worker 2: 4 CPUs, 16 GB RAM, currently using 2 CPU → 2 free → ✅ KEEP
Worker 3: 8 CPUs, 32 GB RAM, currently using 1 CPU → 7 free → ✅ KEEP
Worker 4: NotReady status → ❌ SKIP
STEP 2 — SCORING (rank the suitable nodes)
Worker 2: score 62/100 (has enough, not much spare)
Worker 3: score 88/100 (has plenty of spare capacity)
STEP 3 — BINDING (assign the Pod)
Winner: Worker 3 🏆
Scheduler writes to etcd: "fraud-model-pod-5 → Worker 3"
Kubelet on Worker 3 picks this up and starts the container.
🎯 Scheduler Scoring Factors
- Available CPU & RAM — prefer nodes with more free resources
- Node affinity rules — "only place on GPU nodes" (critical for deep learning models that need OCI GPU shapes)
- Pod anti-affinity — "spread replicas across different nodes" (never put all fraud model copies on the same machine)
- Taints and tolerations — special labels that say "this node is reserved for ML workloads only — no other Pods allowed"
- Data locality — prefer the node closest to the data (e.g., same OCI Availability Domain as the object storage)
💻 Influencing the Scheduler's Decisions
What this shows: How to give the Scheduler hints and rules about where to place your ML model Pods. These YAML snippets tell the Scheduler "put this ML Pod only on GPU nodes" and "never put two copies of this Pod on the same physical machine" — ensuring both performance and fault tolerance for ML deployments on OCI OKE.
# fraud-model-deployment.yaml — Scheduler directives
spec:
template:
spec:
# RULE 1: Only schedule on GPU Worker Nodes (for deep learning models)
nodeSelector:
accelerator: "nvidia-a10" # OCI VM.GPU.A10.1 nodes have this label
# RULE 2: Spread replicas across different Worker Nodes (fault tolerance)
# Never put two fraud-model Pods on the same node
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchLabels:
app: fraud-model # match our own pods
topologyKey: kubernetes.io/hostname
# topologyKey = "different hostnames" = different physical nodes
# RULE 3: Resource requirements (Scheduler filters nodes without this)
containers:
- name: fraud-api
resources:
requests:
memory: "2Gi" # Need minimum 2GB RAM
cpu: "1000m" # Need minimum 1 full CPU core
limits:
memory: "4Gi" # Max 4GB RAM
cpu: "2000m" # Max 2 CPU cores
# See where the Scheduler placed each Pod (which Worker Node)
kubectl get pods -n ml-production -o wide
# NAME NODE STATUS
# fraud-model-7d8f9b-x2k4p worker-1 Running ← assigned to worker-1
# fraud-model-7d8f9b-m9p3q worker-2 Running ← assigned to worker-2
# fraud-model-7d8f9b-n7r8s worker-3 Running ← assigned to worker-3
# (all on DIFFERENT nodes due to podAntiAffinity rule above ✅)
# Check why a Pod is stuck in Pending (Scheduler can't find suitable node)
kubectl describe pod fraud-model-7d8f9b-pending -n ml-production | grep -A5 "Events"
# Events:
# Warning FailedScheduling "0/3 nodes are available:
# 3 Insufficient memory"
# → All Worker Nodes are full! Need to add more nodes or reduce resource requests.
🔄 Master Service 4: Controller Manager — The Self-Healing Watchdog
🏥 The Hospital Monitor Analogy
Imagine a hospital patient connected to monitors. Every monitor beeps when a reading goes outside the normal range: heart rate too low, blood pressure too high, oxygen dropping. Nurses and doctors respond immediately to bring everything back to normal.
The Controller Manager is the hospital monitoring system for your Kubernetes cluster. It continuously watches the actual state of every resource and compares it to the desired state stored in etcd. When they don't match — it acts. Immediately. Automatically.
🔬 The Controller Manager is Actually Many Controllers
The "Controller Manager" is a single process that runs many specialized mini-controllers inside it. Each one watches a different type of Kubernetes resource.
Deployment Controller
"Desired: 5 fraud-model pods. Actual: 3 running."
→ Creates 2 new Pods. Sends request to API Server → Scheduler picks nodes.
ReplicaSet Controller
"A fraud-model Pod crashed on worker-2."
→ Immediately creates a replacement Pod. No human needed.
Node Controller
"Worker-3 hasn't sent a heartbeat in 40 seconds."
→ Marks Worker-3 as 'NotReady'. After 5 min: evicts all Pods from that node.
Service Account Controller
"New namespace ml-staging was created but has no default service account."
→ Creates default service account automatically.
Job Controller
"ML training batch job has been running for 6 hours — should be done in 4."
→ Marks job as Failed. Optionally retries based on backoffLimit setting.
♾️ The Control Loop — How Controllers Work
Every controller runs an infinite loop called a reconciliation loop:
1. READ desired state from etcd
"fraud-model deployment should have 5 Pods"
2. READ actual state from cluster
"Only 4 Pods are currently running"
3. COMPARE: Is actual == desired?
NO — there's a gap (5 desired, 4 actual)
4. ACT to close the gap
"Create 1 new Pod via API Server"
5. WAIT (typically a few seconds)
6. GO BACK TO STEP 1
This loop runs FOREVER. This is why Kubernetes is "self-healing". 🔄
💻 Watching the Controller Manager in Action
What this shows: How to observe the Controller Manager doing its job in real time. We deliberately break something (delete a Pod manually) and watch the Controller Manager notice and fix it automatically — demonstrating the reconciliation loop.
# ── DEMONSTRATE CONTROLLER MANAGER SELF-HEALING ──────────────────────
# First, check current state (5 Pods running as desired)
kubectl get pods -n ml-production
# fraud-model-pod-1 Running
# fraud-model-pod-2 Running
# fraud-model-pod-3 Running
# fraud-model-pod-4 Running
# fraud-model-pod-5 Running
# DELIBERATELY delete a Pod (simulating a crash!)
kubectl delete pod fraud-model-pod-3 -n ml-production
# pod "fraud-model-pod-3" deleted
# Watch what happens IMMEDIATELY (within 5 seconds!)
kubectl get pods -n ml-production -w
# NAME READY STATUS REASON
# fraud-model-pod-1 1/1 Running
# fraud-model-pod-2 1/1 Running
# fraud-model-pod-3 0/1 Terminating ← being deleted
# fraud-model-pod-4 1/1 Running
# fraud-model-pod-5 1/1 Running
# fraud-model-pod-6 0/1 ContainerCreating ← NEW! Controller created it!
# fraud-model-pod-6 1/1 Running ← RESTORED in ~20 seconds ✅
# The Controller Manager noticed: "Desired=5, Actual=4 → CREATE 1 MORE!"
# You never typed a single recovery command. It just happened.
# See Controller Manager decisions in the events log
kubectl get events -n ml-production | grep "fraud-model" | tail -5
# SuccessfulCreate Created pod: fraud-model-pod-6 ← Controller Manager created this
💪 WORKER NODE SERVICES
Worker Nodes run exactly 2 core services (plus optional extras like kube-proxy). These 2 services receive instructions from the Master and execute them by actually starting, stopping, and monitoring your ML model containers.
⚙️ Worker Service 1: Container Runtime — The Engine Room
🚗 The Car Engine Analogy
A car has a dashboard full of controls — steering wheel, gear stick, pedals. But none of that moves the car. The engine does. You use the controls to tell the engine what to do, but the engine is the actual power source.
The Container Runtime is the engine of the Worker Node. The Kubelet (controls) tells it what to do. The Container Runtime (engine) actually does it: pulling images, creating containers, allocating memory, starting processes, and stopping them.
🔬 The 3 Most Common Container Runtimes
| Runtime | Full Name | Used By | Status |
|---|---|---|---|
| containerd ⭐ | Container Daemon | OCI OKE default, most K8s clusters | Industry standard ✅ |
| CRI-O | Container Runtime Interface - OCI | Red Hat OpenShift | Lightweight alternative |
| Docker Engine | Docker | Standalone Docker, Docker Desktop | Dev/testing (K8s deprecated it) |
docker build commands still work fine —
the change is invisible to developers. 🔄
📦 What the Container Runtime Does Step by Step
Container Runtime does this:
Step 1: Check if image exists locally
"ap-mumbai-1.ocir.io/mytenancy/fraud-model:2.1.0 — not cached locally"
Step 2: Pull image from OCIR
Authenticate with OCIR using imagePullSecrets
Download each image layer (Python base + pip packages + your code)
Cache layers for future use
Step 3: Create container namespace
Create isolated process space (PID namespace)
Create isolated network namespace (IP address)
Create isolated filesystem (container root)
Step 4: Mount volumes
Attach /app/models → OCI Block Volume (persistent model storage)
Step 5: Start the container process
Run: /app/entrypoint.sh
Container PID 1 starts. Python/FastAPI starts. Model loads.
Step 6: Report to Kubelet
"Container is running, PID=1234, IP=10.244.1.15"
💻 Inspecting the Container Runtime
What this shows: How to verify which container runtime is installed on your OCI OKE Worker Nodes, inspect running containers at the container runtime level (bypassing Kubernetes), and check image caching status. Useful when debugging containers that Kubernetes can't start.
# Check which container runtime is running on each Worker Node
kubectl get nodes -o wide
# NAME STATUS VERSION CONTAINER-RUNTIME
# worker-1 Ready v1.29.1 containerd://1.7.11 ← containerd on OKE
# worker-2 Ready v1.29.1 containerd://1.7.11
# SSH into a Worker Node to inspect containerd directly
# (In OCI OKE, Worker Nodes are YOUR OCI Compute instances)
ssh -i ~/.ssh/oci_key opc@
# On the Worker Node: list ALL containers via containerd
# (crictl = CLI for container runtimes, works with containerd and CRI-O)
sudo crictl ps
# CONTAINER IMAGE STATUS NAME
# a9f3b2c1 fraud-model:2.1.0 Running fraud-api
# b8d4e7f2 pause:3.9 Running POD (sandbox)
# List images cached on this Worker Node
sudo crictl images
# IMAGE TAG SIZE
# ap-mumbai-1.ocir.io/.../fraud-model 2.1.0 612MB ← cached ✅
# python 3.10 125MB ← base layer cached
# See container logs at the runtime level (bypasses Kubernetes)
sudo crictl logs a9f3b2c1
# ── Check container runtime health from Kubernetes side ──────────────
kubectl describe node worker-1 | grep -A3 "Container Runtime"
# Container Runtime Version: containerd://1.7.11
# Operating System: Oracle Linux Server 8.9
# Architecture: amd64
👂 Worker Service 2: Kubelet — The Worker's Loyal Agent
📞 The Branch Manager Analogy
Imagine a bank with a head office (Master Node) and many branches (Worker Nodes). Each branch has a branch manager who:
- Receives instructions from head office ("open 3 new accounts today")
- Carries out those instructions locally ("I've opened them")
- Sends regular status reports back ("all accounts active, no issues")
- Alerts head office immediately if something goes wrong ("ATM is down!")
The Kubelet is exactly that branch manager for each Worker Node. It is the only Kubernetes component that runs on Worker Nodes. Everything else on the worker is managed by the Kubelet.
🔬 Everything Kubelet Does
-
Registers itself with the Master
When a Worker Node starts up, its Kubelet introduces itself to the API Server: "Hi, I'm Worker-3. I have 8 CPUs and 32 GB RAM. I'm available for Pods." -
Watches for new Pod assignments
Kubelet watches the API Server continuously for any Pod that has been scheduled onto its node. The moment the Scheduler assigns a Pod to Worker-3, Kubelet on Worker-3 sees it and begins starting the container. -
Instructs the Container Runtime
Kubelet reads the Pod specification (image name, resources, env vars, volumes) and tells the Container Runtime: "Start this container with these settings." -
Runs health checks
After starting a container, Kubelet continuously runs the readinessProbe and livenessProbe defined in your YAML. "Is the ML API responding on/health? Is it still alive?" -
Restarts failed containers
If a container crashes and the Pod's restart policy allows it, Kubelet restarts it locally — without even asking the Master first. Fast local recovery. -
Sends heartbeats to the Master
Every 10 seconds (by default), Kubelet reports its health and resource usage back to the API Server → saved in etcd. If heartbeats stop, the Node Controller marks it as NotReady.
💻 Kubelet in Action — Health Probes for ML Models
What this shows: How to configure the two health probes that Kubelet runs on your ML containers — readinessProbe (is it ready for traffic?) and livenessProbe (is it still alive?). These are the most critical settings for ML models in production because large models take time to load at startup and sometimes hang silently. Kubelet uses these probes to make automatic decisions.
# Kubelet health probe configuration for your ML model Pod
# These probes are run by the Kubelet on the Worker Node
containers:
- name: fraud-api
image: ap-mumbai-1.ocir.io/mytenancy/fraud-model:2.1.0
ports:
- containerPort: 8083
# ── READINESS PROBE ───────────────────────────────────────────────
# Kubelet asks: "Is this container READY to serve traffic?"
# Before readiness passes: Pod gets NO requests from load balancer
# After readiness passes: Pod starts receiving traffic
readinessProbe:
httpGet:
path: /health # Kubelet calls GET http://pod-ip:8083/health
port: 8083
initialDelaySeconds: 20 # Wait 20s before first check (model loads in ~15s)
periodSeconds: 10 # Check every 10 seconds
successThreshold: 1 # 1 successful response = READY
failureThreshold: 3 # 3 consecutive failures = NOT READY (removed from LB)
# ── LIVENESS PROBE ────────────────────────────────────────────────
# Kubelet asks: "Is this container still ALIVE and healthy?"
# If liveness fails: Kubelet KILLS and RESTARTS the container
# Use case: catches frozen ML models that stopped responding but didn't crash
livenessProbe:
httpGet:
path: /health
port: 8083
initialDelaySeconds: 45 # Wait longer — don't kill during model loading!
periodSeconds: 30 # Check every 30 seconds
failureThreshold: 3 # 3 failures in a row = KILL AND RESTART
# ── STARTUP PROBE (for slow-starting ML models) ───────────────────
# New in K8s 1.18 — protects slow-starting containers
# Disables readiness and liveness until startup is complete
# Perfect for: models that take 60+ seconds to load (large LLMs!)
startupProbe:
httpGet:
path: /health
port: 8083
initialDelaySeconds: 10
periodSeconds: 10
failureThreshold: 30 # Give 30 * 10s = 5 minutes to start up
# Once startupProbe passes, liveness and readiness take over
# ── CHECK KUBELET STATUS ON A WORKER NODE ────────────────────────────
# SSH into an OCI Worker Node
ssh -i ~/.ssh/oci_key opc@
# Check if Kubelet is running
systemctl status kubelet
# ● kubelet.service - Kubernetes Kubelet
# Loaded: loaded (/etc/systemd/system/kubelet.service)
# Active: active (running) since Mon 2025-03-10 08:23:11 UTC ← running ✅
# Main PID: 1234 (kubelet)
# View Kubelet logs (to debug container startup issues)
journalctl -u kubelet -f --since "10 minutes ago"
# Mar 10 08:45:01 worker-1 kubelet[1234]: Successfully pulled image from OCIR
# Mar 10 08:45:03 worker-1 kubelet[1234]: Created container fraud-api
# Mar 10 08:45:03 worker-1 kubelet[1234]: Started container fraud-api
# Mar 10 08:45:25 worker-1 kubelet[1234]: Container is ready ← readinessProbe passed!
# ── CHECK HEALTH PROBE STATUS FROM KUBERNETES SIDE ───────────────────
# Back on your laptop (not the Worker Node):
kubectl describe pod fraud-model-pod-1 -n ml-production
# Liveness: http-get http://:8083/health delay=45s timeout=1s period=30s #success=1 #failure=3
# Readiness: http-get http://:8083/health delay=20s timeout=1s period=10s #success=1 #failure=3
# Conditions:
# Ready True ← readinessProbe passing — Pod is in the load balancer ✅
# Events:
# Normal Pulled Successfully pulled image "fraud-model:2.1.0" in 8.3s
# Normal Started Started container fraud-api
Setting
initialDelaySeconds too short on the livenessProbe.
Your fraud model takes 25 seconds to load.
If initialDelaySeconds=10, Kubelet checks health at 10 seconds,
the model hasn't loaded yet, health returns 503,
Kubelet counts as a failure, kills the container, starts again —
and gets stuck in an infinite restart loop (CrashLoopBackOff)! 🔁✅ Always set
initialDelaySeconds on livenessProbe
to at least 2x your model's loading time.
Use startupProbe for models that take over 60 seconds to load.
🔄 Section 3: All 6 Services Working Together — Full ML Deployment Walkthrough
Let's trace exactly what happens when you run
kubectl apply -f fraud-model-deployment.yaml —
following the request through every single service.
━━━ MASTER NODE SERVICES ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
STEP 1 — API Server receives the request
• Authenticates: "Is this a valid OCI/K8s credential?" ✅
• Authorizes: "Is this user allowed to create Deployments?" ✅
• Validates: "Is the YAML well-formed? Does the image exist?" ✅
• Saves to etcd: "Desired state: 5 fraud-model Pods."
STEP 2 — etcd stores the desired state
• Writes to disk: Deployment spec with replicas=5
• Available to all other components via API Server
STEP 3 — Controller Manager detects work to do
• Deployment Controller: "Desired=5 Pods, Actual=0 Pods"
• Creates 5 Pod objects in etcd (status: "Pending")
STEP 4 — Scheduler assigns Pods to Nodes
• Sees 5 unscheduled Pods in etcd
• Filters nodes: all 3 Worker Nodes have capacity
• Scores: Worker-3 has most free RAM → gets 2 Pods
• Updates etcd: Pod-1→Worker-1, Pod-2→Worker-2, Pod-3→Worker-3,
Pod-4→Worker-1, Pod-5→Worker-3
━━━ WORKER NODE SERVICES ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
STEP 5 — Kubelet on each Worker detects its assigned Pods
• Worker-1 Kubelet sees: "I have 2 new Pods to start"
• Reads Pod spec: image, resources, env vars, probes, volumes
STEP 6 — Container Runtime pulls and starts containers
• containerd checks: "Is fraud-model:2.1.0 cached?" → No
• Pulls from OCIR (authenticated via imagePullSecret)
• Creates container with isolated filesystem + network
• Starts /app/entrypoint.sh → FastAPI + ML model loading
STEP 7 — Kubelet runs health checks
• After 20s: calls GET /health → 200 OK → readinessProbe PASSES
• Reports to API Server: "Pod-1 is Ready on Worker-1" ✅
STEP 8 — etcd updated with actual state
• etcd now shows: Desired=5, Actual=5, all Running ✅
• Controller Manager reconciliation loop confirms: "gap = zero, nothing to do"
STEP 9 — Traffic flows to your ML model 🎉
• OCI Load Balancer → Worker Nodes → Pods → ML predictions!
• Total time from kubectl apply to first request served: ~45 seconds
💥 Section 4: Service-by-Service Failure Scenarios
| Service Fails | Immediate Effect | What Still Works | Recovery |
|---|---|---|---|
| API Server | No new deployments, no scaling, no kubectl commands | Existing containers keep running normally | OKE auto-restarts it; HA has 3 replicas |
| etcd | Cluster "freezes" — no state changes possible | Existing containers keep running | OKE runs 3-node etcd cluster; OCI backs up daily |
| Scheduler | New Pods stay in "Pending" — not assigned to nodes | Already-running Pods continue unaffected | Kubernetes restarts it automatically |
| Controller Manager | No self-healing — crashes not auto-recovered | Running containers keep serving traffic | Kubernetes restarts it automatically |
| Container Runtime | No new containers can start on that Worker Node | Already-running containers may continue briefly | systemd restarts containerd; or drain the node |
| Kubelet | Node marked NotReady after 40s; Pods evicted after 5min | Containers already running stay up temporarily | systemd restarts Kubelet; Controller reschedules Pods |
✅ Best Practices & Common Mistakes
- ✅ Always set
readinessProbe— prevents traffic to Pods still loading your ML model - ✅ Set
initialDelaySecondsonlivenessProbelonger than your model loading time - ✅ Use
startupProbefor large ML models (LLMs, heavy CNNs) that take 60+ seconds to load - ✅ Always define resource
requestsandlimits— the Scheduler needs them to make smart decisions - ✅ Use
podAntiAffinity— force the Scheduler to spread your ML replicas across different Worker Nodes - ✅ On OCI: use OKE — Oracle manages the 4 Master services so you focus entirely on your ML deployment
- ✅ Monitor etcd backup status on OKE — it's your cluster's insurance policy
- ❌ Don't skip health probes — Kubelet will send traffic to containers still loading the model
- ❌ Don't set
livenessProbetoo aggressively — CrashLoopBackOff will restart your ML model repeatedly - ❌ Don't ignore "Pending" Pod status — it means the Scheduler can't find a suitable node (resource shortage!)
- ❌ Don't set resource
limitswithoutrequests— the Scheduler can't plan correctly - ❌ Don't run ML workloads directly on Master Nodes — they're reserved for control plane services only
- ❌ Don't ignore etcd warnings in OKE logs — a corrupted etcd = total cluster state loss
🏆 High-Level Summary: All 6 Services Mastered!
🧠 MASTER NODE — 4 Services (The Brain):
- 📡 API Server — The front door. Every request enters here. Authenticates, validates, routes. Nothing bypasses it.
- 📚 etcd — The cluster's memory. Stores desired + actual state. The source of truth for everything. Distributed, durable, backed up.
- 📋 Scheduler — The smart matchmaker. Decides which Worker Node runs which Pod. Filters unsuitable nodes, scores the rest, picks the winner.
- 🔄 Controller Manager — The self-healing watchdog. Watches for desired vs actual state gaps. Creates, deletes, and restarts resources automatically. Infinite reconciliation loop.
💪 WORKER NODE — 2 Services (The Muscle):
- ⚙️ Container Runtime (containerd) — The engine. Actually pulls images from OCIR, creates containers, allocates CPU/RAM, starts processes. The Kubelet's hands.
- 👂 Kubelet — The loyal agent. Registers with Master, watches for Pod assignments, instructs Container Runtime, runs health probes, reports heartbeats. The Worker Node's only voice to the Master.
You've gone from "what is a Master Node?" to understanding every component, what it does, why it exists, and how it fails.
Keep learning, keep deploying 🐼✨
Comments
Post a Comment