Skip to main content

AI Model Serving with Kubernetes: Deploy, Scale, and Manage ML Models

Calculating read time…

AI model serving on Kubernetes means running inference workloads under a scheduler and autoscaler that were originally designed for stateless web services, then deliberately adapting both — GPU-aware scheduling, request-based rather than CPU-based autoscaling, and slower, heavier "pod startup" — to fit a workload where loading a model into memory can take longer than most web requests take to complete. 🧠

A standard Kubernetes Deployment with a CPU-based Horizontal Pod Autoscaler will technically run a model server. It just won't run it well: it can't tell the difference between an idle GPU and a busy one, it scales on the wrong signal, and it has no concept of a canary rollout for a new model version versus a new code version. Purpose-built model-serving platforms on Kubernetes — KServe chief among them — exist specifically to close that gap. Skipping them doesn't mean model serving becomes simpler; it means you end up quietly rebuilding pieces of them yourself, usually under production pressure. ⚡

Note: this post focuses on content, not a diagram, consistent with the recent posts in this series — happy to add one if it would help for this topic.

🔀 Quick Comparison: GPU Sharing Strategies on Kubernetes

Choosing how to share GPUs across model-serving pods is one of the first decisions a Kubernetes-based serving platform forces you to make explicitly.

Strategy Isolation Best fit
Exclusive GPU request Full — one pod, one GPU Large models needing all of a GPU's memory and compute
Time-slicing None — no memory or fault isolation between replicas Many small, bursty workloads on older GPUs without MIG support
MIG (Multi-Instance GPU) Hardware-level memory and fault isolation per instance Multiple independent workloads needing guaranteed isolation on supported GPUs

🎯 Use this table before deciding how to size and share GPU capacity across your serving pods — the wrong choice here is expensive to discover after the fact.

1. Foundations — Why Model Serving Needs Its Own Kubernetes Abstraction

🧸 Kid-friendly analogy: A standard Kubernetes Deployment treats every pod like a light switch — flip it on, and it's instantly ready. A model-serving pod is more like a stove burner that needs a minute to heat up before it's actually useful. Scheduling and scaling logic built only for light switches will keep flipping the burner on and off, never noticing it needed time to get hot first.

A generic Kubernetes Deployment with a CPU-based Horizontal Pod Autoscaler assumes a workload where CPU usage tracks load reasonably well, where pod startup is fast, and where every replica is functionally interchangeable. Model serving breaks each of these assumptions:

  • GPU utilization, not CPU utilization, is usually the meaningful load signal — a GPU-bound inference pod can sit at low CPU usage while its GPU is fully saturated, making CPU-based autoscaling nearly blind to real load.
  • Startup is slow — loading a multi-gigabyte model into GPU memory takes meaningfully longer than starting a typical stateless web service, which changes how aggressively you can scale down and back up.
  • Rollouts need model-aware semantics — a canary for a new model version is really a quality comparison, not just a code-health check, and needs to route a controlled slice of real traffic to be meaningful.

KServe, a Kubernetes-native project, addresses this by providing an InferenceService custom resource that encapsulates autoscaling, networking, health checking, and server configuration behind a standardized interface across ML frameworks (TensorFlow, XGBoost, Scikit-learn, PyTorch, and Hugging Face Transformer/LLM models, among others), while supporting GPU autoscaling, scale-to-zero, and canary rollouts as first-class features rather than something each team reimplements per project.

2. GPU Scheduling — Exclusive Access, Time-Slicing, and MIG

🧸 Kid-friendly analogy: Exclusive GPU access is like renting an entire practice room for yourself. Time-slicing is like several bands sharing one room in shifts, with no wall between them — if one band leaves their amps too loud, the next band's soundcheck gets affected. MIG is like installing real soundproof walls to split one big room into several small, truly separate ones.

What it does: Kubernetes exposes GPUs to pods through a device plugin, most commonly as the nvidia.com/gpu resource for NVIDIA hardware. By default, a pod requesting this resource gets exclusive access to a whole GPU. NVIDIA's GPU Operator adds two mechanisms for sharing a GPU across more than one workload: time-slicing and Multi-Instance GPU (MIG).

Why it is needed: not every inference workload needs — or can afford — a full GPU to itself. A small model serving light, bursty traffic can leave most of a GPU's capacity idle if given exclusive access, while a cluster with many such workloads and few GPUs quickly becomes capacity-constrained without some form of sharing.

How it works, step by step (per NVIDIA's own documentation on GPU sharing in Kubernetes):

  1. Time-slicing lets an administrator define a set of "replicas" for a GPU, each independently assignable to a pod; internally, workloads on these replicas are multiplexed onto the same physical GPU, interleaving execution. There is no memory or fault isolation between replicas — a request for a time-sliced GPU is explicitly documented as providing shared, not exclusive, access.
  2. MIG instead partitions a GPU at the hardware level into several smaller, predefined instances, each behaving like an independent mini-GPU with its own isolated memory and fault domain — a workload on one MIG instance cannot exhaust the memory of, or crash alongside, a workload on another.
  3. Time-slicing and MIG can also be combined, sharing access to individual MIG instances via time-slicing when even more subdivision is needed.
  4. After configuring time-slicing, the GPU Operator applies a node label reflecting the replica factor (nvidia.com/<resource-name>.replicas) and marks the resource's product label as shared, so the scheduler and operators can see which nodes have oversubscribed GPUs.

What fails without understanding this trade-off: using time-slicing for a workload that actually needs guaranteed memory headroom risks one pod's memory usage starving another sharing the same physical GPU, since time-slicing provides no memory isolation — a failure mode that looks like random, hard-to-reproduce out-of-memory errors rather than a clear resource-limit violation.

💡 Trade-off: per NVIDIA's own documentation, GPU metrics tooling such as DCGM-Exporter does not support associating metrics back to individual containers when time-slicing is enabled — a real observability cost that comes bundled with the sharing flexibility, and one that's easy to discover only after you've already lost per-workload GPU visibility in production.

3. Autoscaling — Concurrency, Scale-to-Zero, and Cold Starts

🧸 Kid-friendly analogy: Scaling a model server on request concurrency is like a restaurant host counting how many tables are currently occupied, not how loudly the kitchen sounds — occupied tables is the number that actually tells you whether to open another section.

What it does: rather than scaling on CPU utilization, KServe's default autoscaling (built on Knative Serving) scales based on request concurrency — the number of in-flight requests per pod — against a configurable target, with support for scaling all the way down to zero replicas when there's no traffic, including for GPU-backed pods.

Why it is needed: concurrency is a far more direct signal of whether a model-serving pod needs help than CPU usage, especially for GPU-bound inference where CPU can stay low even under heavy load. Scale-to-zero, in turn, means a rarely-used model doesn't have to permanently reserve GPU capacity it's not using.

How it works, step by step (using KServe's own documented configuration):

  1. An InferenceService is configured with a scaleTarget and scaleMetric (for example, a concurrency target of 1 request per pod), or equivalently via a Knative annotation such as autoscaling.knative.dev/target.
  2. As concurrent requests per pod approach that target, the autoscaler adds replicas; as concurrency drops, it removes them — down to zero if traffic stops entirely and scale-to-zero is enabled.
  3. When a new request arrives to a scaled-to-zero service, a replica has to start and, for a model server, load the model into memory before it can respond — this cold-start latency is the direct cost of scaling to zero.
  4. KServe's documentation notes this target is a soft limit: a sudden burst of requests can temporarily exceed it while new replicas are still starting.

What fails without accounting for cold starts: a service that scales to zero aggressively, with no minimum replica count, can make a user's very first request after an idle period wait for a full model load — acceptable for some internal, occasional-use tools, but a poor experience for anything latency-sensitive or user-facing at unpredictable times.

✅ Worked example: KServe's own documentation walks through creating a TensorFlow-model InferenceService with a concurrency target of 1 and then load-testing it with sustained concurrent requests, observing the autoscaler add replicas to keep per-pod concurrency near the configured target — a concrete illustration of concurrency-based, not CPU-based, scaling in action.

4. Worked Example — A Canary Rollout for a Fine-Tuned Model (hypothetical, for illustration)

Consider a hypothetical team that has fine-tuned a new version of a model already serving production traffic and wants to validate it against real usage before a full cutover.

How it works, step by step:

  1. The team deploys the new model version as a second revision of the same InferenceService, rather than a separate, unrelated deployment.
  2. Traffic is split so a small percentage — for example, 5% — routes to the new revision, with the rest continuing to the current, proven version.
  3. The team compares latency, error rate, and (separately) output-quality signals between the two revisions over a representative traffic window, rather than judging on gut feel after a few manual test queries.
  4. If the new revision performs at least as well as the old one across those signals, traffic is gradually shifted further, ending in a full cutover; if not, traffic is shifted back to the proven revision with no user-facing disruption.

What this illustrates: the platform-level canary mechanism (traffic splitting between revisions) is necessary but not sufficient — it has to be paired with the same representative evaluation discipline that any model change deserves, or the canary just becomes a slower way to ship a regression instead of a way to catch one.

🎯 Use this pattern for any model version change reaching real production traffic, not just for application code changes.

5. Implementation — InferenceService and GPU-Sharing Configuration

A minimal InferenceService with a concurrency-based autoscaling target and a GPU resource request:

# Illustrative example — verify current API version and fields against KServe's own documentation
apiVersion: "serving.kserve.io/v1beta1"
kind: "InferenceService"
metadata:
  name: "example-model"
spec:
  predictor:
    scaleTarget: 4
    scaleMetric: concurrency
    minReplicas: 1     # avoids a full cold start on every idle period
    model:
      modelFormat:
        name: pytorch
      storageUri: "gs://example-bucket/models/example-model"
      resources:
        limits:
          nvidia.com/gpu: "1"

And an illustrative time-slicing configuration applied at the node level through the NVIDIA GPU Operator, allowing four pods to share one physical GPU:

# Illustrative example — verify current ConfigMap schema against NVIDIA's own documentation
version: v1
sharing:
  timeSlicing:
    resources:
      - name: nvidia.com/gpu
        replicas: 4

Illustrative examples only — confirm current API fields, resource names, and ConfigMap schemas against KServe's and NVIDIA's own documentation before use.

Two implementation details are easy to overlook: setting a minimum replica count for anything latency-sensitive avoids paying the full model-load cold start on every idle period, and matching the GPU-sharing strategy to the workload's actual memory needs (time-slicing only where memory isolation genuinely isn't required) prevents a resource-sharing decision from becoming a stability problem later.

6. Enterprise Rollout — Multi-Model Density, Observability, and Governance

Beyond a single model's autoscaling and GPU configuration, running many models in production on Kubernetes raises its own set of concerns:

  • High-density multi-model serving. For organizations serving many models, each with modest individual traffic, KServe's ModelMesh component is designed specifically to pack more models onto shared infrastructure with intelligent routing, rather than dedicating a full pod (and GPU allocation) per model regardless of its actual traffic.
  • Composed inference pipelines. Real serving flows often need pre-processing, the model call itself, and post-processing as connected steps; KServe's InferenceGraph is designed for exactly this kind of pipeline, ensemble, or multi-step inference flow rather than treating every step as an unrelated service.
  • Per-workload GPU observability. Given the documented gap in per-container GPU metrics under time-slicing, teams sharing GPUs this way need a deliberate plan for what observability they're willing to trade away, and where exclusive GPU access or MIG might be worth the cost for workloads that need clear per-workload visibility.
  • Rollout and rollback discipline. Canary traffic splitting between model revisions should be paired with defined promotion criteria and a fast path back to the previous revision, treated with the same rigor as any other production release process.
  • Capacity planning around cold starts. Decide deliberately, per service, whether scale-to-zero's GPU savings are worth its cold-start latency cost, rather than applying one blanket autoscaling policy to every model regardless of how latency-sensitive it is.
  • Access control over InferenceService creation and GPU quotas. Because GPU capacity is finite and expensive, treat who can create or resize GPU-backed InferenceServices as a governed permission, not an open action available to any cluster user.

7. Common Mistakes

  • Using CPU-based autoscaling for GPU-bound inference. A GPU-saturated, CPU-idle pod looks "fine" to a CPU-based HPA, so the autoscaler never reacts to the load that's actually happening.
  • Enabling scale-to-zero without considering cold-start impact. This trades GPU cost savings for unpredictable first-request latency; fine for some internal tools, a poor choice for anything latency-sensitive without a minimum replica floor.
  • Using time-slicing for workloads that need memory isolation. Since time-slicing provides no memory or fault isolation between replicas, a memory-hungry neighbor on the same physical GPU can cause confusing, intermittent failures in an otherwise well-behaved workload.
  • Losing per-container GPU observability without realizing it. Enabling time-slicing without accounting for the documented DCGM-Exporter metric-association gap can leave a team unable to tell which specific workload is actually driving GPU load during an incident.
  • Treating a canary rollout as sufficient without a real evaluation plan. Splitting traffic to a new model revision only helps if latency, error rate, and quality are actually compared against the current revision — routing traffic without measuring anything defeats the purpose of canarying in the first place.
  • Not setting a minimum replica count for latency-sensitive services. Letting every service scale fully to zero by default, regardless of its latency requirements, is a blanket policy applied where a per-service decision was actually needed.

❓ FAQ

Why doesn't a standard Kubernetes HPA work well for model serving?

Because it typically scales on CPU utilization, which can stay low even when a GPU-bound inference pod is fully saturated. Request-concurrency-based autoscaling, as used by KServe, is a more direct signal of real inference load.

What's the difference between GPU time-slicing and MIG?

Time-slicing multiplexes workloads onto the same physical GPU with no memory or fault isolation between them. MIG partitions a GPU at the hardware level into isolated instances, each with its own memory and fault domain — stronger isolation, but requiring supported GPU hardware.

Does scaling a GPU-backed service to zero save real cost?

Yes, since an idle service isn't holding onto GPU capacity it isn't using — but that savings comes with cold-start latency on the next request, since a replica has to start and load the model before it can respond.

What is KServe's ModelMesh for?

It's designed for high-density, multi-model serving — packing many models with intelligent routing onto shared infrastructure, rather than requiring a dedicated pod and GPU allocation for every individual model regardless of its traffic level.

Is a canary rollout enough to safely ship a new model version?

The traffic-splitting mechanism itself is necessary but not sufficient — it needs to be paired with a real comparison of latency, error rate, and output quality between the current and new revisions, not just a smaller blast radius if something goes wrong.

🔗 References & Further Reading

"Kubernetes," "KServe," "NVIDIA," and related marks are trademarks of their respective owners.

📝 Summary

  • Model serving breaks the assumptions generic Kubernetes autoscaling relies on: CPU usage doesn't reflect GPU load, startup is slow, and rollouts need model-aware semantics.
  • GPU sharing comes in two flavors with a real trade-off — time-slicing for flexibility with no isolation, MIG for hardware-level isolation with narrower hardware support.
  • Concurrency-based autoscaling, including scale-to-zero, fits inference load far better than CPU-based scaling, but scale-to-zero trades cost savings for cold-start latency.
  • A worked canary example shows that traffic-splitting infrastructure only pays off when paired with genuine latency, error-rate, and quality comparison between model revisions.
  • A real InferenceService and GPU-sharing configuration show how these decisions are expressed concretely, not just conceptually.
  • Running many models well at enterprise scale adds density (ModelMesh), pipeline composition (InferenceGraph), and governance concerns on top of single-model configuration.
  • Most real incidents trace back to a short list of avoidable mistakes: CPU-based autoscaling on GPU workloads, careless scale-to-zero, and GPU-sharing strategies mismatched to a workload's actual isolation needs.

Hopefully this gives you a clear, first-principles map of what makes AI model serving on Kubernetes genuinely different from serving an ordinary web application — and which setting to check first when scaling or GPU sharing isn't behaving the way you expect. 🙌

Comments