A GPU container is a normal Docker container that has been given a narrow, controlled window into your host's physical GPU — the container never contains a GPU itself; it borrows one through a driver, a runtime hook, and a device file, so the same image can run identically on a laptop, a bare-metal server, or a Kubernetes cluster node. 🧩
Getting this wrong is not a cosmetic bug: a misconfigured GPU container silently falls back to CPU (an LLM that should answer in 300ms takes 40 seconds), or it grabs every GPU on a shared node and starves three other teams' jobs, or it ships a 14GB image because CUDA layers were duplicated across every service. In production AI infrastructure, GPU containerization is the boundary between an ML idea that works on one engineer's workstation and a system that survives real traffic, real cost pressure, and real on-call rotations. ⚙️
Figure 1: The path a workload takes from a container image to silicon — the same chain repeats, node by node, inside a Kubernetes cluster.
📑 In This Post
- 1. Foundations: what a GPU container actually is
- 2. Mechanics: how a container reaches a physical GPU
- 3. Real example: your first GPU-accelerated container
- 4. Implementation: containerizing an LLM inference service
- 5. Scaling out: scheduling GPUs with Kubernetes
- 6. Enterprise rollout: governance, cost, and reliability
- 7. Common mistakes
- 8. FAQ
- 9. References & further reading
- 10. Summary
🔀 Quick Comparison: Three Ways to Give a Workload a GPU
| Approach | What isolates the workload | Best fit | Main weakness |
|---|---|---|---|
| Bare-metal install | Nothing — process runs directly on the host with the driver and CUDA toolkit installed globally | Single-user research boxes | Dependency conflicts between projects; no portability |
| Docker + NVIDIA Container Toolkit | A container namespace; the toolkit injects driver libraries and device files at container start | One GPU host, one or a few services, local dev, CI | No cross-node scheduling, quota, or bin-packing on its own |
| Kubernetes + GPU device plugin | Pods scheduled onto specific GPU nodes based on the nvidia.com/gpu resource | Multi-service, multi-team, multi-node production fleets | More operational surface: node labels, plugin daemonsets, quotas |
🎯 Use this when: choosing how much orchestration your GPU workload genuinely needs before you reach for Kubernetes by default.
1. Foundations: What a GPU Container Actually Is
Kid-friendly analogy: Imagine a shipping container with all your toys packed inside, and one special hole cut in the side with a hose running to your house's water tap. The toys stay boxed and portable, but that one hose lets water flow in from outside whenever you need it. A GPU container is boxed the same way — except the "hose" carries GPU commands, not water, and it only exists because someone deliberately drilled that hole and connected it.
Technically: a container is an isolated Linux process — its own filesystem, its own view of processes, its own network namespace — but it still shares the host machine's kernel. A GPU is a physical PCIe device owned by that same kernel, exposed through character device files such as /dev/nvidia0 and a set of driver libraries. By default, Docker does not expose any host devices to a container, which is a deliberate isolation boundary, not an oversight. A GPU container is simply a container that has been given explicit, narrow permission to see specific GPU device files and load the matching driver libraries from the host.
This matters because the GPU driver itself must match the physical card and cannot be virtualized cheaply the way a CPU instruction set can. According to Docker's own documentation, exposing an NVIDIA GPU to a container requires installing the correct NVIDIA driver on the host first, and then installing the NVIDIA Container Toolkit — the driver is a host-level concern, not something baked into the application image.
What fails without this separation: teams either install CUDA toolkits directly on shared hosts (version conflicts the moment two teams need different CUDA releases), or they try to bundle a driver inside the image itself, which breaks the moment the container is scheduled onto a node with a different physical GPU or host driver version.
2. Mechanics: How a Container Reaches a Physical GPU
Kid-friendly analogy: Picture a hotel bellhop who meets every guest (container) at the door. Before the guest even unpacks, the bellhop quietly slips the room key (device access) and a house map (driver libraries) into their bag. The guest never had to ask the front desk directly — the bellhop does it automatically, every single time, based on a standing instruction.
The "bellhop" here is the NVIDIA Container Toolkit, and the mechanism is a container-runtime hook, not magic baked into Docker itself. Step by step, when you start a GPU-enabled container:
- Docker's CLI parses the
--gpusflag and passes GPU capability requirements down to the container runtime. Including the --gpus flag when starting a container is what exposes GPU resources to it. - The low-level runtime (typically
runc) invokes a pre-start hook registered by the NVIDIA Container Toolkit, before the container's own entrypoint runs. - That hook inspects the host: which driver version is installed, which GPUs exist, and which capability set was requested (compute, utility, video, and so on).
- It bind-mounts the matching driver shared libraries and the requested GPU device files into the new container's filesystem and device namespace.
- Only then does the container's actual process start, and from inside, CUDA-aware libraries detect the GPU exactly as they would on bare metal.
Two details here separate correct architecture from cargo-culted commands. First, the toolkit's own project documentation is explicit that you do not need the CUDA Toolkit installed on the host system — only the NVIDIA driver is required there; the CUDA runtime and libraries the application needs normally live inside the image or are supplied by an official CUDA base image. Second, GPU capability sets are configurable rather than all-or-nothing: you can request only the "utility" capability, for example, which adds tools like nvidia-smi without granting full compute access — useful for diagnostic sidecars that should never run inference workloads.
What fails without the toolkit's hook: nothing crashes loudly. The container starts, the application imports its ML framework, and the framework quietly reports zero visible devices, silently falling back to CPU. This is one of the most expensive silent failures in GPU containerization because throughput drops by an order of magnitude with no error message at all.
✅ Worked example: A team building an inference API noticed p95 latency jump from 280ms to 6.4 seconds after a routine base-image rebuild. The new base image had a newer CUDA runtime than the driver installed on the fleet's older nodes supported. The GPU was technically "attached," but the ML framework refused to initialize the CUDA context and fell back to CPU without raising an exception. The fix was pinning the CUDA runtime inside the image to a version explicitly compatible with the minimum driver version running on any node in the fleet, and adding a startup health check that asserts a GPU device is actually visible before accepting traffic.
3. Real Example: Your First GPU-Accelerated Container
This section is deliberately runnable rather than theoretical. Docker's documentation shows the canonical smoke test as running an interactive container with the --gpus all flag against a base image and calling nvidia-smi to confirm the GPU is visible.
- Confirm the host is ready. The NVIDIA driver must already be installed and working (verified with
nvidia-smion the host itself), and the NVIDIA Container Toolkit must be installed and Docker restarted afterward. - Run a minimal smoke test to prove the toolkit is wired correctly before touching your own image:
docker run -it --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
- Read the output carefully. A correct result lists driver version, CUDA version, and at least one GPU with memory figures — not an error about a missing device or driver mismatch.
- Scope GPU visibility deliberately on multi-GPU hosts rather than defaulting to
all:docker run -it --rm --gpus '"device=0"' nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
On systems with multiple GPUs, individual devices can be exposed by index, and the device value must be quoted when it contains a comma for multi-device selection. - Layer your application on top of an official CUDA base image rather than installing CUDA by hand, so version alignment with the driver stays predictable across the fleet.
💡 Trade-off to know: --gpus all is the fastest way to get unblocked locally, but on a shared host it silently hands one process every GPU on the machine. In any environment with more than one workload, scope the flag to specific device indices or UUIDs from the start — retrofitting device scoping after an incident is far more disruptive than starting with it.
4. Implementation: Containerizing an LLM Inference Service
Kid-friendly analogy: A single home cook (a script) can make one sandwich at a time, well, but slowly if forty people show up at once. A restaurant kitchen (an inference server) is built from the ground up to take many orders at once, batch similar ones together, and keep the grill (the GPU) busy instead of idle between orders. Wrapping a model in a serving framework is the difference between a home cook and a kitchen.
A model file plus a Python script is not a production service — it is a single-request tool that leaves the GPU idle between calls and cannot serve concurrent users safely. Purpose-built inference servers such as vLLM, NVIDIA Triton Inference Server, and Ollama exist specifically to batch concurrent requests, manage GPU memory across a running model, and expose a stable network API, and each ships an official container image built on top of a CUDA-enabled base rather than requiring you to assemble the CUDA stack yourself.
What it does: the serving container loads model weights once at startup, keeps them resident in GPU memory, and accepts a stream of independent requests over HTTP or gRPC, scheduling them onto the GPU efficiently instead of one process per request.
Why it's needed: GPU memory is a fixed, expensive resource; reloading multi-gigabyte weights per request would make latency and cost unworkable, and naive one-request-per-process designs waste the GPU's ability to batch work.
How it works, step by step:
- The container starts and the serving framework reads model weights from a mounted volume or a model cache directory, rather than baking multi-gigabyte weights into the image layer itself.
- The framework allocates a GPU memory pool sized for the model and a configured batch/concurrency ceiling.
- Incoming requests are queued and grouped into batches the GPU can process together, based on the serving framework's own scheduling policy.
- A health endpoint reports readiness only after the model is fully loaded onto the GPU — this is the signal an orchestrator should wait for before sending traffic.
- Metrics such as queue depth, tokens per second, and GPU memory utilization are exposed for the observability stack described later in this post.
What fails without it: without a real serving layer, teams see request queuing that looks like a network problem, GPU out-of-memory crashes under concurrent load, and cold-start latency spikes every time a container restarts and reloads weights from scratch.
Best practices: separate the model-weight cache from the application image (mounted volume or object storage) so a code change doesn't force a multi-gigabyte rebuild; pin the serving framework version explicitly, since batching and memory-management behavior changes between releases; and load-test with realistic concurrent traffic patterns, not single sequential requests, before trusting any latency number.
🎯 Use this when: you are moving a model from a notebook or a single script into anything more than one concurrent user.
5. Scaling Out: Scheduling GPUs with Kubernetes
Kid-friendly analogy: One host with a GPU is one classroom with one whiteboard. A Kubernetes cluster is a whole school building, and the scheduler is the office staff deciding which classroom (node) has a free whiteboard (GPU) before sending a class (a Pod) there — instead of every teacher wandering the halls checking rooms themselves.
A single Docker host runs out of usefulness the moment you have more than one GPU-hungry service or more than one team. Kubernetes has included stable support, since version 1.26, for managing NVIDIA and AMD GPUs across different nodes in a cluster using device plugins. The device plugin is a per-node agent, not part of core Kubernetes itself: an administrator installs GPU drivers from the hardware vendor on each node and runs that vendor's device plugin, after which the cluster exposes a custom schedulable resource such as nvidia.com/gpu.
Requesting a GPU in a Pod spec looks like requesting CPU or memory, with one important restriction: GPUs may only be specified in the limits section of a container's resources — a limit can be set without a matching request, since Kubernetes treats the limit as the request by default, but if both are set they must be equal, and a request cannot be set without a limit. In practice this means GPUs in Kubernetes are never fractionally overcommitted by the scheduler the way CPU can be — a Pod either gets a whole GPU unit it asked for, or it does not get scheduled onto that node at all, unless the cluster additionally layers in a GPU-sharing add-on.
- Label and taint GPU nodes distinctly from CPU-only nodes, so the scheduler and human operators can both reason about where GPU work lands.
- Install the vendor device plugin as a DaemonSet so every GPU node advertises its GPU count as a schedulable resource automatically.
- Set resource limits on every GPU-consuming Pod explicitly — never rely on defaults — and treat an un-set GPU limit as a policy violation, not an oversight to catch later.
- Combine node selectors or node affinity with the GPU resource request when a cluster mixes GPU generations, so a model built for one architecture doesn't land on incompatible hardware.
- Add readiness probes tied to actual model-load completion (see the serving section above), so the Kubernetes Service only routes traffic to Pods that have finished loading weights onto the GPU.
What fails without this discipline: Pods that request no GPU limit can still be scheduled onto a GPU node and consume none of the accelerator while occupying general capacity, and Pods that assume implicit GPU access without a limit will crash or silently run on CPU depending on the base image, exactly as in the single-host case.
6. Enterprise Rollout: Governance, Cost, and Reliability
Running GPU containers reliably in an organization is an operating model question as much as a technical one. The following areas consistently separate teams that scale GPU infrastructure smoothly from teams that firefight it:
- Ownership and access control: define who can request GPU quota, who approves new GPU-backed deployments, and how service accounts are scoped so one team's workload cannot exhaust a shared node pool's capacity.
- Image and driver versioning: pin base image tags and CUDA runtime versions per service, and track the minimum driver version each fleet node must run; treat a driver upgrade as a coordinated change with a rollback plan, not a routine patch.
- Dataset and model-artifact versioning: version model weights and evaluation datasets alongside code, so a rollback of a serving container also rolls back to a known-good model artifact rather than reintroducing a regression.
- CI gates before GPU rollout: require an automated smoke test that confirms GPU visibility inside the built image (not just that the image builds), plus a latency/throughput regression check against a fixed evaluation set before promoting a new serving image.
- Cost and budget controls: GPU capacity is typically the largest single line item in an AI infrastructure budget; enforce quotas per team or namespace, and alert on GPU utilization that stays persistently low, since an idle reserved GPU is a direct, ongoing cost.
- Dashboards and alerting: track GPU memory utilization, queue depth, tokens or requests per second, and error rate together — a spike in errors with flat GPU utilization usually points to a driver or scheduling problem rather than a model problem.
- Canary releases and rollback criteria: roll a new serving image to a small percentage of GPU capacity first, compare latency and output-quality signals against the current production image, and define numeric rollback thresholds in advance rather than deciding reactively during an incident.
- Privacy of production-derived data: logs and traces captured from GPU inference services often contain real user prompts or outputs; apply the same data-handling and retention policies to these logs as to any other production PII surface.
- Incident response: document the specific signature of a "silent CPU fallback" incident (available earlier in this post) as a named runbook entry, since it is a repeatable, previously-seen failure mode rather than a novel one each time.
🎯 Use this when: a GPU-backed service is about to take real user traffic, not just internal or staging traffic.
7. Common Mistakes
These mistakes recur across teams because they are all "quiet" failures — nothing throws an error, so they survive until someone notices a cost or latency anomaly.
- Assuming the CUDA toolkit needs to be installed on the host. Official toolkit documentation states plainly that the CUDA Toolkit does not need to be installed on the host system — only the NVIDIA driver does. Teams that install a full CUDA toolkit on every host anyway create an extra, unmanaged version surface that has to be kept in sync with every image separately, multiplying upgrade risk for no benefit.
- Defaulting to
--gpus allon shared or multi-tenant hosts. This works fine on a single-user dev box and becomes a resource-starvation incident the moment a second workload lands on the same host, because the first process silently claims every GPU. - Treating a missing GPU as a silent CPU fallback instead of a hard failure. Many ML frameworks degrade to CPU without raising an exception, which turns a configuration bug into a mysterious performance regression discovered by users rather than by monitoring.
- Baking multi-gigabyte model weights into the container image. This makes every code change trigger a full model re-upload, slows every deployment and rollback, and makes image registries expensive to operate; weights belong in a mounted volume or object storage, referenced at startup.
- Leaving GPU resource limits unset on Kubernetes Pods. Because GPU requests cannot be set without a matching limit, an unset limit is not a minor omission — it usually means the Pod either won't schedule as intended or won't get the isolation the team assumes it has.
- Skipping a real concurrency load test before production sign-off. A model that answers instantly to one sequential request can still queue, throttle, or OOM under ten concurrent ones; sequential testing hides exactly the failure mode production traffic will trigger first.
❓ FAQ
Do I need to install the CUDA Toolkit on my Docker host?
No. The NVIDIA Container Toolkit's own documentation confirms that only the NVIDIA driver needs to be installed on the host — the CUDA Toolkit itself does not. The CUDA runtime your application needs typically comes from the container's base image instead.
What's the actual difference between --gpus all and requesting a specific device?
--gpus all exposes every GPU on the host to that one container. Specifying a device by index or UUID instead exposes only that particular GPU, which is the safer default on any host running more than one workload.
Can Kubernetes split one GPU across multiple Pods?
Not through the core scheduling mechanism described here. Standard GPU scheduling in Kubernetes allocates whole GPU units through the nvidia.com/gpu resource, specified only in a container's limits; fractional or shared-GPU access requires an additional vendor-specific add-on layered on top of this base mechanism.
Why did my model silently run on CPU instead of throwing an error?
Most ML frameworks check for GPU availability and fall back to CPU rather than crashing, as a usability default. The fix is operational, not code-level: add a startup check inside the container that explicitly asserts a GPU device is visible and fails the health check if it is not, so the failure surfaces immediately instead of as a slow-motion latency mystery.
Should model weights live inside the container image?
Generally no for anything beyond small models. Keeping weights in a mounted volume or object storage, separate from the application image layers, keeps image builds fast, keeps rollbacks cheap, and lets you version model artifacts independently from application code — which also supports the CI and rollback practices covered in the enterprise rollout section.
🔗 References & Further Reading
- Docker Engine documentation — GPU access
- NVIDIA Container Toolkit — official GitHub repository
- Kubernetes documentation — Schedule GPUs
NVIDIA, CUDA, Docker, and Kubernetes are trademarks of their respective owners.
📝 Summary
- Foundations: a GPU container is a normal container given narrow, explicit access to host GPU device files and driver libraries.
- Mechanics: the NVIDIA Container Toolkit's runtime hook injects driver libraries and device access before the container's process starts.
- Real example: validate GPU access with a minimal
nvidia-smismoke test before layering on an application. - Implementation: a real serving framework, not a bare script, is what turns a model into a production-grade inference service.
- Kubernetes: GPUs are scheduled as whole units via the nvidia.com/gpu resource, specified only in limits.
- Enterprise rollout: ownership, versioning, CI gates, cost controls, and canary rollback criteria are what make GPU infrastructure operable at scale.
- Common mistakes: nearly all of them are silent failures — unset limits, silent CPU fallback, and oversized images — that surface as cost or latency problems, not errors.
Thanks for reading — build your first smoke-tested GPU container this week, and let that habit carry all the way into your production rollout checklist. 🚀
Comments
Post a Comment