CUDA is NVIDIA's parallel computing platform and programming model for running general-purpose code on GPUs, while the NVIDIA Container Toolkit is the set of components that lets a Docker or Kubernetes container reach a host GPU without the container needing its own copy of the NVIDIA driver. Together they answer the two questions every GPU-accelerated container has to solve: which parallel-computing runtime does my code link against, and how does a sandboxed container process get a device handle to hardware that technically belongs to the host kernel?
Getting this stack wrong shows up as some of the most confusing failures in production AI infrastructure: a container that runs perfectly on one node and reports "no CUDA-capable device" on another with the identical image, a Kubernetes pod that schedules successfully but silently sees zero GPUs, or a costly A100 sitting at 12% utilization because nothing shares it. Understanding how the driver, the toolkit, and the CUDA runtime inside the image actually divide responsibility is what turns those into predictable, diagnosable problems instead of mysteries. ⚙️
Figure 1 — Original diagram: the driver lives on the host once; the CUDA Toolkit version travels inside each container image.
📑 In This Post
- 1. Foundations: CUDA, the Driver, and the Toolkit Are Three Different Things
- 2. Mechanics: How the NVIDIA Container Toolkit Actually Injects a GPU
- 3. CUDA Compatibility: Why the Container's CUDA Version Doesn't Have to Match the Driver
- 4. Worked Example: Running a CUDA Container with Docker
- 5. Implementation: GPU Scheduling in Kubernetes
- 6. Sharing a GPU: MIG and Time-Slicing
- 7. Enterprise Rollout: Fleet Management, Observability, and Incident Response
- 8. Common Mistakes
- 9. FAQ
- 10. References & Further Reading
- 11. Summary
🔀 Quick Comparison: Where Does GPU Sharing Actually Happen?
| Sharing Approach | Isolation Level | Hardware Requirement | Typical Fit |
|---|---|---|---|
| Whole GPU per container | Full — one process owns the device | Any NVIDIA GPU | Heavy training, latency-sensitive serving |
| Multi-Instance GPU (MIG) | Hardware-level — separate memory, cache, SMs | Ampere-and-later data-center GPUs (e.g. A100, H100) | Multi-tenant inference needing guaranteed QoS |
| Time-slicing (device plugin) | Software — GPUs advertised multiple times | Any CUDA-capable GPU | Bursty, low-priority, or dev/test workloads |
| One GPU, one job (K8s default) | Full, enforced by the scheduler | Any NVIDIA GPU | Simple clusters, cost-insensitive early stage |
🎯 Use this when deciding whether a workload needs a dedicated GPU, a hardware-isolated slice, or a time-shared slice.
1. Foundations: CUDA, the Driver, and the Toolkit Are Three Different Things
Kid-friendly analogy: Think of the GPU as a power outlet built into the wall of a building (the host). The driver is the building's wiring — installed once, permanently, by the building owner. CUDA is the standard plug shape every appliance (container) is built around. The NVIDIA Container Toolkit is the extension cord that safely reaches from a room's outlet (the container's isolated filesystem) to the wiring behind the wall, without the appliance needing to know anything about how the wiring was installed.
Three names get conflated constantly, and the confusion causes real production bugs:
- The NVIDIA driver is kernel-level software installed on the host machine. It is the only component that talks directly to the physical GPU, and it is installed once per node — never inside a container image.
- The CUDA Toolkit is a set of libraries, a compiler (nvcc), and a runtime that applications link against to run parallel code on the GPU. This is what typically lives inside a container image, at whatever version the application was built for.
- The NVIDIA Container Toolkit is the bridge: a container runtime library (libnvidia-container) and utilities that automatically configure a container to see the host's GPU devices and driver libraries at start time.
A frequently overlooked but important consequence of this split: the host does not need the CUDA Toolkit installed at all — only the NVIDIA driver. The CUDA Toolkit is entirely the container's responsibility, which is exactly why the same host can run one container built against CUDA 12.4 and another built against CUDA 11.8 side by side, as long as the host driver is new enough to support both.
Why it is needed. Without this separation, every application on a shared GPU host would need to agree on one CUDA Toolkit version installed system-wide — an impossible constraint the moment a team runs both a legacy training pipeline and a newer inference service on the same fleet. Containerizing the CUDA Toolkit, while leaving the driver on the host, is what makes mixed-version GPU fleets practical at all.
2. Mechanics: How the NVIDIA Container Toolkit Actually Injects a GPU
Kid-friendly analogy: Normally a shipping container (the Linux container) arrives sealed, with no doors cut into its walls to the outside world. The NVIDIA Container Toolkit is like a specialized dockworker who, right as the container is opened for use, cuts precisely the doors needed to reach specific machines on the dock (the GPU devices) and the tools those machines require (the driver libraries) — and reseals everything else.
What it does. By default, a container has no access to host devices. The NVIDIA Container Toolkit modifies the container creation process so that specific GPU device files and the host driver's shared libraries are made visible inside the container's filesystem namespace at start time — without baking driver binaries into the image itself.
How it works, step by step:
- The host installs the NVIDIA driver and the NVIDIA Container Toolkit packages; no CUDA Toolkit installation is required on the host itself.
- A container is launched with a GPU request — Docker's --gpus flag, or the NVIDIA_VISIBLE_DEVICES environment variable together with the nvidia runtime.
- The container runtime (Docker, containerd, or CRI-O) hands off to nvidia-container-runtime, which wraps the standard OCI runtime.
- libnvidia-container inspects which GPUs were requested and injects the matching device nodes and the host's driver shared libraries into the container's mount namespace.
- The application inside the container, linked against its own bundled CUDA Toolkit version, calls into those injected driver libraries exactly as it would on a bare-metal host.
What fails without it. A container started without the toolkit's runtime hook sees no GPU devices at all — nvidia-smi inside the container reports no devices found, even though the host has working GPUs, because the container's isolated namespace simply never had the device files or driver libraries mapped into it.
Enumeration and driver capabilities. The environment variable NVIDIA_VISIBLE_DEVICES controls exactly which GPUs a container can see — a comma-separated list of indexes or UUIDs, all (the default in base CUDA images), or none. A companion variable, NVIDIA_DRIVER_CAPABILITIES, controls which driver feature sets (compute, video, graphics, utility, and others) get mapped in, so a headless inference container is not carrying graphics-driver surface area it will never use.
✅ Practical example. On a host with four GPUs, requesting docker run --gpus '"device=1,2"' myimage makes exactly GPUs 1 and 2 visible inside the container — GPUs 0 and 3 remain completely invisible to that container's process, which is how multiple isolated jobs share one physical host safely.
💡 Trade-off. Leaving NVIDIA_VISIBLE_DEVICES=all as the default inside a base CUDA image means a container launched without an explicit GPU request may still see every GPU on the host — the documented device-plugin warning for Kubernetes exists precisely because unrequested containers otherwise get unrestricted access, which is both a scheduling and a security concern on shared nodes.
🎯 Use this when debugging a "no CUDA-capable device" error or deciding how narrowly to scope GPU visibility on a shared host.
3. CUDA Compatibility: Why the Container's CUDA Version Doesn't Have to Match the Driver
Kid-friendly analogy: A newer video game (a newer CUDA Toolkit) usually still runs on last year's game console (an older driver) as long as the console got its firmware updates within the same console generation. Jump to a console from a completely different generation, though, and the game needs a special compatibility patch just to boot at all.
What it does. NVIDIA's CUDA Compatibility guarantees define exactly when a container's bundled CUDA Toolkit version and the host's installed driver version are allowed to differ, so teams do not have to upgrade every node's driver in lockstep with every application's CUDA dependency.
Why it is needed. Driver upgrades on production data-center hosts follow qualification and maintenance schedules that rarely match an individual team's release cadence. Without a defined compatibility model, shipping a container built against a newer CUDA Toolkit would force an immediate, fleet-wide driver upgrade before it could run anywhere.
How it works, step by step:
- Backward compatibility is implicit: a newer host driver always supports containers built against older CUDA Toolkit versions, with no extra configuration.
- Minor Version Compatibility (MVC), available from CUDA 11 onward, lets a container built with a newer CUDA Toolkit minor release run on an older driver, as long as both share the same CUDA major version — for example, a CUDA 12.9 container on a driver that only natively supports CUDA 12.8.
- Forward Compatibility extends this across a major version boundary, but only by installing an explicit cuda-compat-<major>-<minor> package alongside the older driver, and only where the platform and GPU generation support it.
- Outside those documented paths — for example, a CUDA major version genuinely newer than anything the installed driver family supports, with no forward-compatibility package installed — the container simply cannot initialize CUDA, regardless of how well-formed the image is.
What fails without understanding this. A team rebuilds its inference image on a newer CUDA base without checking the fleet's driver versions, and pods that worked in staging (on freshly imaged nodes) fail to initialize CUDA in production on older nodes whose driver predates the required minimum — a failure mode that looks like a code bug but is actually a driver/toolkit version mismatch outside the documented compatibility window.
✅ Practical example. A platform team standardizes on driver 570 across its GPU fleet. Because that driver corresponds to CUDA 12.8, application teams can freely ship containers built against CUDA Toolkit 12.8 or 12.9 under Minor Version Compatibility, without coordinating a driver upgrade for every application release — only a jump to CUDA 13 would require revisiting the driver baseline.
🎯 Use this when planning how often to upgrade node drivers versus how freely application teams can bump their container's CUDA Toolkit version.
4. Worked Example: Running a CUDA Container with Docker
This walkthrough follows the documented installation and runtime steps for the NVIDIA Container Toolkit, using a labeled hypothetical team as the scenario.
What it does. A single Ubuntu GPU host is prepared so that any correctly built CUDA image can run on it, without installing the CUDA Toolkit system-wide.
How it works, step by step:
- Install the NVIDIA GPU driver for the Linux distribution — via the distribution's package manager, per NVIDIA's own installation guidance.
- Configure the NVIDIA Container Toolkit's package repository and install the nvidia-container-toolkit package through the OS package manager.
- Confirm the container runtime (Docker, in this example) is configured to use the nvidia runtime, either as the default or invoked per run.
- Verify with a minimal smoke test before deploying anything real.
# Illustrative example — install toolkit (Debian/Ubuntu), then verify curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \ sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg sudo apt-get update && sudo apt-get install -y nvidia-container-toolkit # Smoke test: confirm the container can see the host GPU docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
Choosing the right base image tag matters here. NVIDIA's official CUDA images are published in three flavors: base images include only the CUDA runtime (cudart); runtime images build on that and add the CUDA math libraries and NCCL, with a cuDNN-included variant also published; devel images build on runtime and add headers and the nvcc compiler for building CUDA code, which makes them well suited to multi-stage builds.
# Illustrative multi-stage build: compile against 'devel', ship 'runtime' FROM nvidia/cuda:12.4.1-devel-ubuntu22.04 AS build WORKDIR /app COPY . . RUN make FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04 COPY --from=build /app/my-cuda-app /usr/local/bin/ ENV NVIDIA_DRIVER_CAPABILITIES=compute,utility CMD ["my-cuda-app"]
Why this two-stage pattern matters. The final image never carries the compiler toolchain or development headers the devel image bundles, which keeps the shipped image smaller and reduces the tools available to an attacker inside a compromised container — a routine hardening step once past initial prototyping.
💡 Trade-off. NVIDIA's supported CUDA image tags are periodically retired and rebuilt for CVE fixes, and the latest tag for these images has been deprecated by NVIDIA — pinning an explicit version tag, and re-validating on a schedule, is necessary rather than optional for a reproducible production build.
🎯 Use this when standing up the first GPU-enabled Docker host or restructuring a Dockerfile to separate build-time and run-time CUDA dependencies.
5. Implementation: GPU Scheduling in Kubernetes
Kid-friendly analogy: If Docker's --gpus flag is one person asking a single room's front desk for a key, Kubernetes' device plugin is a hotel-wide reservation system: every floor (node) reports how many rooms (GPUs) it has to a central desk (the scheduler), which then matches incoming guests (pods) to a floor that actually has a free room.
What it does. The NVIDIA device plugin for Kubernetes is a DaemonSet — a controller-managed pod running on every GPU node — that advertises each node's GPU count to the kubelet as an extended resource, nvidia.com/gpu, and monitors GPU health.
Why it is needed. Kubernetes' scheduler has no native concept of a GPU; it schedules based on declared resources. The device plugin framework is the extension point that teaches the scheduler a new resource type exists, how much of it each node has, and how to allocate it exclusively to the pod that requested it.
How it works, step by step:
- Each GPU node has the NVIDIA driver and NVIDIA Container Toolkit installed, with the container runtime configured to use the nvidia runtime.
- The device plugin DaemonSet runs on every GPU node and registers with the kubelet over a local gRPC socket, reporting the node's available GPU count.
- The kubelet publishes that count as allocatable capacity under nvidia.com/gpu for the node.
- A pod spec requests GPUs as a resource limit; the scheduler places the pod only on a node with enough unallocated nvidia.com/gpu capacity.
- At container start, the kubelet informs the device plugin which specific GPU device IDs were allocated, and the NVIDIA Container Toolkit injects exactly those devices into the container — the same mechanism used by plain Docker, now driven by Kubernetes' allocation decision.
# Illustrative Pod spec requesting one GPU
apiVersion: v1
kind: Pod
metadata:
name: gpu-inference-pod
spec:
containers:
- name: server
image: myorg/inference:1.0
resources:
limits:
nvidia.com/gpu: 1
What fails without it. Requesting nvidia.com/gpu on a cluster without the device plugin installed leaves the resource permanently unknown to the scheduler, and the pod stays in Pending with an "Insufficient nvidia.com/gpu" scheduling failure, even on nodes with physically idle GPUs.
At fleet scale, teams rarely install the driver, toolkit, and device plugin as three separate manual steps per node. NVIDIA's GPU Operator packages the driver, container toolkit, device plugin, DCGM metrics exporter, GPU Feature Discovery (which labels nodes with their GPU model and capabilities), and MIG management into one Helm-installed operator, so provisioning a new GPU node becomes a declarative, repeatable process rather than a manual runbook.
✅ Practical example. A cluster autoscaler adds a fresh GPU node; the GPU Operator's DaemonSets automatically install the driver and toolkit and register the node's GPUs with the device plugin, so pods with a pending nvidia.com/gpu request schedule onto it within minutes, with no engineer SSHing into the new node.
🎯 Use this when moving from manually configured single-node GPU Docker hosts to a self-service, autoscaling GPU cluster.
6. Sharing a GPU: MIG and Time-Slicing
Kid-friendly analogy: MIG is like a landlord physically subdividing one large apartment into several smaller, separately locked units, each with its own kitchen and bathroom — true walls, not just a schedule. Time-slicing is more like several roommates sharing one full apartment on a rotating schedule: nothing is physically divided, they just take turns, and if one roommate overstays their turn the others wait.
What MIG does. Multi-Instance GPU, available on Ampere-generation and later data-center GPUs such as the A100 and H100, uses spatial partitioning to carve one physical GPU into as many as seven independent GPU instances, each with its own dedicated memory, cache, and streaming multiprocessors, running simultaneously with hardware-level fault isolation.
Why it is needed. A single inference request for a small model rarely saturates a modern data-center GPU's full compute and memory capacity. Without partitioning, that idle capacity is simply wasted; MIG lets several independent workloads — different teams, different models, different tenants — each get a guaranteed, isolated slice with predictable quality of service instead of competing unpredictably for the whole device.
How it works, step by step:
- An administrator enables MIG mode on the physical GPU, which requires a GPU reset.
- The GPU is partitioned into instances of chosen sizes — for example, several equal 10 GB slices on an 80 GB A100, or a mix of smaller and larger slices to match different workloads.
- Each MIG instance behaves like a standalone GPU to the CUDA platform above it — no application-level code changes are required.
- In Kubernetes, the device plugin and GPU Feature Discovery advertise each MIG instance as its own schedulable unit, so pods can request a specific slice size.
- Because isolation is physical, a runaway or crashing workload on one instance cannot affect the memory, cache, or throughput of workloads on other instances of the same GPU.
What time-slicing does instead. Where MIG is unavailable — older GPU generations, or GPUs that do not support hardware partitioning — the NVIDIA Kubernetes device plugin supports time-slicing, first introduced starting with device plugin version 0.12.0, which advertises a single physical GPU to Kubernetes as multiple schedulable replicas. Multiple pods can then be scheduled onto what the cluster sees as separate GPU resources, while the underlying hardware time-shares the actual device among their contexts.
What fails without the right choice. Applying time-slicing to a latency-sensitive, SLA-bound inference service can introduce unpredictable tail latency, because one tenant's burst of compute-heavy requests can delay another tenant's requests on the same physical device — there is no hardware wall between them, only a scheduling illusion at the Kubernetes resource-accounting layer.
💡 Trade-off. MIG gives real isolation but fixes GPU memory into predefined slice sizes chosen at partition time — a workload needing more memory than its assigned slice cannot simply burst into a neighboring instance's memory, unlike time-slicing where the full GPU memory remains a shared (and shareable) pool.
🎯 Use this when GPU utilization metrics show significant idle capacity per device but the workloads sharing it need either guaranteed isolation (choose MIG) or simple best-effort sharing (choose time-slicing).
7. Enterprise Rollout: Fleet Management, Observability, and Incident Response
Once GPU containers move from one engineer's workstation to a shared production fleet, the driver/toolkit/CUDA stack becomes infrastructure that needs the same operational discipline as networking or storage.
Ownership and versioning. Assign a platform team ownership of the node-level stack — driver version, container toolkit version, and the Kubernetes device plugin version — separately from application teams' ownership of their container's CUDA Toolkit version. Document a supported CUDA compatibility matrix (which driver baselines support which CUDA major/minor versions under Minor Version Compatibility) so application teams can self-serve version bumps without filing a ticket every time.
CI gates. Build a GPU smoke test into CI for any image change touching the CUDA base layer — at minimum, an automated run of nvidia-smi and a trivial CUDA kernel launch inside a GPU-enabled CI runner — so a broken CUDA/driver combination is caught before a canary rollout, not during one.
Access controls. Scope who can install or modify cluster-wide GPU components (the device plugin, GPU Operator configuration, MIG partition layouts) tightly, since a misconfigured MIG layout or device plugin version can silently take an entire node's GPU capacity offline for every tenant scheduled there.
Budget controls. Track GPU utilization, not just GPU allocation — a pod holding a full nvidia.com/gpu request at 10% actual compute utilization is a strong signal to evaluate MIG partitioning or time-slicing for that workload class before provisioning additional GPU nodes.
Dashboards and alerts. NVIDIA's DCGM exporter, deployable as part of the GPU Operator, is the standard source for GPU-level metrics — utilization, memory usage, temperature, and ECC error counts — feeding Prometheus-based dashboards. Alert on driver/toolkit version drift across nodes (a common source of "works on some nodes, not others" incidents), rising ECC error counts (an early hardware-failure signal), and device-plugin pod crash-loops, which silently remove a node's entire GPU capacity from the schedulable pool.
Incident response. Maintain separate runbooks for "a node's GPUs disappeared from the schedulable pool" (usually a device plugin or driver issue — check the DaemonSet and driver logs first) versus "a specific pod can't initialize CUDA" (usually a CUDA/driver compatibility mismatch — check the image's CUDA version against that node's driver version first). Conflating these two failure classes routinely wastes on-call time chasing the wrong layer of the stack.
💡 Trade-off. Standardizing the entire fleet on one driver version simplifies support and CI, but it also means every application team's maximum usable CUDA version is capped by that one driver baseline until the whole fleet is upgraded together — a real coordination cost worth deciding on deliberately, not by default.
🎯 Use this checklist when a GPU stack that one team stood up manually is being handed off to a platform team as a shared, multi-tenant service.
8. Common Mistakes
Installing the full CUDA Toolkit on the host "just in case." Since only the driver is required on the host, installing the full Toolkit system-wide adds an unnecessary, hard-to-keep-current dependency that can drift out of sync with what containers actually need — and gives engineers a false signal that host-level CUDA version matters for container compatibility, when it does not.
Shipping a CUDA image without pinning the base tag. Building against an unpinned or "latest"-style tag — a tag NVIDIA has explicitly deprecated for these images — means a rebuild months later can silently pull a different CUDA version than originally tested, producing compatibility failures that are difficult to bisect after the fact.
Requesting nvidia.com/gpu without the device plugin installed and healthy. The pod does not fail with a clear "device plugin missing" error; it simply sits in Pending with an insufficient-resource scheduling failure, which looks identical to a genuinely full cluster and sends on-call engineers looking for capacity that was never the real problem.
Leaving every container with unrestricted GPU visibility. Defaulting to NVIDIA_VISIBLE_DEVICES=all for workloads that only need one GPU widens the blast radius of a compromised or buggy container to every GPU on the node, and undermines any scheduler-level isolation the platform is otherwise trying to enforce.
Applying time-slicing to latency-SLA workloads. Time-slicing has no hardware isolation, so a noisy neighbor sharing the same physical GPU can introduce tail-latency spikes that are extremely difficult to diagnose from the affected service's own metrics alone, since nothing in that service's own container appears to have changed.
Upgrading application CUDA versions without checking the fleet's driver baseline. A CUDA Toolkit bump that looks like a routine dependency update can silently exceed what Minor Version Compatibility or Forward Compatibility support on older nodes, turning a routine release into a partial fleet outage that only affects pods scheduled onto not-yet-upgraded nodes — making it intermittent and confusing to triage.
❓ FAQ
Do I need to install the CUDA Toolkit on my host machine to run GPU containers?
No. NVIDIA's own guidance is explicit that only the NVIDIA GPU driver needs to be installed on the host — the CUDA Toolkit ships inside the container image itself, which is what lets one host run containers built against different CUDA versions.
What's the difference between the base, runtime, and devel CUDA image tags?
Base images include only the CUDA runtime library; runtime images add on top of that the CUDA math libraries and NCCL, with a cuDNN variant also published; devel images build on runtime and add the compiler and development headers needed to build CUDA code, which is why they're typically used only in the build stage of a multi-stage Dockerfile.
Can a container use a newer CUDA version than my host driver officially supports?
Often yes, within documented limits: Minor Version Compatibility lets a container use a newer CUDA Toolkit minor release than the driver's native version as long as both share the same CUDA major version, and Forward Compatibility can extend this across major versions with an explicit compatibility package, subject to platform and GPU support.
Should I use MIG or time-slicing to share a GPU across multiple workloads?
MIG provides real hardware-level isolation with guaranteed memory and compute per instance, so it fits multi-tenant or SLA-bound workloads, but it requires Ampere-generation or later data-center GPUs. Time-slicing works on broader hardware but only simulates sharing at the scheduling layer, with no isolation guarantee against noisy neighbors.
Is this article eligible for a rich FAQ result in search engines?
This post includes FAQPage structured data matching the questions and answers shown above. Whether any given search engine displays a rich result is decided entirely by that search engine and cannot be guaranteed by this article or its publisher.
🔗 References & Further Reading
- NVIDIA Container Toolkit — Installation Guide
- NVIDIA Container Toolkit — Specialized Configurations with Docker (GPU enumeration, driver capabilities)
- NVIDIA — CUDA Compatibility Guide (Minor Version Compatibility, Forward Compatibility)
- NVIDIA — Official CUDA Docker Hub Repository (base/runtime/devel image tags)
- NVIDIA — Kubernetes Device Plugin documentation (nvidia.com/gpu resource, DaemonSet architecture)
- NVIDIA — Using MIG on NVIDIA DGX A100 (Multi-Instance GPU)
- NVIDIA — Multi-Instance GPU product overview and specifications
- NVIDIA Developer Blog — Getting the Most Out of the A100 GPU with Multi-Instance GPU
Product and project names above (NVIDIA, CUDA, Docker, Kubernetes, A100, H100) are trademarks of their respective owners and are used here only to identify the referenced technologies.
📝 Summary
- Foundations: the driver lives once on the host; the CUDA Toolkit ships inside each container; the NVIDIA Container Toolkit is the bridge between them.
- Mechanics: libnvidia-container and nvidia-container-runtime inject GPU devices and driver libraries into a container's namespace at start time, controlled by NVIDIA_VISIBLE_DEVICES and NVIDIA_DRIVER_CAPABILITIES.
- Compatibility: Minor Version Compatibility and Forward Compatibility define exactly when a container's CUDA version can differ from the host driver's native version, avoiding fleet-wide driver upgrades for every release.
- Docker example: install the driver and toolkit on the host, pin an explicit CUDA base image tag, and use base/runtime/devel deliberately across a multi-stage build.
- Kubernetes implementation: the NVIDIA device plugin advertises nvidia.com/gpu to the scheduler; the GPU Operator automates driver, toolkit, and plugin rollout across a fleet.
- Sharing a GPU: MIG gives hardware-isolated slices on supported GPUs; time-slicing gives software-scheduled sharing on broader hardware, with real latency trade-offs.
- Enterprise rollout: version-baseline ownership, GPU smoke tests in CI, DCGM-based observability, and separate runbooks for scheduler-level versus in-container CUDA failures keep the stack operable at scale.
- Common mistakes: unpinned base images, unrestricted GPU visibility, misapplied time-slicing, and driver/CUDA version drift across nodes are the recurring root causes behind GPU container incidents.
The CUDA and container stack looks intimidating mostly because three separately versioned components — driver, toolkit, and container runtime — are easy to mistake for one thing. Once that separation is clear, most GPU container failures become a matter of checking the right layer first. Happy shipping!
Comments
Post a Comment