LLM Inference Optimization with Docker: Optimize GPU Performance, Throughput, and Latency
LLM inference optimization in Docker means getting continuous batching, paged KV-cache memory, quantization, and speculative decoding to actually work inside a container boundary — which depends as much on how you configure the container (GPU passthrough, shared memory, volumes, resource limits) as on the inference engine itself. 🐳
A serving engine like vLLM or TensorRT-LLM already implements these optimizations in software. But a container that hides the GPU from the process, caps shared memory too low for tensor-parallel inference, or re-downloads a 15GB model on every restart will quietly throw away most of that engineering — and the failure often looks like "the model is slow" when the real cause is one missing Docker flag. 🚨
📑 In This Post
- Foundations: what the container boundary changes about inference
- GPU passthrough, shared memory, and continuous batching
- Model volumes, image size, and quantized builds
- Speculative decoding as a multi-container pattern
- Worked example: containerizing a vLLM-served chat API
- Implementation: Dockerfile, Compose, and health checks
- Enterprise rollout: Kubernetes GPU scheduling and governance
- Common mistakes and why they hurt in production
- FAQ
- References & further reading
- Summary
🔀 Quick Comparison: Container-Level Concerns per Optimization
Each inference optimization technique has its own Docker-specific dependency — miss it, and the technique either silently degrades or fails outright.
| Technique | Container-level dependency | Symptom if misconfigured |
|---|---|---|
| Continuous batching / paged KV cache | GPU visible via --gpus/NVIDIA runtime; enough GPU memory headroom for the container's configured concurrency |
Requests queue or fail under load that the GPU itself could handle |
| Tensor-parallel serving | Sufficient container shared memory (--ipc=host or --shm-size) |
Crashes or hangs specific to multi-GPU tensor parallelism |
| Quantized model weights | Build-time choice of base image/kernels supporting the quantization format | Model loads but ignores GPU acceleration for that precision, or fails to load |
| Model weight persistence | Named volume mounted for the model cache directory | Every container recreation re-downloads multi-gigabyte weights |
🎯 Use this table as a pre-flight checklist before assuming a "slow" containerized model is an engine problem rather than a container configuration gap.
1. Foundations — What the Container Boundary Changes About Inference
🧸 Kid-friendly analogy: A container is like a diver's wetsuit — it's meant to fit closely and keep everything self-contained, but if you forget to cut holes for the diver's air hose (the GPU) or their weight belt pouch (shared memory), the wetsuit doesn't make the diver better at swimming; it just gets in the way.
A Docker container does not change the fundamental physics of inference: decode is still sequential and memory-bandwidth bound, and batching, paging, quantization, and speculative decoding are still what make it fast at scale, exactly as they would be on bare metal. What the container adds is a boundary that everything — the GPU driver, shared memory, model weights, and the network port — has to explicitly cross. Every inference optimization technique has a corresponding "does the container actually expose this correctly" question, and that question is usually where production incidents start.
Concretely, three container-level decisions determine whether an inference-optimized engine performs as designed once containerized:
- Does the container see the GPU? Without an explicit GPU-aware runtime, the process inside falls back to CPU, and every optimization technique above becomes moot.
- Does the container have enough shared memory? Multi-process, multi-GPU serving patterns (tensor parallelism) depend on inter-process shared memory that Docker restricts by default.
- Where do model weights and cache live? Baking weights into the image or leaving them in the container's writable layer defeats the point of treating the image as a disposable, rebuildable artifact.
2. GPU Passthrough, Shared Memory, and Continuous Batching
🧸 Kid-friendly analogy: Continuous batching inside a container is like a shared kitchen where several cooks (GPU processes) pass ingredients back and forth on a common counter (shared memory). If the counter is too small, the cooks can't pass things efficiently no matter how good the kitchen equipment is.
What it does: to actually run continuous batching and a paged KV cache — the core throughput techniques covered by vLLM and TensorRT-LLM — the container needs direct GPU access, and if the engine runs tensor-parallel across multiple GPUs, it needs enough shared memory for the underlying process communication.
Why it is needed: the NVIDIA driver lives on the host, not inside the container image, so GPU access has to be explicitly granted per container. Separately, Docker's default shared-memory allocation (a small /dev/shm) is too small for the inter-process communication that libraries like PyTorch use during multi-GPU tensor-parallel inference, which vLLM's own deployment documentation calls out explicitly.
How it works, step by step (vLLM's official Docker deployment guidance as a concrete, verifiable example):
- Install and configure the NVIDIA Container Toolkit on the host, as with any GPU-accelerated container.
- Run the container with the NVIDIA runtime and GPU devices requested:
docker run --runtime nvidia --gpus all --ipc=host -p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:latest --model <model-id> - Use either
--ipc=hostor an explicit--shm-sizeflag so the container's shared memory matches what tensor-parallel inference needs, rather than Docker's small default. - Once running, the engine's own scheduler handles continuous batching and paged KV-cache allocation inside the GPU memory it was granted — the container's job was simply to not block that access.
What fails without it: a container started without a GPU-aware runtime will often start successfully and serve requests on CPU — no crash, no obvious error, just dramatically higher latency per token. A container with GPU access but default shared memory can crash or hang specifically when tensor parallelism is enabled, which is a confusing failure mode because single-GPU runs of the same image work fine.
Best practices: verify GPU visibility inside the container explicitly (rather than inferring it from response speed) before troubleshooting engine-level configuration; size shared memory deliberately for multi-GPU deployments instead of defaulting to --ipc=host everywhere, since that flag also relaxes container isolation from the host.
💡 Trade-off: --ipc=host is the simplest fix for shared-memory limits, but it shares the host's IPC namespace with the container, which is a meaningfully looser isolation boundary than a scoped --shm-size value. In multi-tenant or security-sensitive environments, sizing --shm-size explicitly is usually the more defensible default.
3. Model Volumes, Image Size, and Quantized Builds
🧸 Kid-friendly analogy: Baking model weights into a Docker image is like sewing your groceries into your jacket instead of putting them in the fridge — the jacket becomes huge and heavy, and every time you get a new jacket (image update) you have to buy groceries all over again.
What it does: keeping model weights in a mounted volume (as with the official vLLM image mounting a Hugging Face cache directory, or Ollama mounting /root/.ollama) decouples the multi-gigabyte model artifact from the much smaller, more frequently updated engine image.
Why it is needed: engine images get updated far more often than any single model version is retrained. If weights live inside the image, every engine patch forces a full model re-download and a much larger image to store, transfer, and scan.
How quantization intersects with this: a quantized model (INT4/INT8 weight-only, or FP8 where hardware supports it) is itself a smaller artifact to store and pull, which directly reduces both volume size and cold-start pull time — a second, often overlooked, container-level benefit of quantization beyond GPU memory savings.
What fails without it: teams that bake a specific model into a custom image for a "self-contained" deployment often end up with multi-gigabyte images that are slow to build, slow to push to a registry, and slow to pull onto new nodes — undermining exactly the fast-scaling behavior containers are supposed to enable.
Best practices: mount model weights from a named volume, network storage, or an object-storage-backed init step rather than baking them into the image, unless you have a specific, deliberate reason (such as an air-gapped or serverless-style deployment) to accept the trade-off; if you do bake a model in, keep that step in its own build stage so an engine code change doesn't force a redundant model re-download during the build.
4. Speculative Decoding as a Multi-Container Pattern
🧸 Kid-friendly analogy: Running a draft model and a target model together is like having two chefs share one kitchen — a fast prep cook (the draft model) and a head chef who checks the prep cook's work (the target model). They need to work in the same kitchen, at the same time, which is a container resource-planning question as much as a modeling one.
What it does: speculative decoding — using a small draft model to propose tokens that a larger target model verifies in one pass, as introduced in the original 2023 papers by Leviathan et al. (Google Research) and Chen et al. (DeepMind) — requires both models to be loaded and available to the same serving process at inference time.
Why it is a container-level concern: the draft and target model together need to fit in the GPU memory the container was granted, on top of whatever KV-cache memory continuous batching also needs. This is a capacity-planning decision made at the container/pod resource-limit level, not something the engine can invent extra memory for on its own.
How it works in practice: most serving engines that support speculative decoding load both models within a single container process (rather than as separate containers exchanging predictions over the network, which would add latency defeating the point of the optimization) — so the relevant Docker-level decisions are sizing the container's GPU memory allocation for both models together, and, in Kubernetes, requesting GPU resources that reflect that combined footprint rather than the target model alone.
✅ Worked example: a team enabling speculative decoding in their containerized serving setup found their GPU ran out of memory under production traffic, even though it had run fine in testing with fewer concurrent requests. The cause was that their Kubernetes GPU resource request had been sized for the target model alone; it needed to account for the draft model's footprint plus the KV cache needed for real concurrency, not just a single test request.
5. Worked Example — Containerizing a vLLM-Served Chat API (hypothetical, for illustration)
Consider a hypothetical team moving an internal chat assistant from a hand-rolled Flask container running a general-purpose deep learning library to vLLM's official Docker image.
How the migration worked, step by step:
- They replaced their custom Dockerfile with the official
vllm/vllm-openaiimage, keeping the same model weights. - They mounted their existing Hugging Face model cache as a volume instead of copying weights into the image, so the switch didn't require re-downloading anything.
- They added
--gpus alland the NVIDIA runtime, which their original container had never explicitly requested — meaning their "GPU-accelerated" service had, in fact, been running partly on CPU without anyone noticing until they compared logs. - They set an explicit
--shm-sizeafter their first multi-GPU test run hung, tracing the hang back to Docker's default shared-memory limit. - They benchmarked the new container against their old one on the same representative internal question set before fully cutting traffic over.
What this illustrates: most of the migration effort was container configuration — GPU passthrough, volumes, shared memory — not inference-engine tuning. That is typical: the engine usually does the hard optimization work already; the container's job is to not accidentally block it.
🎯 Use this approach when adopting a purpose-built inference engine to replace an ad hoc containerized model server.
6. Implementation — Dockerfile, Compose, and Health Checks
A Compose file makes the GPU, volume, and shared-memory decisions explicit and version-controlled:
# Illustrative example — adapt image tag, model, and resource values before use
services:
vllm:
image: vllm/vllm-openai:latest
ports:
- "8000:8000"
volumes:
- hf_cache:/root/.cache/huggingface
ipc: host # or set a specific shm_size instead, for tighter isolation
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
command: >
--model example-org/example-7b-instruct
--max-model-len 4096
--gpu-memory-utilization 0.90
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
volumes:
hf_cache:
Illustrative example only — verify current flags and defaults against your chosen engine's own documentation before use.
Two implementation details matter more than they first appear:
- Health checks should probe the inference endpoint, not just process liveness. A container can be "up" while the engine is still loading a multi-gigabyte model into GPU memory; a liveness probe that only checks the process, not the API, will route traffic to a pod that isn't ready to serve.
- GPU memory utilization settings interact with container/pod memory limits. A GPU memory fraction configured inside the engine (to leave headroom for the KV cache) should be considered alongside any Kubernetes resource limits set at the pod level, not treated as an independent setting.
7. Enterprise Rollout — Kubernetes GPU Scheduling and Governance
Moving a containerized, optimization-aware inference server to Kubernetes adds scheduling and governance concerns on top of the single-container ones:
- GPU-aware scheduling. Pods requesting GPU resources need the cluster's GPU device plugin correctly configured so the scheduler places them only on nodes with available GPU capacity — an inference pod scheduled onto a GPU-less node fails the same way an under-provisioned Docker container does, just at a different layer.
- Resource requests sized for the real workload. As with the speculative-decoding example above, GPU memory requests should reflect the combined footprint of the model(s), KV cache at target concurrency, and any draft model — not just the base model's size.
- Image provenance and scanning. Treat inference-engine images the same as any other production container: pull from a private, scanned registry, pin digests rather than floating tags, and gate promotion on a vulnerability scan passing.
- Canary rollout for engine or quantization changes. Route a small percentage of traffic to a new engine version, quantization scheme, or draft model configuration, and compare latency, error rate, and quality against the current deployment before a full rollout — the container orchestration layer (traffic splitting, rollback) is what makes this practical at scale.
- Autoscaling tied to the right signal. Because inference is memory-bound rather than CPU-bound, autoscaling on CPU utilization alone often under- or over-reacts; queue depth, GPU utilization, and latency percentiles are more representative signals for scaling inference pods.
- Access control and change ownership. Changes to engine flags, quantization configuration, or GPU resource requests are production changes with real cost and quality impact, and should go through the same review process as any other deployment change, not be tunable directly by anyone with cluster access.
- Cost dashboards per environment. GPU time is the dominant cost driver; track it separately for dev, staging, and production so a forgotten always-on GPU pod in a non-production namespace doesn't become an invisible ongoing cost.
8. Common Mistakes
- Assuming a container has GPU access because it starts successfully. A container missing the NVIDIA runtime or
--gpusflag still runs — it just falls back to CPU, silently discarding every GPU-dependent optimization technique. - Leaving Docker's default shared memory in place for multi-GPU serving. Tensor-parallel inference depends on inter-process shared memory that Docker's default
/dev/shmsize does not provide, producing hangs or crashes that look unrelated to the actual cause. - Baking model weights into the image. This bloats the image, slows every build and pull, and forces a full re-download on every engine version bump — undermining the separation between a frequently-updated engine and a stable model artifact that containers are meant to provide.
- Sizing GPU resource requests for the model alone. Ignoring KV-cache memory at real concurrency, or a draft model's footprint with speculative decoding, leads to out-of-memory failures that only appear under production load, not in light testing.
- Using process liveness as a readiness signal. A pod can be "alive" while still loading a model into GPU memory; routing traffic to it before the inference endpoint itself is ready produces avoidable request failures during every rollout or restart.
- Scaling containerized inference on CPU utilization. Since decode-phase inference is memory-bandwidth bound rather than CPU bound, CPU-based autoscaling policies often fail to reflect real load, leaving GPUs either overloaded or idle relative to what the autoscaler believes is happening.
❓ FAQ
Why does my containerized model run slower than the same model outside Docker?
The most common cause is that the container never actually got GPU access — it starts and serves requests, just on the CPU. Confirm GPU visibility inside the container explicitly rather than assuming a flag worked.
Do I need --ipc=host to run an inference container?
Only if you're running multi-GPU tensor-parallel inference, where shared memory between processes matters. Single-GPU deployments typically don't need it; when you do need more shared memory, a scoped --shm-size value is a tighter alternative than sharing the host's IPC namespace.
Should I bake model weights into my inference image?
Generally no — mounting weights from a volume keeps the image small and lets you update the engine without re-downloading the model. Baking weights in can make sense for specific cases like air-gapped or serverless-style deployments where a self-contained image is the priority.
How should I size Kubernetes GPU resource requests for an inference pod?
Account for the model's own memory footprint, KV-cache memory at your expected concurrency, and any additional models such as a speculative-decoding draft model — not just the base model size alone.
Is a container health check enough to know an inference pod is ready?
Process liveness alone is not enough, since a container can be running while its model is still loading into GPU memory. A readiness check against the actual inference endpoint is a more reliable signal before routing traffic.
🔗 References & Further Reading
- vLLM official documentation — "Deploying with Docker"
- NVIDIA — NVIDIA Container Toolkit repository and documentation
- Kwon, W., Li, Z., Zhuang, S., et al. — "Efficient Memory Management for Large Language Model Serving with PagedAttention," ACM SIGOPS 29th Symposium on Operating Systems Principles (SOSP), 2023
- Leviathan, Y., Kalman, M., Matias, Y. — "Fast Inference from Transformers via Speculative Decoding," arXiv:2211.17192 / ICML, 2023
Product and project names mentioned (vLLM, Docker, NVIDIA, Kubernetes) are trademarks or projects of their respective owners.
📝 Summary
- Containerizing an inference engine doesn't add new optimization techniques — it adds a boundary those techniques have to correctly cross: GPU access, shared memory, and model storage.
- GPU passthrough and shared-memory sizing determine whether continuous batching and tensor-parallel serving actually get GPU acceleration inside the container.
- Model weights belong in a mounted volume, not baked into the image, to keep engine updates decoupled from multi-gigabyte model artifacts.
- Speculative decoding adds a second model's memory footprint that container and Kubernetes resource requests must explicitly account for.
- A worked migration example shows most containerization effort going into configuration — GPU flags, volumes, shared memory — rather than engine tuning.
- Kubernetes rollout adds GPU-aware scheduling, image provenance, canary releases, and load-appropriate autoscaling signals on top of single-container concerns.
- Most real incidents trace back to a short list of avoidable container misconfigurations: missing GPU runtime, default shared memory, baked-in weights, and liveness-only health checks.
Hopefully this gives you a clear map of where inference optimization meets the container boundary — and which Docker or Kubernetes setting to check first when something that should be fast isn't. 🙌
Comments
Post a Comment