Docker Model Runner (DMR) lets developers pull, run, and distribute AI models using familiar Docker commands and OCI-standard packaging — but, unusually for anything with "Docker" in the name, the model itself does not run inside a container; it runs as a native process on the host, with Docker handling distribution, tooling, and API access around it. 🧩
That architectural choice is the single most important thing to understand about Docker Model Runner, and it's easy to miss if you assume it works like every other "run this in Docker" tool. Assume the model is sandboxed the way a normal container is, and you'll misjudge its isolation properties. Assume it behaves like a typical detached background container, and its host-socket and internal-DNS access patterns will look confusing. Understanding why Docker built it this way — and where that trades performance for different security and portability properties — is the key to using it well. ⚙️
📑 In This Post
- Foundations: what Docker Model Runner is, and why it doesn't containerize the model
- Mechanics: host-based execution, sockets, and endpoints
- Model distribution: OCI artifacts and the docker model CLI
- Backends and GPU acceleration: llama.cpp, vLLM, and Diffusers
- Worked example: local development without a cloud API dependency
- Implementation: enabling, calling, and packaging a model
- Enterprise rollout: registries, capacity planning, and governance
- Common mistakes and why they hurt in production
- FAQ
- References & further reading
- Summary
🔀 Quick Comparison: Docker Model Runner vs. Containerized Model Servers
DMR and a fully containerized engine like Ollama-in-Docker or vLLM-in-Docker solve overlapping problems with different trade-offs.
| Aspect | Docker Model Runner | Fully containerized engine (e.g. vLLM/Ollama in Docker) |
|---|---|---|
| Where inference runs | Native host process (llama.cpp by default) | Inside the container, using the host GPU via passthrough |
| Isolation model | Model process shares the host directly; distribution and tooling are containerized concepts, the runtime is not | Model process runs inside the container's namespace and filesystem boundary |
| Model packaging | GGUF/Safetensors as OCI Artifacts, pulled like an image | Weights mounted from a volume or baked into a custom image |
| Best fit | Local developer inner loop, quick iteration, Docker Desktop workflows | Production serving with full control over isolation and orchestration |
🎯 Use this table to decide whether you want the fastest local iteration loop (DMR) or the strongest container isolation boundary (a fully containerized engine).
1. Foundations — What Docker Model Runner Is, and Why It Doesn't Containerize the Model
🧸 Kid-friendly analogy: Imagine a library (Docker) that's brilliant at cataloging, shelving, and lending out books (models) using one consistent system. But instead of making you read every book inside a soundproof booth (a container), the library hands you the book and lets you read it comfortably at your own desk (the host), where the lighting and chair are already set up just right.
Docker Model Runner is Docker's own tool, built into Docker Desktop and Docker Engine, for pulling, running, and serving AI models using Docker-style tooling and OCI-standard distribution. According to Docker's own documentation, it streamlines pulling, running, and serving large language models directly from Docker Hub, any OCI-compliant registry, or Hugging Face.
The distinctive design decision, explained directly in Docker's own announcement of the feature, is that the model itself does not run inside a container. Instead, running a model delegates the inference work to an inference server — llama.cpp by default — running as a native process on the host, specifically to avoid the performance overhead that containerization or virtualization would add to GPU-bound inference. Docker's containerization tooling still governs how the model is distributed, versioned, and invoked; it just isn't the runtime boundary around the inference process itself.
This matters for two very different reasons:
- Performance: a host-native process gets direct access to GPU acceleration (Apple's Metal API on Apple Silicon, CUDA on NVIDIA GPUs, or Vulkan on a broader range of GPUs) without the overhead containerizing the workload could add.
- Security assumptions: the strong per-tool isolation boundary that containerizing an MCP server or an inference engine provides does not automatically apply to the model process itself in this architecture — a distinction worth understanding clearly before assuming a "Dockerized" model carries the same isolation guarantees as a normal container.
2. Mechanics — Host-Based Execution, Sockets, and Endpoints
🧸 Kid-friendly analogy: Think of Docker's usual container boundary as a walled garden. Docker Model Runner's inference engine works more like a helpful neighbor who lives just outside the wall but has a dedicated gate (a socket or internal address) that anyone inside the garden can knock on whenever they need something.
What it does: when you run docker model run, Docker does not spin up a container for the model. It delegates the request to the inference server process, which loads the requested model into memory if it isn't already resident, and serves the response back through an OpenAI-compatible (and Ollama-compatible) API.
Why it is needed: without a host-native path, GPU-accelerated inference inside a Docker Desktop virtual machine layer (particularly on macOS, where Docker Desktop itself runs inside a lightweight VM) would add exactly the kind of overhead the whole feature exists to eliminate.
How it works, step by step (Apple Silicon as the original, most-documented case):
- Docker Desktop runs the llama.cpp-based inference server as a process on the host, not inside any container.
- From the host, the model runner is reachable through the Docker socket; from inside another container, it's reachable via a special internal address,
model-runner.docker.internal, without needing the model itself to be containerized. - You can also expose it over a specific TCP port for host processes (for example, code using an OpenAI SDK pointed directly at a local endpoint) with a command such as
docker desktop enable model-runner --tcp <port>. - Requests to the inference server follow the OpenAI-compatible chat completions format, so existing OpenAI-client code generally works by pointing it at the local endpoint instead of a cloud API.
What fails without understanding this: developers sometimes expect to find the running model as a container in docker ps output and are confused when it isn't there — it's running as part of Docker Desktop's own host-level process, not as a workload you'd manage with ordinary container commands.
💡 Trade-off: host-based execution buys real inference performance and simpler GPU access, at the cost of the strong per-workload isolation boundary a normal container provides. For local development and fast iteration this trade generally favors performance; for a production serving path where isolation, resource limits, and blast-radius containment matter more, a fully containerized engine (as covered in earlier posts on Ollama and vLLM in Docker) remains the more conservative choice.
3. Model Distribution — OCI Artifacts and the docker model CLI
🧸 Kid-friendly analogy: Packaging a model as an OCI Artifact is like putting a book in a standard-sized shipping box instead of an oddly shaped one — any shipping company (any container registry) can already handle that box without special accommodations.
What it does: Docker Model Runner packages model weights — typically GGUF format, with Safetensors also supported — as OCI Artifacts, the same standardized packaging format used for container images, and can pull or push them to Docker Hub or any OCI-compliant registry, as well as pull GGUF-format models from Hugging Face (which the Hugging Face integration packages as OCI Artifacts on the fly).
Why it is needed: before this kind of standardization, teams needed separate tooling for distributing model weights than for distributing application containers — different registries, different versioning conventions, different access-control mechanisms. Reusing the OCI standard means a model can be versioned, tagged, and access-controlled with the same infrastructure already in place for container images.
How it works, step by step:
- Pull a model by name and tag, much like a container image:
docker model pull ai/smollm2:360M-Q4_K_M - Run it interactively or send it a single prompt:
docker model run ai/smollm2:360M-Q4_K_M "Give me a fact about whales." - The model is cached locally, so subsequent runs don't require re-pulling it, similar to how Docker caches image layers.
- To distribute a custom or fine-tuned model, package it as an OCI Artifact and push it to any container registry your organization already uses — no separate model-hosting infrastructure required.
What fails without it: teams that keep model weights in ad hoc file shares or bespoke storage buckets end up maintaining a parallel set of versioning, access-control, and distribution practices alongside the ones they already have for containers — extra operational surface for no real benefit once a standardized packaging format exists.
4. Backends and GPU Acceleration — llama.cpp, vLLM, and Diffusers
Docker Model Runner's default inference backend is llama.cpp, which runs across the widest range of hardware — Apple Silicon via Metal, NVIDIA GPUs via CUDA, and, per Docker's own documentation, a broader set of GPUs (including AMD and Intel integrated GPUs) via Vulkan support. Docker's documentation also lists support for vLLM and Diffusers as additional inference engines, but specifically scoped to Linux systems with NVIDIA GPUs — these are not available on macOS or Windows in the same way llama.cpp is.
Why the backend choice matters: llama.cpp is the general-purpose, broadly-compatible default suited to quantized GGUF models on a developer's own machine. vLLM brings the continuous-batching and paged-KV-cache throughput techniques described in dedicated inference-optimization coverage, but its narrower platform support reflects that it's aimed at higher-throughput scenarios closer to production serving rather than single-user local development.
✅ Worked example: a developer on a Windows machine with an NVIDIA GPU can install the vLLM backend for Docker Model Runner specifically to test higher-throughput serving locally, while a teammate on an Apple Silicon Mac uses the default llama.cpp backend for the same models — the docker model commands stay the same, but the underlying engine and its platform requirements differ.
5. Worked Example — Local Development Without a Cloud API Dependency (hypothetical, for illustration)
Consider a hypothetical team building a feature that calls an LLM as part of their application, who want to iterate on prompts and application logic without paying per-token cloud API costs during development, or being blocked by network access in a restricted environment.
How it works, step by step:
- Developers pull a small, quantized model with
docker model pulland point their application's existing OpenAI-client code at the local endpoint instead of a cloud provider. - Because the API is OpenAI-compatible, most of the application's existing request and response handling code works unmodified.
- Prompt iteration happens entirely locally, with no per-request cost and no dependency on external network availability.
- Before deploying, the team switches the endpoint back to their production LLM provider or a fully containerized serving engine, treating the local model as a development-time stand-in rather than the final production dependency.
What this illustrates: Docker Model Runner's value here is squarely in the local developer inner loop — fast iteration, no cost, no network dependency — rather than as a drop-in replacement for a production-grade, fully containerized serving stack.
🎯 Use this approach when your goal is fast, cost-free local iteration on LLM-backed application logic, not production-scale serving.
6. Implementation — Enabling, Calling, and Packaging a Model
Enabling Docker Model Runner and calling it from a containerized application follows a small, consistent set of steps:
# Illustrative example — verify current flags against Docker's own documentation
# Enable Model Runner (host machine)
docker desktop enable model-runner
# Pull and run a small model directly
docker model pull ai/smollm2:360M-Q4_K_M
docker model run ai/smollm2:360M-Q4_K_M "Summarize this in one sentence."
# From inside another container on the same Docker network,
# the model is reachable at: http://model-runner.docker.internal/engines/llama.cpp/v1/chat/completions
Illustrative example only — command flags, endpoint paths, and defaults can change between Docker versions; confirm against current Docker documentation before use.
To package a custom or fine-tuned model for distribution, the workflow mirrors building and pushing a container image: package the GGUF or Safetensors weights as an OCI Artifact and push that artifact to whichever container registry your organization already trusts and manages access for, rather than standing up separate model-hosting infrastructure.
7. Enterprise Rollout — Registries, Capacity Planning, and Governance
Because Docker Model Runner is oriented primarily toward the local developer experience, "enterprise rollout" here mostly means governing how teams use it consistently and safely, rather than operating it as a shared production service:
- Registry governance for custom models. If your organization publishes fine-tuned or internal models as OCI Artifacts, apply the same access-control and scanning practices to that registry namespace as you would to application container images.
- Capacity awareness on developer machines. Docker's own documentation notes that Model Runner does not currently include safeguards to prevent launching a model too large for the host machine's available resources, which can cause severe slowdowns; teams should set expectations about which model sizes are appropriate for typical developer hardware.
- Consistent local-vs-production boundaries. Make it an explicit team convention that Docker Model Runner is a local development and iteration tool, not a substitute for the production serving stack's isolation and scaling properties — this avoids anyone accidentally treating a host-native local process as if it had the same operational guarantees as a properly containerized production service.
- Platform-specific backend awareness. Since vLLM and Diffusers backend support is currently scoped to Linux with NVIDIA GPUs, don't assume every developer's machine can reproduce the same backend behavior; document which platforms support which backends to avoid confusing, hardware-specific bug reports.
- CI considerations. If extending Model Runner-style local inference into CI environments, verify GPU or backend availability in that environment explicitly, since CI runners often lack the GPU hardware a developer's own machine has.
8. Common Mistakes
- Assuming the model runs inside a container. Looking for it in
docker ps, or assuming it carries the same isolation boundary as a normal container, leads to incorrect assumptions about both its behavior and its security properties. - Treating it as a production serving solution by default. Docker Model Runner is designed around the local developer experience; using it as-is for production traffic skips the isolation, scaling, and orchestration properties a fully containerized serving engine provides.
- Pulling a model too large for the host without checking first. Since there's currently no built-in safeguard against this, an oversized model can make a development machine unusable rather than simply failing to load.
- Expecting vLLM or Diffusers backend behavior on macOS or Windows without an NVIDIA GPU. These backends are documented as scoped to Linux with NVIDIA GPUs; expecting identical behavior across all platforms leads to confusing, hardware-specific failures.
- Confusing the Docker socket path with the TCP endpoint. Code expecting a host-accessible TCP port needs Model Runner explicitly enabled with a TCP flag; code running inside a container should instead use the internal Docker-network address — mixing these up produces connection failures that look like the service isn't running at all.
- Not distinguishing custom model registries from public ones in access policy. Publishing an internal fine-tuned model as an OCI Artifact to a registry without applying the same access controls used for application images can expose proprietary model weights more broadly than intended.
❓ FAQ
Does Docker Model Runner run models inside a Docker container?
No. According to Docker's own documentation and announcement, the inference engine runs as a native process on the host — typically llama.cpp — specifically to avoid the performance overhead containerization would add to GPU-accelerated inference.
How do models get distributed with Docker Model Runner?
Models are packaged as OCI Artifacts — the same standardized format used for container images — and can be pulled from or pushed to Docker Hub or any OCI-compliant registry, or pulled from Hugging Face in GGUF format.
Can I run vLLM as a backend with Docker Model Runner?
Yes, but currently only on Linux systems with NVIDIA GPUs, per Docker's own documentation. On macOS and Windows, or on non-NVIDIA hardware, the default llama.cpp backend is what's available.
Is Docker Model Runner meant to replace a production inference server?
It's designed primarily for the local developer experience — fast iteration, no per-request cost, no cloud dependency — rather than as a production serving solution with the isolation and orchestration properties a fully containerized engine provides.
What happens if I try to run a model too large for my machine?
Docker's own documentation notes there is currently no built-in safeguard against this; attempting it can cause severe slowdowns or make the system temporarily unusable, so checking a model's resource requirements against your hardware beforehand is worthwhile.
🔗 References & Further Reading
- Docker official documentation — "Docker Model Runner"
- Docker official blog — announcement of Docker Model Runner and its host-based execution architecture
"Docker," "Docker Model Runner," and related marks are trademarks of Docker, Inc.
📝 Summary
- Docker Model Runner brings Docker-style tooling and OCI-standard distribution to AI models, but the model itself runs as a native host process, not inside a container.
- That host-based execution trades some of the isolation a normal container provides for direct GPU access and less inference overhead.
- Models are packaged as OCI Artifacts, letting existing container-registry infrastructure, versioning, and access control cover them without separate tooling.
- The default llama.cpp backend runs broadly across platforms; vLLM and Diffusers backends are currently scoped to Linux with NVIDIA GPUs.
- A worked example shows its clearest fit: fast, cost-free local development iteration rather than production-scale serving.
- Enterprise use mostly means governance — registry access control for custom models, capacity awareness on developer machines, and a clear team convention distinguishing local iteration from production serving.
- Most confusion traces back to one root cause: assuming the model is containerized the way everything else in Docker usually is.
Hopefully this gives you a clear picture of what makes Docker Model Runner different from a typical containerized model server — and where that difference should shape how you use it. 🙌
Comments
Post a Comment