Skip to main content

Ollama with Docker: Run Local LLMs in Containers

Calculating read time…

Ollama with Docker packages a local large-language-model server into a portable container so anyone with Docker can pull, run, and serve open-weight models (Llama, Gemma, Mistral, and others) through a single REST API on port 11434, without hand-installing drivers or runtimes on every machine. 📦

For a hobbyist, that convenience is nice to have. For a team shipping a production RAG pipeline, an internal coding assistant, or an agentic MCP workflow, it is the difference between "works on my laptop" and a service that survives a redeploy, a GPU driver upgrade, or a Friday-afternoon incident. Get the container boundary wrong — the wrong volume, a missing GPU runtime, an unpinned image tag — and you inherit slow cold starts, silently-CPU-bound inference, or a model store that vanishes on the next docker rm. 🚨

Diagram showing a client application sending HTTP requests on port 11434 to an Ollama server process inside a Docker container, which reads model weights from a mounted named volume and optionally uses a GPU through the NVIDIA Container Toolkit or ROCm device mapping.

Figure 1 — Original composition: request path from a client into a containerized Ollama server, its model volume, and its GPU passthrough path.

🔀 Quick Comparison: Ways to Run an Open-Weight LLM Locally

Ollama in Docker is one of several ways to self-host an open-weight model. The right choice depends on how much serving sophistication you need versus how quickly you need something running.

Approach Setup effort GPU support Best fit
Ollama native install Lowest Direct on macOS/Windows/Linux Single developer laptop
Ollama in Docker Low NVIDIA via Container Toolkit; AMD via ROCm image Reproducible dev/test, small internal services
vLLM in Docker Medium NVIDIA, tuned for throughput High-concurrency inference APIs
Triton / TensorRT-LLM High NVIDIA, deeply optimized Large-scale, latency-sensitive production serving

🎯 Use this when you want a single, low-friction way to try, demo, or lightly serve open-weight models without standing up a dedicated inference stack.

1. Foundations — What Ollama Is, and Why Docker Changes How You Run It

🧸 Kid-friendly analogy: Think of a large language model as a very heavy, very smart encyclopedia. Ollama is the librarian who knows how to open that encyclopedia, find the right page, and read the answer out loud when you ask a question. Docker is the sealed moving box that carries the librarian, their tools, and their reading glasses to any new room — so the librarian works the same way in every room, without you refitting the room first.

Technically: Ollama is a model-serving daemon. It exposes a REST API (default port 11434), manages a local cache of model weights in the GGUF format, and handles loading a model into memory, running inference, and evicting it when idle. It ships both a CLI and an HTTP API, so the same binary can be driven interactively or called programmatically from an application.

Docker changes what "running Ollama" means in three concrete ways:

  • Environment isolation. The container bundles Ollama's exact runtime dependencies, so the behavior on a developer's laptop, a CI runner, and a production host is defined by one image tag rather than by whatever happens to be installed on each machine.
  • Portability of the serving unit. The same image can be scheduled by Docker Compose on a single host or by Kubernetes across a cluster, without changing how the application talks to it — it is still an HTTP client hitting port 11434.
  • Separation of model data from the runtime. Model weights live in a mounted volume rather than inside the image, so you can update the Ollama binary, roll back a bad release, or move the container to new hardware without re-downloading multi-gigabyte model files every time.

Ollama has shipped an official, Docker-sponsored open-source image on Docker Hub as ollama/ollama since October 2023, with separate guidance for CPU-only and NVIDIA-GPU runs on Linux, and a note that Mac users should generally run Ollama natively rather than inside Docker Desktop because Docker Desktop on macOS does not pass GPUs through to containers.

2. Mechanics — Ports, Volumes, and the Model Lifecycle

🧸 Kid-friendly analogy: Imagine a food truck (the container) parked on a street. The serving window (the port) is where customers place orders. The pantry in the back (the volume) is where the ingredients are kept — and smart truck owners keep the pantry in a storage unit that survives even if the truck itself gets swapped for a newer model.

What it does: when you start the official image, the container's entrypoint runs ollama serve, which opens the REST API on port 11434 inside the container. You publish that port to the host, mount a named volume at /root/.ollama for model blobs and manifests, and then pull or run models against the running server.

Why it is needed: containers are meant to be disposable. If model weights were baked into a container's writable layer instead of a volume, deleting or recreating the container would silently delete every downloaded model, and every image rebuild would need to re-fetch gigabytes of weights.

How it works, step by step:

  1. Start the container with a published port and a named volume:
    docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
  2. Ollama's entrypoint launches ollama serve and begins listening on port 11434 inside the container.
  3. You pull a model by name, either by executing the CLI inside the running container or by calling the REST pull endpoint from a client:
    docker exec -it ollama ollama run llama3
  4. The server writes the model's layered blobs into the mounted volume, keyed by content hash, and records a manifest describing the model's tag.
  5. On each inference request, the server loads the requested model into memory (CPU RAM, or GPU memory when GPU acceleration is configured), keeps it resident for a short idle window to serve follow-up requests cheaply, and evicts it under memory pressure or after the idle timeout.

What fails without it: skip the volume mount and every container recreation re-downloads every model from scratch, which is slow and, on metered or shared networks, expensive. Skip the port publish and no client outside the container's network namespace can reach the API at all — a common first-timer failure that looks like "Ollama is hanging" but is actually "nothing is listening where the client is looking."

Best practices: always use a named volume (not an anonymous one) so it survives docker rm; pin the image to a specific tag rather than floating on latest in anything beyond local experimentation; and put the API behind a reverse proxy or network policy in any environment where the host is reachable from outside your own machine, since the default API has no built-in authentication layer of its own.

✅ Worked example: a two-person team building an internal documentation search tool ran ollama/ollama on a shared Linux VM with a named volume and a fixed image tag. When they later needed to move the service to a bigger VM, they only had to reattach the same volume — no model re-downloads, no re-configuration of clients, because the API contract on port 11434 never changed.

3. GPU Acceleration — NVIDIA Container Toolkit and ROCm

🧸 Kid-friendly analogy: A container is like a guest staying in your house who brought their own suitcase of clothes (the software inside the image) but did not bring their own car. If they need to drive somewhere fast (do GPU math), they have to borrow your car (the host's GPU and driver) — and there has to be a house rule (the NVIDIA Container Toolkit) that lets a guest actually use the garage.

A container does not ship its own GPU driver. The driver stays on the host; what a GPU-aware container runtime does is expose the host's GPU devices and driver libraries into the container's namespace so the process inside can talk to the GPU directly.

What it does: the NVIDIA Container Toolkit registers a GPU-aware runtime with the Docker daemon so that --gpus=all (or a specific device list) becomes a meaningful flag instead of a no-op.

Why it is needed: without it, Ollama inside the container falls back to CPU inference. The container will still run and still answer requests — it will just be dramatically slower per token, which is a common source of "Ollama in Docker is slow" reports that are actually "Ollama in Docker never got the GPU."

How it works, step by step, on Linux with an NVIDIA GPU:

  1. Install the NVIDIA driver on the host (not inside the container).
  2. Install the NVIDIA Container Toolkit package for your distribution.
  3. Register the toolkit as a Docker runtime:
    sudo nvidia-ctk runtime configure --runtime=docker
    sudo systemctl restart docker
  4. Start the container requesting GPU devices:
    docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
  5. Verify from inside the container that the GPU is visible before assuming acceleration is active, rather than inferring it from response speed alone.

For AMD GPUs, Ollama publishes a separate ollama/ollama:rocm image tag, started with the ROCm device files mapped in instead of the NVIDIA runtime flag:

docker run -d --device /dev/kfd --device /dev/dri -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama:rocm

💡 Trade-off: GPU passthrough ties your container to a specific host's driver stack and hardware. A container built and tested on a machine with GPU acceleration will still start on a GPU-less machine — it just silently reverts to CPU inference, so latency regressions from a bad scheduling decision (landing on a node without GPUs) can be easy to miss without explicit checks.

4. Worked Example — A Containerized RAG Assistant (hypothetical, for illustration)

Consider a hypothetical internal tool: an engineering team wants a chat assistant that answers questions against their own runbooks, with nothing leaving the company network.

What it does: a small multi-container stack — an Ollama container for the language model, a vector-capable database container for storing embedded runbook chunks, and a thin API container that retrieves relevant chunks and forwards a prompt to Ollama's REST API.

Why it is needed: keeping generation, retrieval, and the application layer as separate containers lets each be scaled, updated, and restarted independently — a slow embedding job does not need to take the chat API down, and upgrading the model server does not require touching the database.

How it works, step by step:

  1. A user question hits the API container.
  2. The API container embeds the question and queries the vector store for the closest runbook chunks.
  3. The API container builds a prompt containing those chunks and calls Ollama's /api/generate endpoint over the internal Docker network.
  4. Ollama loads the requested model (if not already resident) and streams the response back to the API container.
  5. The API container returns the answer to the user, optionally including which runbook chunks were used.

What fails without careful design: if the API container calls Ollama by its published host port instead of the internal container network, the setup becomes fragile outside the exact machine it was built on. If retrieval quality is never evaluated separately from generation quality, teams often "fix" hallucinated answers by tweaking the prompt when the real defect is in retrieval.

Best practices: put all three containers on a user-defined Docker network and address Ollama by its service name; keep the vector store's data on its own named volume; and evaluate retrieval (did it find the right chunks) separately from generation (did the model use them correctly), since conflating the two makes failures much harder to diagnose.

🎯 Use this when you are prototyping or lightly running an internal, retrieval-augmented assistant on infrastructure you already control.

5. Implementation — Compose, Dockerfiles, and Image Hygiene

A minimal Compose file for the RAG-style stack above illustrates the network and volume relationships without any vendor-specific configuration:

services:
  ollama:
    image: ollama/ollama:0.x.y   # example tag — pin a real, tested version
    volumes:
      - ollama_models:/root/.ollama
    networks: [assistant_net]
    # add: deploy.resources.reservations.devices for GPU access

  api:
    build: ./api
    depends_on: [ollama]
    environment:
      - OLLAMA_HOST=http://ollama:11434
    networks: [assistant_net]

volumes:
  ollama_models:

networks:
  assistant_net:

Illustrative example only — adapt image tags, GPU reservation syntax, and environment variables to your verified setup before use.

If you build a custom image on top of Ollama — for example, to pre-bake a specific model into the image for a serverless-style deployment — treat that image like any other production ML artifact: pin a base image digest, scan it for known vulnerabilities before pushing to a private registry, and keep the model-baking step in its own build stage so a source-code change does not force a full model re-download.

6. Enterprise Rollout — Governance, Scaling, and Observability

Moving from "it runs on a VM" to "it is an owned production service" changes what you need to control:

  • Ownership and access control. Decide who can push new image tags, who can trigger a model change in production, and put the Ollama API behind network policy or an authenticating proxy rather than exposing port 11434 directly to untrusted networks.
  • Image and model versioning. Pin both the Ollama image tag and the specific model tag used in production, and record which combination was validated together — an unannounced model update behaves like a silent behavioral change to anyone downstream.
  • Scaling beyond one host. A single Ollama container serves one host's worth of GPU capacity. Scaling further means running multiple Ollama instances behind a load balancer or moving to a Kubernetes Deployment with GPU-aware scheduling, at which point the same container image and API contract carry over — only the orchestration changes.
  • Observability. Track request latency, tokens per second, GPU memory utilization, and model load/eviction events. A model being evicted and reloaded on every request is a common, quiet cause of latency spikes that plain uptime monitoring will not surface.
  • Evaluation and regression testing. Before promoting a new model tag, run it against a fixed, representative set of prompts and compare outputs (and, for RAG use cases, retrieval-plus-generation quality together) against the currently deployed version, rather than judging a new model purely on a handful of manual spot checks.
  • Rollback criteria. Because the model volume is decoupled from the image, rolling back is normally a matter of redeploying the previous image tag and pointing it at the same volume — define in advance what latency or quality regression triggers that rollback.
  • Cost and budget controls. GPU time is the dominant cost driver; track it per environment (dev/staging/prod) so a forgotten always-on GPU container in a test environment does not become an invisible ongoing cost.
  • Privacy of production-derived data. If real user prompts or retrieved documents are logged for debugging or evaluation, treat those logs with the same access controls as the source data they came from, especially in regulated environments.

7. Common Mistakes

  • Running without a named volume. This seems harmless until the first container recreation, at which point every downloaded model disappears and must be re-pulled — costly on slow networks and disruptive during an otherwise routine redeploy.
  • Assuming GPU acceleration is active because the container "just started." A container without the NVIDIA runtime configured, or scheduled onto a node without a GPU, still starts successfully and still serves requests — just on the CPU. Teams sometimes spend hours tuning prompts to "fix slowness" that is actually a missing --gpus=all flag or an unconfigured toolkit.
  • Floating on the latest image tag in production. An unannounced upstream update can change default behavior or resource usage between two deployments that a team believed were identical, breaking reproducibility exactly when you most need it.
  • Exposing the API without any access control. The Ollama API does not include its own authentication layer by default; publishing port 11434 to a network reachable by untrusted clients turns an internal convenience into an open door.
  • Treating model swaps as configuration changes rather than releases. Swapping the model tag a production service points to changes its behavior as much as a code deploy would, and skipping evaluation before that swap is how quality regressions reach users unnoticed.
  • Conflating "the container is healthy" with "the model server is ready." A container can pass a basic process-alive check while the model is still loading or has just been evicted; use an API-level readiness check against the actual inference endpoint, not just container liveness.

❓ FAQ

Does running Ollama in Docker require a GPU?

No. The official image supports a CPU-only run command, and Ollama will serve requests using CPU inference. GPU acceleration is an optional path that requires the NVIDIA Container Toolkit (for NVIDIA GPUs) or the ROCm image tag (for AMD GPUs).

Where are models stored when Ollama runs in a container?

In the directory the official image writes to inside the container, typically mounted from a named Docker volume so model blobs and manifests persist independently of the container's own lifecycle.

Why can Docker Desktop on a Mac not accelerate Ollama with the GPU?

Docker Desktop on macOS does not pass the host GPU through to containers, so GPU-accelerated inference on a Mac is achieved by running Ollama natively rather than inside a Docker container there.

Can multiple applications share one Ollama container?

Yes — since Ollama exposes a standard REST API, any number of internal services can send requests to the same running container, though at higher concurrency you should measure whether one instance's GPU memory and throughput are sufficient before assuming it scales indefinitely.

Is Ollama in Docker suitable for high-throughput production inference?

It is a reasonable fit for lighter-weight internal tools and moderate traffic. Teams with high-concurrency, latency-sensitive serving needs typically evaluate purpose-built inference servers (such as vLLM or Triton) alongside it, as covered in the comparison table above.

🔗 References & Further Reading

📝 Summary

  • Ollama in Docker packages a local LLM server, its REST API on port 11434, and its model cache into a portable, reproducible unit.
  • Ports and volumes are the two container mechanics that determine whether clients can reach the server and whether models survive container recreation.
  • GPU acceleration requires explicit setup — the NVIDIA Container Toolkit for NVIDIA GPUs, or the rocm image tag with device mapping for AMD GPUs — and silently falls back to CPU if skipped.
  • A worked RAG example shows how Ollama fits into a multi-container application without needing to be the only moving part.
  • Production readiness means pinned image and model tags, network access control, real observability, and defined rollback criteria — not just a running container.
  • Most real incidents trace back to a handful of avoidable mistakes: missing volumes, assumed-but-absent GPU acceleration, floating tags, and unguarded network exposure.

Hopefully this gives you a solid, first-principles map of running Ollama in Docker — from a single laptop container to a service your team can actually operate with confidence. 🙌

Comments