Skip to main content

Dockerized AI Agents: Build, Containerize, and Deploy AI Agent Applications

Calculating read time…

A Dockerized AI agent is an autonomous or semi-autonomous LLM-driven program — the reasoning loop, the model connection, and its external tools — packaged so that each piece runs in its own isolated container instead of as scripts and API keys scattered across a developer's laptop. That separation matters because an agent is not just a chatbot: it decides, plans, and takes actions — searching the web, editing files, calling APIs, sometimes writing and running its own code — and every one of those actions is a place where an untrusted output can turn into a real-world side effect. 

The stakes are different from ordinary LLM serving. A single vLLM container that returns bad text is a quality problem; a single agent container with unrestricted filesystem or network access that gets manipulated by a prompt injection is a security incident. Containerizing the agent, its model connection, and its tool integrations separately is what turns "an LLM that can call functions" into something you can actually govern, audit, and roll back in production. This post walks through how that stack fits together, using Docker's own documented agentic AI tooling as the backbone. ⚙️

🔀 Quick Comparison: Ways to Wire an Agent to Its Tools

Approach Best for Trade-off
Hand-wired SDK calls in the agent process Quick prototypes, a single trusted tool Every new tool means new code, new credentials, no shared isolation
One monolithic container for agent + tools Small demos, single-developer use A compromised tool has the same blast radius as the agent itself
MCP Gateway brokering containerized tool servers Production agents, multiple tools, shared governance An added service to run and monitor; requires the Docker socket to spawn tool containers
Fully hosted agent platform (no local containers) Teams that don't want to operate infrastructure Less control over isolation boundaries and data residency

1. Foundations: What a Dockerized AI Agent Stack Actually Is

Kid-friendly analogy: imagine a school project where one kid does the thinking, one kid runs to the library for facts, and one kid keeps the hall pass that says who's allowed to leave the classroom. Cramming all three jobs into one kid means if that kid makes a bad decision, there's nothing stopping them from also wandering off with the hall pass.

Translated technically, a modern agentic application is described by Docker's own documentation as a stack built from three core components: a model doing the reasoning, an agent holding the orchestration logic, and a gateway that links the agent to external tools and services. Each of those three pieces has a different risk profile and a different update cadence, which is precisely why containerizing them separately — instead of bundling everything into one process — is the foundation of a governable agent deployment rather than a nice-to-have.

What it does: it turns an agent's reasoning loop, its connection to a language model, and its access to external tools into independently deployable, independently restartable, independently auditable units.
Why it's needed: agents act on the world. A tool integration bug or a compromised dependency in one tool shouldn't be able to reach the agent's model credentials or another tool's data just because they happened to share a process.
What fails without it: a single Python process running the agent logic, the model client, and every tool's SDK together means a vulnerability in any one dependency — a malicious package update in a web-scraping library, for instance — has access to everything else in that process, including API keys for the model provider.

🎯 Use this when you're deciding whether to bundle an agent and its tools into one script or split them into services from day one.

2. Mechanics: The Three-Part Agentic Stack

Kid-friendly analogy: think of a restaurant kitchen. The head chef (the agent) doesn't personally grow the vegetables or catch the fish (the model's raw reasoning power) and doesn't personally answer the phone for delivery orders (the tool integrations) — each role is a separate station, and the chef just coordinates between them.

Docker's documentation names the three layers directly. The model layer is the reasoning engine — a locally-run open model or a hosted API — doing the writing and planning. The agent layer is where the orchestration logic lives: it takes a goal, breaks it into steps, and talks to the model, the tools, and the user interface. The MCP gateway layer links the agent to the outside world through the Model Context Protocol, giving the agent a standard way to call external capabilities without hand-rolling a custom integration for every API.

In a containerized deployment, each of these typically becomes its own Compose service: a model-runner service serving an OpenAI-compatible API locally, an agent service holding the orchestration code, and a gateway service brokering tool access. Docker Compose then wires them together with service-level networking, so the agent reaches the model and the gateway by service name rather than hardcoded addresses.

✅ Worked example: in Docker's published fact-checking sample application, an Auditor agent coordinates two sub-agents — a Critic that verifies claims using a web-search tool, and a Reviser that edits the answer based on the Critic's findings. The model, the agent logic, and the tool access each sit behind their own service boundary in the same Compose file.

3. How the MCP Gateway Isolates and Governs Tool Access

Kid-friendly analogy: a hotel concierge doesn't hand guests a master key and point them toward the supply closet. The concierge listens to the request, decides which staff member is authorized to fulfill it, and only that staff member ever touches the closet.

What it does: the MCP Gateway runs Model Context Protocol servers — the "tools" an agent can call — each in its own isolated container with restricted privileges, network access, and resource usage, and centralizes logging and call-tracing across every tool call.
Why it's needed: without a broker in the middle, an agent that wants to use ten different tools either needs ten different SDKs wired into its own process, or ten different sets of credentials it manages itself — both of which widen the attack surface and make it hard to answer "what did this agent actually do" after the fact.
How it works step by step:

  1. The agent sends a tool-call request to the Gateway rather than to the tool directly.
  2. The Gateway identifies which registered MCP server handles that tool and, if it isn't already running, starts it as its own Docker container.
  3. The Gateway injects any credentials the tool needs and applies its configured security restrictions before forwarding the request.
  4. The tool's container processes the request in isolation and returns the result through the Gateway.
  5. The Gateway logs the call, giving a centralized, auditable record of every tool invocation across the whole agent fleet.

What fails without it: each tool integration becomes bespoke code with its own dependency footprint running inside the agent's process boundary, and there's no single place to see, restrict, or revoke what an agent has been doing with its tools.

💡 Trade-off: container isolation limits the blast radius of a compromised tool server, but it is not a substitute for agent-level authorization controls — a legitimately-behaving tool called by a manipulated or over-permissioned agent can still cause damage the container boundary was never designed to stop. Both layers of defense are needed together, not one instead of the other.

🎯 Use this when an agent needs more than one or two external tools, or when several teams need shared, auditable tool governance.

4. Real Example: A Fact-Checking Agent Built With Compose

Docker publishes a working sample that demonstrates this pattern end to end: a web application where an agent verifies a user-submitted claim using a live web search tool, then revises its answer based on what it found. The whole stack is defined in one Compose file.

# Simplified illustration based on Docker's documented compose.yaml pattern
services:
  adk:
    build:
      context: .
    ports:
      - "8080:8080"
    environment:
      - MCPGATEWAY_ENDPOINT=http://mcp-gateway:8811/sse
    depends_on:
      - mcp-gateway
    models:
      gemma3:
        endpoint_var: MODEL_RUNNER_URL
        model_var: MODEL_RUNNER_MODEL

  mcp-gateway:
    image: docker/mcp-gateway:latest
    use_api_socket: true
    command:
      - --transport=sse
      - --servers=duckduckgo

models:
  gemma3:
    model: ai/gemma3:4B-Q4_0
    context_size: 10000

Reading this the way an architect would: the adk service is the agent application, built from a local Dockerfile. It doesn't talk to the search tool directly — it points at the gateway's SSE endpoint through an environment variable. The mcp-gateway service is a Docker-maintained image that spins up the requested tool servers, in this case a search tool, on demand. The top-level models block and the service-level models block together tell Compose to run a local model and automatically inject its connection details into the agent service — no manually-managed endpoint URL required.

In the underlying sample, the agent framework coordinates three named sub-agents: one that verifies factual claims using the search tool, one that edits the answer based on what was found, and a top-level agent that sequences the two and acts as the entry point for the whole flow.

🎯 Use this when you want a documented, working starting point for a tool-using agent before building your own from scratch.

5. Implementation: Building the Agent's Own Container

Kid-friendly analogy: packing a lunchbox the night before means you're not scrambling in the morning, and you know exactly what's in it. Building the agent's dependencies into the image ahead of time means the container that starts in production has exactly what it was tested with — nothing installed on the fly, nothing missing.

The agent service itself is a normal application container, built with the same discipline as any other production image: dependencies installed and layer-cached during the build, application code copied in afterward so code changes don't invalidate the dependency layer, and the process run as a non-root user rather than the image default.

# Simplified illustration based on Docker's documented agent Dockerfile pattern
FROM python:3.13-slim
ENV PYTHONUNBUFFERED=1
WORKDIR /app

COPY pyproject.toml uv.lock ./
RUN pip install uv && uv pip install --system .

COPY agents/ ./agents/

RUN useradd --create-home --shell /bin/bash app     && chown -R app:app /app
USER app

ENTRYPOINT ["/entrypoint.sh"]

Two details in the documented pattern are easy to miss and matter operationally. First, dependency files are copied and installed in a separate layer before the application code, so Docker's build cache can reuse that layer on every rebuild that only changes agent logic — a meaningful iteration-speed difference once the dependency list is large. Second, the container drops root privileges before the entrypoint runs, so an agent process that gets tricked into executing something it shouldn't still can't write outside its own user's permissions.

The entrypoint script itself is where routing logic between a hosted API and a local model typically lives — covered next, because it's the mechanism that decides where the agent's intelligence actually comes from at runtime.

6. Local Models, Secrets, and the Hosted-API Fallback

Kid-friendly analogy: think of a backup generator that only kicks in when the main power is unavailable — the house doesn't need to know or care which source is actually powering the lights, as long as the switch happens automatically.

A documented entrypoint pattern checks for a Docker secret containing a hosted API key at startup. If a real key is present, it's read from the secret file and exported as an environment variable, and the agent talks to the hosted provider. If no key is present, the entrypoint instead points the agent's OpenAI-compatible client at a locally-running model server, so the exact same application code can run against either a hosted model or a fully local one depending only on which secret was provided at deploy time.

Why it's needed: teams often want to develop and test against a free local model, then flip to a hosted provider for production quality without maintaining two separate codebases.
What fails without it: hardcoding a single model provider means every environment switch is a code change, and local development either requires a paid API key or diverges meaningfully from what runs in production.

✅ Worked example: a Compose secret named for an API key is mounted into the container at a fixed path at runtime. The entrypoint script reads that file if it exists, exports it as the API key environment variable, and the same container image seamlessly targets a hosted provider in staging and a local model during a developer's offline testing, with no image rebuild required either way.

7. From One Laptop to a Fleet: Scaling Agent Stacks

A Compose file that runs an agent, a gateway, and a model on one machine is the right starting point for development, but production agent fleets generally need an orchestrator on top for the same reasons any production service does: automatic restarts, horizontal scaling, and scheduling across multiple hosts rather than one. The same container images built and tested with Compose become the deployable units in Kubernetes or another orchestrator — the agent, gateway, and model-serving containers don't need to be rebuilt, only redeployed with orchestrator-specific manifests for networking, secrets, and scaling policy.

💡 Trade-off: the Compose-file approach documented above works cleanly for a homogeneous stack — one agent framework, one model provider, one gateway. Organizations running a polyglot agent ecosystem, with multiple agent frameworks, multiple model providers, and multiple identity systems side by side, need integration work that a single Compose file doesn't abstract away on its own. Plan for that complexity rather than assuming it scales for free.

🎯 Use this when a working local agent stack needs to support real user traffic with uptime guarantees rather than a single developer's session.

8. Enterprise Rollout: Governance for Autonomous Containers

Agent containers earn a stricter rollout process than a typical stateless service because they take actions, not just return responses. A durable governance process for agent deployments typically covers:

  1. Tool allow-listing. Configure the gateway's server list explicitly rather than granting an agent open-ended access to every catalog tool; an agent should only ever be able to reach the specific tools its use case requires.
  2. Credential and secret isolation. API keys and tool credentials belong in the orchestrator's native secrets mechanism, injected into the gateway or the tool container at runtime rather than the agent's own process — the agent should generally never hold raw credentials for tools it merely calls through the gateway.
  3. Image provenance. Only run gateway and tool-server images that carry verifiable build provenance and signing, since a tampered tool image sitting inside an agent's trusted call path is a direct path to data exfiltration or unauthorized action.
  4. Access controls on the gateway's logs. Tool-call logs can contain sensitive request and response content; restrict who can read them the same way you would restrict access to application logs containing user data.
  5. Change control on the tool allow-list. Adding a new tool to an agent's available set is a permission change, not a configuration tweak — route it through the same review process as a new production credential.
  6. Incident response for agent actions. Document how to immediately revoke a specific tool's access or halt an agent's gateway connection if it's observed taking unintended actions, separate from a general service rollback.

🎯 Use this when an agent moves from a personal experiment to something that can take actions affecting real users, data, or systems.

9. Observability and Evaluation for Agent Behavior

Evaluating an agent is harder than evaluating a single model response, because the thing under test is a multi-step sequence of decisions, not one output. A few practices carry over from general LLM evaluation, and a few are specific to agents:

  • Tool-call tracing. Because the gateway logs every tool invocation centrally, use that log as the basis for tracing what an agent actually did in response to a given input — not just what it said.
  • Task success rate on representative scenarios. Build a fixed, versioned set of realistic tasks the agent should complete, and re-run it against every new agent version or prompt change before promoting it, the same discipline as a regression test suite for code.
  • Human review of edge cases. Automated task-success metrics miss subtle failures — an agent that technically completes a task but takes an unnecessary or risky path through its tools. Periodic human review of transcripts catches what pass/fail metrics don't.
  • Canary rollout for new agent or tool versions. Route a small share of traffic to an updated agent container or a newly added tool before a full cutover, and watch for a rise in tool-call errors, unexpected tool usage patterns, or task failures.
  • Cost and latency per completed task. An agent that calls tools repeatedly to reach an answer can be far more expensive and slower than the model cost alone suggests; track cost and latency per successfully completed task, not just per model call.

🎯 Use this when you need confidence that an agent will behave predictably before letting it run against real user requests unsupervised.

10. Common Mistakes

  • Giving the agent's own container direct tool credentials instead of routing through the gateway. This defeats the entire point of centralized governance — there's no single place left to revoke access, audit calls, or apply a restriction, because the agent process now holds the keys itself.
  • Granting a broad tool allow-list "to be safe" instead of a narrow one. Every additional tool an agent can reach is an additional action a manipulated or buggy agent could take; unused tool access is pure risk with no corresponding benefit.
  • Bundling the agent, model client, and every tool SDK into one container. This collapses the isolation the three-part stack is designed to provide, so a vulnerability anywhere in that dependency tree has access to everything else in the process, including model credentials.
  • Running the agent process as root inside its own container. If the agent is ever manipulated into writing files or executing something unintended, running as root removes the one filesystem-level barrier that would have limited the damage.
  • Skipping tool-call logging review until after an incident. The gateway's centralized logs are only useful if someone actually looks at them; treat unusual tool-call patterns as an active signal to investigate, not an archive to check retroactively.
  • Testing only the happy path before shipping a new agent version. An agent that completes the demo scenario flawlessly can still take a wildly different, unsafe path through its tools on a slightly different input — a fixed evaluation scenario set catches this only if it's actually run before every release, not just once during initial development.

❓ FAQ

Is the MCP Gateway required to run a Dockerized AI agent, or can an agent call tools directly?

It's not strictly required — an agent can call an SDK or API directly from its own container. The Gateway exists to solve the problems that approach creates at scale: isolating each tool in its own container, centralizing credential handling, and logging every tool call in one place, rather than repeating that work per tool inside the agent's own code.

Does container isolation alone make an AI agent safe to run unsupervised?

No. Container isolation limits the blast radius of a compromised tool server, but it does nothing to stop a legitimately-behaving tool from being misused by an agent that has been manipulated or over-permissioned. Agent-level authorization controls and a narrow tool allow-list are a separate, necessary layer of defense.

Can the same agent container run against a hosted model API and a local model without changes?

Yes, in the documented pattern this is handled entirely by the entrypoint script and environment variables set at deploy time — if a hosted API key secret is present, the agent uses it; otherwise it falls back to a locally-running model server automatically, with no code or image change required.

Do I need Kubernetes to run a Dockerized agent stack, or is Docker Compose enough?

Docker Compose is enough for local development, testing, and small-scale single-host deployments. Once you need automatic restarts across multiple hosts, horizontal scaling of the agent or gateway, or scheduling across a shared GPU fleet, the same container images typically move to Kubernetes or a similar orchestrator on top of the Compose-defined stack.

What's the difference between the MCP Toolkit and the MCP Gateway?

The Toolkit is the broader system for discovering, configuring, and managing MCP servers, including a catalog of pre-built, signed tool images. The Gateway is the runtime component within that system that actually brokers requests between an agent and the tool containers it manages, applying security restrictions and logging along the way.

🔗 References & Further Reading

Docker, Docker Compose, Docker Desktop, Docker Model Runner, and MCP Gateway are trademarks of Docker, Inc.; the Model Context Protocol (MCP) is an open standard originated by Anthropic. These names are referenced for identification only.

📝 Summary

  • A Dockerized AI agent stack splits reasoning (the model), orchestration (the agent), and tool access (the MCP gateway) into separately containerized components.
  • The MCP Gateway runs each tool server in its own isolated container, injecting credentials and logging every call centrally rather than trusting the agent's process with direct tool access.
  • Docker's documented sample application shows this pattern concretely: an agent service, a gateway service, and a model block wired together in one Compose file.
  • The agent's own container follows standard production Dockerfile discipline — cached dependency layers and a non-root user.
  • An entrypoint pattern lets the same agent image target a hosted model API or a local model automatically, based on which secret is present at deploy time.
  • Scaling beyond one host generally means moving the same container images to Kubernetes or a similar orchestrator on top of the Compose-defined stack.
  • Enterprise governance for agents centers on narrow tool allow-lists, credential isolation, image provenance, and access-controlled logs.
  • Evaluation has to cover multi-step agent behavior, not just single responses: tool-call tracing, task success rate, human review, and canary releases for new agent or tool versions.
  • Most agent-container incidents trace back to skipped isolation — direct tool credentials in the agent process, over-broad tool access, or bundling everything into one container.

That's the full path from three loose Python scripts to a governed, containerized agent stack — happy shipping! 🚀

Comments