Skip to main content

MCP Server Deployment Steps: From Local to Production

Calculating read time…

📚 New to MCP itself? Start one step back with Build Your First MCP Server (Enterprise Edition) before coming back here — this post picks up right after you've got a working server on your own laptop.

You built an MCP server that works on your laptop. Now what? A server that only runs when you personally type a command isn't a tool your team, your agent host, or a Fortune 500 rollout can depend on — it's a demo. Turning it into something real means five concrete moves: putting the code somewhere durable (Git), picking a place for it to live day and night (a remote host), getting it running there (deployment), making sure only the right callers can reach it (secure exposure), and proving all of that actually works from outside your own machine (remote testing). 🧭

Why this matters: The gap between "runs on my machine" and "runs in production" is where most MCP rollouts quietly stall or, worse, quietly leak access. An MCP server that talks to a ticketing system, a database, or an internal API is a credential-holding, tool-executing process — if it's exposed carelessly, it's not a convenience, it's a new attack surface. Enterprises adopting MCP (Anthropic's own Claude Desktop and Claude Code, Microsoft's Copilot Studio, Block's internal tooling, Sourcegraph's coding agents) all converge on the same shape: source control, a hardened host, a locked-down network path, and a repeatable way to verify the thing actually does what it claims before anyone plugs it into a real agent. 🔐

Five-step journey from a local MCP server to a deployed, secured, tested one

🔀 Quick Comparison: stdio vs. Streamable HTTP

Everything below hinges on one fork in the road: an MCP server built for stdio only talks to a process on the same machine that launched it. To let anything remote — a teammate's laptop, a hosted Claude integration, a CI job — reach it, the server needs to speak over the network using the Streamable HTTP transport. That single distinction decides almost every choice in this post.

Dimension stdio Streamable HTTP
Who can connect Only the process that spawned it, on the same machine Any client that can reach the endpoint's URL
Auth needed? Usually none — the OS process boundary is the security boundary Yes — treat it like any public API (bearer tokens / OAuth)
Typical host Your own laptop, launched by Claude Desktop or an IDE A cloud VM, container platform, or serverless endpoint
Network exposure to secure None TLS, firewall rules, Origin validation, token checks
Good fit for Personal tools, local file access, quick prototypes Team or company-wide tools, hosted agent integrations

🎯 Use this when: you're deciding whether "deploy" even applies to you — if your server will only ever be launched by you locally, you can stop at stdio. Everything past this point assumes you're moving to Streamable HTTP.

1️⃣ Put the Project in Git

Kid analogy: imagine your MCP server is a recipe you keep improving. Git is the recipe box where every version you've ever written is saved, labeled, and never lost — even if tonight's batch turns out burnt, yesterday's recipe is still sitting safely in the box. 🍪

Why it's the actual first step, not a formality: every later step in this post — building a container image, running a deploy script on a remote host, rolling back a bad release — assumes there's a single, versioned source of truth for the server's code. GitHub's own remote MCP servers, Sentry's hosted MCP server, and Cloudflare's reference MCP-on-Workers examples are all published as ordinary Git repositories precisely so that "what's running in production" and "what's in the repo" can be kept honest against each other.

git init git add . git commit -m "Initial commit: working MCP server" git remote add origin https://github.com/<you>/<your-mcp-server>.git git push -u origin main
✅ Practical example: tag every release you deploy — git tag v1.0.0 && git push --tags — so that when Step 3's deployment script checks out code on the remote host, it's pulling an exact, named version instead of "whatever main happened to be."
💡 Key warning: never commit secrets — API keys, database passwords, OAuth client secrets — into the repository, even a private one. Use a .gitignore entry for your .env file and load secrets on the host instead. A secret committed once stays in Git history forever, even after you delete the file in a later commit.

🎯 Use this when: you're about to touch a remote host for the first time — do this before Step 2, not after, so the very first thing that lands on the server is a clean checkout, not a folder you dragged over by hand.

2️⃣ Choose a Remote Host (OCI, Step by Step for Beginners)

Kid analogy: your recipe box (Git) is safe, but nobody can taste the cookies until someone actually bakes them somewhere. A remote host is the oven that's always on, even while you're asleep, so anyone who wants a cookie can get one. 🏠

There are several reasonable places to run a small MCP server — a container platform, a serverless function runtime, or a plain virtual machine. This walkthrough uses Oracle Cloud Infrastructure's Always Free tier because it's one of the few options that gives a beginner a real, always-on Linux VM (not a time-boxed trial) at zero cost, which matters when you're still learning and don't want a surprise bill. The same steps map onto AWS, Azure, or GCP with different menu names.

🧪 Hands-on lab: stand up a disposable OCI VM for your MCP server

1
Create a free OCI account at cloud.oracle.com. You'll be asked for a card for identity verification, but Always Free resources are never charged. Pick a home region carefully — it can't be changed later for that account.
2
In the console, go to Networking → Virtual Cloud Networks and click Start VCN Wizard → Create VCN with Internet Connectivity. Accept the defaults — this gives you a public subnet with an internet gateway, which your VM needs to be reachable at all.
3
Go to Compute → Instances → Create Instance. Give it a name, choose an Always Free–eligible shape (either the tiny AMD micro shape or, for more headroom, the Ampere A1 flexible ARM shape at 1–4 OCPUs / up to 24 GB RAM), and pick an Ubuntu LTS image.
4
Under Add SSH keys, choose "Generate a key pair for me" and immediately download the private key — it's shown exactly once. Leave networking on the VCN and public subnet you just made, then click Create.
5
Once the instance shows a running state, note its public IP, then connect: chmod 400 your-key.key && ssh -i your-key.key ubuntu@<public-ip>. Expect to see: an Ubuntu welcome banner and a shell prompt — that's confirmation the VM is alive and reachable.

Common first-timer snag: SSH refuses to connect. Nine times out of ten this is the two-layer firewall: OCI's Security List (or NSG) has to allow inbound TCP/22 from your IP and the instance's own OS firewall has to allow it too. Fix the cloud-level rule under Networking → your VCN → Security Lists first, then check sudo iptables -L or ufw status on the instance itself.

This toy instance is exactly the shape production deployments use, just smaller: a Fortune 500 team runs the same "VM behind a VCN with a security list" pattern, just with more OCPUs, a load balancer in front, and infrastructure-as-code (Terraform/OpenTofu) instead of clicking through the console by hand.

🎯 Use this when: you want a real, persistent Linux box to learn on without a recurring bill — swap in a managed container platform once your team needs autoscaling, multiple environments, or zero-downtime deploys.

3️⃣ Deploy the MCP Server

Kid analogy: the oven (your VM) is on, but the cookie dough (your code) still needs to get from your kitchen counter into the oven, and someone needs to press "start" and keep an eye on it so it doesn't switch off overnight. 🍪🔥

"Deploy" really means three things happening in order: get the code onto the host, make sure it keeps running (including after a crash or a reboot), and switch its transport from local stdio to network-reachable Streamable HTTP. On the small OCI VM from Step 2, the simplest reliable pattern is a systemd service — no container platform required for a first deployment, though Docker is a fine alternative once you want to package dependencies more portably.

# On the VM sudo apt update && sudo apt install -y git nodejs npm git clone https://github.com/<you>/<your-mcp-server>.git cd your-mcp-server && npm ci && npm run build # /etc/systemd/system/mcp-server.service [Unit] Description=My MCP Server After=network.target [Service] ExecStart=/usr/bin/node /home/ubuntu/your-mcp-server/build/index.js --transport http --port 8787 Restart=always User=ubuntu EnvironmentFile=/home/ubuntu/your-mcp-server/.env [Install] WantedBy=multi-user.target
✅ Practical example: after writing the unit file, run sudo systemctl enable --now mcp-server. Because it's registered with systemd, the same server that answered your Step 1 Git tag will restart automatically on a reboot or a crash — the difference between a demo and something a team can rely on.
💡 Key warning: at this point the server is listening on the VM's internal port (8787 above), reachable only inside the instance or over the raw OCI network — it is not yet safely exposed to the internet. Do not open that port directly in your Security List; Step 4 puts a reverse proxy in front of it first.

🎯 Use this when: your server is stable enough to leave unattended — if you're still actively debugging tool logic, keep iterating locally over stdio and only push to the VM once behavior is settled.

4️⃣ Expose the Server Securely

Kid analogy: now that the cookies are ready, you don't just leave the front door wide open for anyone walking by — you put in a mail slot (one narrow, controlled opening) and you check ID before handing out cookies to strangers. 🚪🔑

The current MCP specification is explicit that a Streamable HTTP server must validate the incoming Origin header to block DNS-rebinding attacks, should bind to localhost when running purely locally, and should implement real authentication for any connection that isn't local. In practice, three layers stack on top of each other on a real deployment:

  1. TLS termination — a reverse proxy (Caddy, Nginx, or a managed load balancer) holds the certificate and only ever forwards decrypted traffic to the MCP process over the loopback interface, never raw internet traffic straight to your Node or Python process.
  2. Network-level firewalling — on OCI this is the Security List/NSG plus the instance's own ufw/iptables; only ports 443 and a tightly-scoped 22 are ever opened outward.
  3. Application-level authorization — OAuth 2.1 with PKCE, following the discovery pattern real hosted MCP servers use today: a client hits your protected resource, gets a 401 pointing at a /.well-known/oauth-protected-resource document, follows that to your authorization server's metadata, and only then exchanges a code for a token.

GitHub's and Sentry's own remote MCP servers, along with Cloudflare's published workers-oauth-provider pattern, all follow this same shape — an OAuth-aware layer sitting in front of the MCP endpoint rather than trusting the network alone.

Original diagram showing an MCP client, the public internet, a reverse proxy handling TLS and origin checks, and an OCI VM running the MCP server alongside an identity provider
# Minimal Caddy reverse proxy — /etc/caddy/Caddyfile mcp.yourcompany.com { reverse_proxy 127.0.0.1:8787 } # Caddy requests and renews the TLS certificate automatically
✅ Practical example: point your domain's DNS A record at the OCI instance's public IP, install Caddy, drop in the three-line Caddyfile above, and reload it. The systemd-managed server from Step 3 is now reachable only through an HTTPS endpoint with a real certificate — the internal port never touches the public internet directly.
💡 Key warning: a bearer token or API key alone is not the same as OAuth-based authorization, and it doesn't scale to "revoke access for one user without rotating a shared secret for everyone." For anything beyond a personal single-user server, plan for per-user tokens from day one — retrofitting OAuth onto a server your whole team already depends on is far more painful than building it in from Step 4 the first time.

🎯 Use this when: more than one trusted person or one hosted agent platform needs to reach this server — a single developer testing against their own always-local server can defer full OAuth, but nothing that leaves your own machine should skip TLS and Origin validation.

5️⃣ Test It Remotely

Kid analogy: before you tell the whole neighborhood the cookies are ready, you taste one yourself — from outside the kitchen, the way a guest actually would, not while standing right next to the oven. 🍪✅

The official MCP Inspector is the standard tool for this: it's both an MCP client and a small local web UI, and it can point at a URL exactly the way a hosted agent platform would, rather than spawning your server as a local subprocess.

# From your own laptop, NOT the VM npx -y @modelcontextprotocol/inspector # In the Inspector UI: # - Transport: Streamable HTTP # - URL: https://mcp.yourcompany.com # - Paste your bearer token / complete the OAuth flow if prompted # - Click Connect, then List Tools
✅ Practical example: a green "Connected" state plus a populated tools list confirms the whole chain from Steps 1–4 actually works end to end: DNS resolves, TLS is valid, the reverse proxy forwards correctly, the MCP process answers, and (if configured) the OAuth handshake completes.
💡 Key warning: testing only from the VM itself (e.g. curl localhost:8787) proves the process runs — it proves nothing about DNS, certificates, firewall rules, or the reverse proxy. Always run the Inspector, or a real client like Claude, from a genuinely separate network before calling deployment "done."

🎯 Use this when: right before you hand the URL to a teammate or wire it into a hosted agent — treat a clean remote Inspector session as your release gate, the same way a health check gates a normal web deployment.

🏢 Enterprise Rollout at Scale

Everything above scales from "one server, one developer" to "dozens of servers, whole company" by formalizing the same five steps rather than replacing them:

  • Server governance and allow-listing. Maintain an internal registry of which MCP servers are approved, and require agent hosts (Claude Code, Claude Desktop, Copilot Studio) to connect only to registry-listed URLs — an unlisted server is treated the same as an unreviewed third-party dependency.
  • Credential and OAuth handling. Centralize the OAuth authorization server rather than letting every team stand up its own; this is exactly why patterns like Cloudflare's shared OAuth-provider library exist — one hardened identity layer in front of many MCP servers, instead of many bespoke ones.
  • Versioning of servers. Pin Git tags to deployed releases (Step 1) and expose a version identifier from the server itself so a client can detect a stale or unexpectedly-rolled-back deployment.
  • CI enforcement. Run the MCP Inspector's underlying SDK in CI against a staging deployment on every merge, so a broken tool schema is caught before it reaches the production URL your whole company points at.
  • Observability for tool calls. Log every tools/call invocation — caller identity, tool name, latency, and outcome — the same way you'd log any other API endpoint that touches sensitive systems; without this, a compromised or misbehaving agent is invisible until real damage is done.

🎯 Use this when: a second team, or a second server, is about to join your setup — the moment MCP stops being "your project" is the moment these five practices stop being optional.

⚠️ Common Mistakes

Treating a remote MCP server as trusted by default. Because local stdio servers inherit the OS process boundary for free, teams sometimes carry that same trust assumption into a network-exposed server. Reasoning: a remote endpoint has no such inherent boundary — anyone who can reach the URL is, by default, a stranger, and needs to prove otherwise via auth.
Skipping input/output schema validation. A tool that trusts whatever JSON a client sends, or returns whatever a downstream API hands back unfiltered, turns a schema mismatch into a runtime failure — or worse, a way to smuggle unexpected data into a model's context. Reasoning: MCP tool schemas exist specifically so both sides can reject malformed data before it does damage.
Granting over-broad tool permissions. A "database" tool that can run arbitrary SQL, when the actual need was "look up an order by ID," gives an agent (and anyone who compromises it) far more reach than the task requires. Reasoning: MCP tools should mirror the principle of least privilege exactly like any other API scope.
Ignoring transport security details. Deploying Streamable HTTP without Origin validation, or binding a "just testing" server to 0.0.0.0 on a machine that's actually internet-facing, reopens exactly the DNS-rebinding risk the specification calls out. Reasoning: these protections are cheap to add and expensive to retrofit after a real incident.
Confusing "it runs" with "it's deployed." A process that's up right now but not managed by systemd (or an equivalent) will not survive a reboot, a crash, or a host reclaim — and nobody finds out until the tool silently stops answering mid-workflow. Reasoning: process supervision is what turns "currently running" into "reliably available."

❓ FAQ

Do I need OCI specifically, or will any cloud provider work?

Any provider works — OCI is used here because its Always Free tier gives a persistent VM at no cost, which lowers the barrier for a first deployment. The Git → host → deploy → expose → test sequence is identical on AWS, Azure, GCP, or a bare-metal box.

Can I skip OAuth if only I will ever use this server?

A single-user server can get by with a long random bearer token instead of full OAuth as a temporary measure, but it should still sit behind TLS and Origin validation. The moment a second person or a hosted platform needs access, move to OAuth so access can be granted and revoked per user.

Do I still need a reverse proxy if my MCP framework has built-in HTTPS?

Usually yes. A dedicated reverse proxy centralizes certificate renewal, can enforce Origin and rate-limit rules independently of your application code, and lets you swap or scale the MCP process behind it without touching your TLS setup.

How is this different from just using ngrok to expose my local server?

A tunnel like ngrok is fine for a five-minute demo, but the URL and your laptop's uptime are tied together, and you typically don't control firewalling or certificate lifecycle. A real deployment decouples the server's availability from your own machine being on.

What's the single most common reason a "successful" deployment fails remote testing?

Firewall mismatch — the cloud-level rule (Security List/NSG) and the host-level rule (ufw/iptables) both have to allow the same port, and beginners frequently fix only one of the two.

🔗 References & Further Reading

Official / primary sources (used for fact-checking protocol behavior and transport security requirements):

Companion reading on this blog:

Product and company names (Anthropic, Claude, Claude Desktop, Claude Code, GitHub, Cloudflare, Sentry, Oracle Cloud Infrastructure, Microsoft Copilot Studio, Block, Sourcegraph, and others) are trademarks of their respective owners, referenced here for identification purposes only. All explanations above are original synthesis based on publicly documented behavior, not reproductions of any vendor's text.

📝 Summary

  • Git first: version the server before it ever touches a host, and keep secrets out of the repo.
  • Choose a host: OCI's Always Free tier gives beginners a persistent, no-cost VM to deploy on.
  • Deploy: run the server under systemd (or a container) so it survives crashes and reboots, and switch it to Streamable HTTP.
  • Expose securely: TLS via a reverse proxy, a two-layer firewall, and OAuth 2.1 once more than one caller needs access.
  • Test remotely: use the official MCP Inspector from a separate network as your release gate, not a localhost curl.
  • At scale: the same five steps become governance, centralized identity, versioning, CI, and observability.

That's the full path from "it works on my laptop" to "it's a dependable tool other people and other agents can use." Happy shipping! 🚀

Comments