The Model Context Protocol (MCP) doesn't have one wire format — it has three, and only two of them are still recommended. stdio runs a server as a local subprocess and swaps messages over standard input/output; HTTP+SSE was the original remote transport and used two separate endpoints; Streamable HTTP replaced it with a single endpoint that can answer with a plain reply or open into a live stream. Picking the wrong one for the job is the single most common reason a working MCP integration falls over the moment it leaves a developer's laptop. 🔌
This isn't a cosmetic choice. GitHub's own hosted MCP server, Anthropic's remote connectors in Claude, and most enterprise MCP gateways today are built on Streamable HTTP specifically because the older SSE design couldn't survive a load balancer. Meanwhile Claude Desktop and Claude Code still default local, developer-machine servers to stdio, because for a single trusted process on one machine, a network transport is pure overhead. Get the mapping backwards — stdio across a network, or a stateful SSE stream behind an auto-scaling group — and you'll spend a weekend debugging "flaky" connections that were never flaky, just structurally wrong. 🏗️
The three MCP transports, as three ways two kids might pass notes.
📑 In This Post
- What a "transport" actually means in MCP
- stdio: the local subprocess pipe
- HTTP+SSE: the deprecated first attempt at remote MCP
- Streamable HTTP: one endpoint, two response modes
- Choosing a transport for a real deployment
- Enterprise rollout at scale
- Common mistakes
- FAQ
- References & further reading
- Summary
🔀 Quick Comparison
| Trait | stdio | HTTP+SSE (deprecated) | Streamable HTTP |
|---|---|---|---|
| Where it runs | Same machine, local subprocess | Remote, over a network | Remote, over a network |
| Endpoints | n/a (pipes, not URLs) | Two — a POST endpoint and a separate SSE endpoint | One — a single URL, answering to both request verbs |
| Statefulness | Implicit — tied to the process | Tied to a pinned, long-lived connection | Optional — session ID header, can be stateless |
| Scales behind a load balancer? | Not applicable | Poorly — needs sticky sessions | Yes — any instance can answer |
| Status | Current, recommended for local servers | Deprecated since MCP spec 2025-03-26 | Current, recommended for remote servers |
| Typical real user | Claude Desktop / Claude Code local servers | Early 2024–2025 remote MCP servers | GitHub's hosted MCP server, Claude's remote connectors |
1️⃣ What a "Transport" Actually Means in MCP
Kid analogy: imagine you and a friend are passing notes in class. The words on the note — "can I borrow your eraser?" — are the message. But how the note gets from your hand to theirs — sliding it across a shared desk, folding it into a paper airplane and throwing it, or handing it to a kid in the row ahead to pass along — that's the transport. The note reads the same either way. Only the delivery mechanism changes.
In MCP, every message — regardless of transport — is JSON-RPC 2.0: a request has a method name, an id, and parameters; a response carries a result or an error keyed to that same id; a notification is a one-way message with no id at all. The transport layer's only job is getting these JSON-RPC envelopes from client to server and back, reliably, in order, without corrupting them. Everything above that — initialize, tools/list, tools/call — is identical no matter which of the three transports carries it.
That separation is why an MCP server's tool logic rarely needs to know or care which transport is in use. A well-built server implementation exposes the same handler functions to a stdio loop and to an HTTP router; only the plumbing at the edges differs. This is also exactly why getting the transport choice wrong doesn't corrupt your tool logic — it just makes the plumbing leak, silently, usually under load rather than during a demo.
✅ Worked example: Anthropic's own modelcontextprotocol/servers reference repository ships an "everything" demo server whose tool-handling code is transport-agnostic — the same handlers can be wired up behind stdio or an HTTP listener depending on how the process is started. That's the pattern worth copying: write tools once, mount transports around them.
🎯 Use this when: you're deciding how to structure a new MCP server codebase — keep transport code in its own thin layer from day one, even if you're only shipping stdio today.
2️⃣ stdio: The Local Subprocess Pipe
Kid analogy: you and your sibling share a bedroom, and you've strung a tin-can telephone between your beds. It only works because you're both in the same room — there's no way for a kid down the street to tap into that string, and if either of you leaves the room, the line goes dead. That's stdio: a private, one-to-one wire that exists only because the two ends were launched together.
Mechanically: the host application (say, Claude Desktop) starts the MCP server as a child process. It doesn't open a socket or a port — it just launches the executable and gets back a handle to that process's standard input and standard output streams. Every JSON-RPC message the client wants to send is written as a single line to the child's stdin. Every reply the server sends back is a single line written to its own stdout. A message can never break across more than one line internally, since the line break itself is what tells the reader where one message ends and the next begins — that's the whole framing scheme, no length prefixes or brackets required. Anything the server wants to log for a human goes to stderr, which the client may show, save, or throw away, since it is never parsed as protocol traffic.
Step-by-step, a stdio session looks like this:
- The host spawns the server binary as a child process and keeps a reference to its stdin/stdout file descriptors.
- The client writes an
initializerequest to stdin, declaring its protocol version and capabilities. - The server writes back its own capabilities on stdout, and the client confirms with an
initializednotification. - Requests and responses continue flowing line-by-line, in either direction, for the life of the process.
- The host closes stdin (or terminates the process directly) to end the session; there's no separate "logout" handshake.
Why enterprises still lean on this for local tooling: there is no network attack surface to defend. No port is opened, no TLS certificate needs rotating, no reverse proxy needs configuring. A company shipping an internal developer tool as an MCP server — say, a proprietary build-system integration — can hand engineers a single binary that Claude Code or Claude Desktop launches on demand, with the operating system's own process permissions doing the access control. That's a meaningfully smaller thing to audit than a hosted HTTP service.
💡 What breaks without it: if a team tries to "share" a stdio server across a team by exposing it over SSH port-forwarding or a shared always-on VM, they lose the entire security model stdio was designed around — one client, one subprocess, one lifetime. At that point you've built a fragile, unauthenticated remote service and are better off simply using Streamable HTTP with real session management and auth, covered next.
🎯 Use this when: the server and the AI client run on the same trusted machine — local dev tools, IDE integrations, personal automation scripts.
3️⃣ HTTP+SSE: The Deprecated First Attempt at Remote MCP
Kid analogy: now imagine you want to send notes to a friend in a different classroom. You can't share a desk anymore, so you set up two separate mailboxes down the hall: one where you drop off notes, and a second one where your friend's replies pile up — except that second mailbox only "delivers" if someone keeps standing next to it the whole time, ready to catch what falls in. Walk away, and you miss everything until you come back and set it up again.
That's roughly what the original HTTP+SSE transport (spec revision 2024-11-05) required. A client would open a persistent Server-Sent Events connection to one endpoint to receive messages, and separately POST outgoing messages to a different endpoint. The SSE connection had to stay open for the server to push anything at all, which meant the client process needed a long-lived, pinned connection to one specific server instance for the whole session.
💡 Key warning: that pinned-connection requirement is precisely what made HTTP+SSE painful in production. Put a fleet of MCP servers behind a standard load balancer, and a request could land on instance A while the matching SSE stream was open on instance B — the reply would never arrive. Teams worked around this with sticky sessions, but that just traded one operational headache (session affinity, uneven load) for another. It's the same lesson as the tin-can-telephone-over-SSH anti-pattern above: stretching a design past the shape it was built for.
The protocol maintainers deprecated this design starting with spec revision 2025-03-26, replacing it with Streamable HTTP. Some servers and SDKs still accept SSE connections for backward compatibility with older clients, but new implementations are steered firmly toward the newer transport — and by the November 2025 spec revision, the ecosystem had largely finished the migration, with several vendors issuing formal sunset notices for their own SSE endpoints.
🎯 Use this when: only if you're maintaining an old client that hasn't migrated yet — otherwise, treat this section as historical context, not a design to copy.
4️⃣ Streamable HTTP: One Endpoint, Two Response Modes
Kid analogy: instead of two separate mailboxes, imagine one smart mailbox that can do either job depending on what you ask for. Drop in a quick question, and it hands you a one-line answer on the spot. Drop in something that takes a while — like "please assemble this 500-piece puzzle and tell me when it's done" — and instead of making you wait silently, it opens a little window and starts posting updates as it works, then finally slides the finished answer through. Same mailbox, same address, two different ways of replying.
Mechanically, Streamable HTTP collapses everything onto a single URL — commonly something like /mcp — which answers to both the POST and GET verbs. GitHub's own remote MCP server for Copilot follows exactly this pattern — a single hosted MCP endpoint that IDEs and agents like Claude Code connect to with an authorization header, no separate streaming URL required.
- The client sends an
initializerequest as an HTTP POST. TheAcceptheader on that request names two content types it can handle in reply —application/jsonfor a plain answer,text/event-streamfor a live one. - The server may respond with a session identifier in an
Mcp-Session-Idheader — a random, unguessable token, not a socket reference. - Every following request from that client includes the same session header, so any server instance behind the load balancer can look up the right state.
- For a quick tool call, the server can just answer with
Content-Type: application/jsonand close out immediately. - For a slower call, the server instead answers with
Content-Type: text/event-stream, streams progress notifications, then sends the final JSON-RPC response as the last event on that same stream before closing it. - Separately, the client may open a plain
GETto the same endpoint to hold a channel open for messages the server wants to push without being asked — a server-initiated notification rather than a reply to a specific call.
✅ Worked example: the same fast-path tool call, sketched as a command a developer might actually run against a local test server, followed by what comes back:
$ curl http://127.0.0.1:8000/mcp \
-H "Mcp-Session-Id: sess_9c02" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":7,"method":"tools/call",
"params":{"name":"lookup_order","arguments":{"order_id":"A1029"}}}'
# Server picks the synchronous path and closes out immediately:
{"jsonrpc":"2.0","id":7,
"result":{"content":[{"type":"text","text":"Order A1029: shipped"}]}}
Nothing here forced a stream to open — the server judged the lookup fast enough to answer in one shot. Bump the same request to something slower (say, a report that takes twenty seconds to assemble) and the only visible difference from the caller's side is that the reply takes longer to arrive as a sequence of stream events instead of one line.
The important detail is that the client never had to guess which mode it would get — the Accept header simply advertises that it can handle either, and the server picks based on how long the work will actually take. That flexibility is the entire reason this design replaced the rigid two-endpoint SSE approach.
💡 Harder case — resumability: networks drop connections mid-stream. Streamable HTTP servers can tag each SSE event with an ID; if a client's connection dies partway through a long tool call, it can reconnect with a Last-Event-ID header and the server can replay only what was missed, rather than restarting the whole call. This resumability option didn't exist at all under the old two-mailbox SSE design — it's a direct consequence of having one endpoint that owns the whole exchange.
🎯 Use this when: the server needs to be reachable over a network — SaaS integrations, internal platform tools shared across a company, anything a load balancer will sit in front of.
🧪 Try It Yourself: A Five-Minute Transport Comparison
Before touching a production server, it helps to actually watch the two current transports behave differently on a disposable example. This lab uses a single throwaway MCP server you can delete afterward.
Install the official Python SDK's example "everything" server in a scratch folder: pip install mcp, then run its built-in demo server with the stdio transport (the SDK's quickstart script defaults to this). Expect to see: the process starts and simply waits — no port opens, confirmed by netstat showing nothing new listening.
Point a local MCP client at it — Claude Desktop's developer settings, or the SDK's own test client — and add it as a stdio server. Expect to see: the tool list appears almost instantly, since there's no network round trip at all.
Now re-run the same example server with its Streamable HTTP option enabled (most SDK quickstarts expose a flag or environment variable for this) and open http://127.0.0.1:PORT/mcp with curl -i. Expect to see: a 4xx response to a bare GET with no proper Accept header — that's correct behavior, not a bug, since the spec allows a server to answer GET with 405 if it doesn't offer a standalone stream.
Common first-timer mistake: forgetting the Accept: application/json, text/event-stream header on the POST request. Without it, spec-compliant servers are entitled to reject the call outright — this is one of the most frequently filed "my MCP server doesn't work" bugs, and it's a client-side header, not a server bug.
The bridge to production: everything above used 127.0.0.1 with no auth, which is fine for a five-minute lab and dangerous for anything real. A production Streamable HTTP deployment adds TLS termination, Origin validation, and a real auth layer in front of that same single endpoint — the rollout section below covers exactly what changes.
5️⃣ Choosing a Transport for a Real Deployment
Most SDKs make it straightforward to support more than one transport from the same codebase — the tool-handling logic stays identical, and only the startup path differs depending on how the process is launched or configured. A common real-world pattern, visible across several open-source MCP servers, is defaulting to stdio for local development and switching to Streamable HTTP behind a feature flag for production deployment, without touching the underlying tool implementations at all.
| Situation | Recommended transport |
|---|---|
| IDE plugin, CLI tool, local dev assistant | stdio |
| Company-wide internal tool shared by many users | Streamable HTTP behind auth |
| Public SaaS product exposing an MCP integration | Streamable HTTP with OAuth 2.1 |
| Legacy client that hasn't updated its SDK yet | HTTP+SSE, only as a compatibility fallback |
🏢 Enterprise Rollout at Scale
Shipping a single MCP server for a demo and rolling one out to a whole organization are different problems. Once a Streamable HTTP server is meant to serve every engineer, or every customer, a handful of concerns move from "nice to have" to mandatory:
- Server governance and allow-listing. Treat every MCP server as an external dependency with an owner, a version, and a review — not something any team can point a client at unreviewed. A central registry of approved servers, keyed by URL and expected capability set, keeps "shadow MCP" from spreading the way shadow IT once did.
- Credential and OAuth handling. Streamable HTTP servers exposed beyond a single trusted machine should sit behind OAuth 2.1 with PKCE, not a long-lived static bearer token pasted into a config file. Session IDs identify a conversation, not a user — authorization still has to happen on every request via a proper token, validated server-side.
- Origin validation and network binding. The spec treats checking the incoming
Originvalue as non-negotiable, and pushes local deployments toward binding only the loopback address (127.0.0.1) instead of every network interface. Skip either step and a malicious web page can use DNS rebinding to reach a server that was only ever meant to answer localhost. - Versioning of servers. Tool schemas change. Pin a specific server version or protocol revision in CI, and treat a schema change in a production tool the same way you'd treat a breaking API change — with a changelog and a deprecation window, not a silent swap.
- CI enforcement. Run automated checks that every tool exposed by a server has an input schema, that destructive tools are flagged as such, and that the server responds correctly to malformed input, before it's ever added to the approved registry.
- Observability for tool calls. Because every tool call is a distinct JSON-RPC method with an id, they're naturally loggable as discrete events — capture latency, error rate, and which identity invoked which tool, the same way you'd instrument any other internal API.
✅ Worked example: GitHub's hosted MCP server requires a bearer token on every request to its single Streamable HTTP endpoint and documents distinct scopes for read versus write operations — a concrete illustration of pairing one endpoint with real, per-request authorization rather than trusting the session alone.
🎯 Use this when: an MCP server is about to be used by more than one team, or exposed outside a single developer's machine — that's the line where governance stops being optional.
⚠️ Common Mistakes
- Treating any MCP server as trusted by default. A tool call is a request for the server to do something — connecting to an unreviewed server and letting a model invoke its tools automatically is functionally similar to running unvetted code. The reasoning: MCP intentionally puts a lot of trust in whatever server a client is pointed at, and that trust has to be earned, not assumed.
- Skipping input and output schema validation. A tool that accepts freeform arguments without a schema, or returns unstructured text a model has to guess the shape of, invites both malformed calls and prompt-injection-style attacks hidden in tool output. The reasoning: schemas are the contract that lets a client sanity-check what it's sending and receiving before acting on it.
- Over-broad tool permissions. A single "run_sql" tool with unrestricted database access is far riskier than five narrow tools each scoped to one table or one read-only view. The reasoning: the blast radius of a single bad or manipulated tool call should be as small as the task actually requires.
- Ignoring transport security details. Skipping Origin validation on a Streamable HTTP server, or binding it to all interfaces instead of localhost during local development, is exactly the gap that enables DNS-rebinding attacks described in the spec's own security guidance. The reasoning: these protections exist because the attack has already been demonstrated in practice, not as theoretical hardening.
- Mixing up stateful assumptions across transports. Code written assuming stdio's implicit one-process-per-client model breaks in unexpected ways if naively lifted onto Streamable HTTP, where a fleet of interchangeable server processes may sit behind one shared address. The reasoning: state that lived safely in process memory under stdio needs to move into the session store once a server goes remote.
❓ FAQ
Is HTTP+SSE completely gone from MCP?
No — it's deprecated, not removed. Some servers and SDKs still accept it for backward compatibility with older clients, and the spec even documents a fallback handshake a client can use to detect which transport an unfamiliar server speaks. New servers, though, should target Streamable HTTP.
Can one server support both stdio and Streamable HTTP?
Yes. Because the transport layer only handles moving JSON-RPC messages, most SDKs let the same tool-handling code be wired up behind either entry point, choosing one at startup through a config setting — stdio for local development, Streamable HTTP for a shared deployment.
Does Streamable HTTP require the server to keep a connection open the whole time?
No, and that's the point. A server is free to hand back one plain JSON reply and be done, or shift into a live event stream only for the specific calls that genuinely need to trickle out progress before the final answer. Nothing forces a long-lived connection for every interaction.
Is the session ID in Streamable HTTP the same thing as authentication?
No — they solve different problems. The Mcp-Session-Id header ties related requests together across possibly-different server instances; it does not by itself prove who the caller is. Production deployments still need a separate authorization layer, typically OAuth 2.1, on top of session tracking.
Why not just use plain REST instead of designing a new transport layer?
MCP needs bidirectional messaging — a server can send requests and notifications to a client mid-session, not just replies. Plain request/response REST doesn't have a native way for the server to initiate that side of the conversation, which is exactly the gap Streamable HTTP's optional SSE upgrade fills.
🔗 References & Further Reading
Official/primary sources relied on:
- Model Context Protocol specification, Transports — modelcontextprotocol.io/specification/2025-03-26/basic/transports
- modelcontextprotocol GitHub organization (spec and SDK source) — github.com/modelcontextprotocol
Additional practitioner background (used for context, not quoted or closely followed):
- GitHub's own documentation for its hosted remote MCP server configuration
- Public vendor migration notices describing SSE-to-Streamable-HTTP transitions
Model Context Protocol, MCP, GitHub, Claude, Claude Desktop, and Claude Code are trademarks of their respective owners, referenced here for identification only. All explanations above are original synthesis written from the primary sources listed; no text has been reproduced verbatim from any source.
📝 Summary
- A transport moves JSON-RPC messages; it never changes what a tool call, resource, or prompt looks like.
- stdio is a private pipe between a host and a subprocess it launched — perfect for local, single-machine tools.
- HTTP+SSE was the original remote transport, split across two endpoints, and is now deprecated because it couldn't scale behind a load balancer.
- Streamable HTTP replaces it with one endpoint that can answer instantly or open into a stream, tracked by a session ID rather than a pinned connection.
- Enterprise rollouts need governance, OAuth, Origin validation, versioning, CI checks, and observability layered on top of whichever transport is chosen.
- Most production mistakes come from trusting a server by default or carrying stdio's single-process assumptions into a shared HTTP deployment.
That's the map: three transports, one still meant for daily use locally, one retired, and one carrying almost all new remote MCP traffic today. Pick the transport that matches where your server actually runs, not the one that's fastest to prototype. Happy building! 🚀
Comments
Post a Comment