Skip to main content

Monolithic vs Microservices in Model Serving: Architecture, Scalability, and Deployment

Calculating read time…

Post training of LLM , Now comes the real question: how do you actually serve it to thousands of users?

Do you bundle everything into one big server? Or do you split it into many small, independent pieces?

This is the monolithic vs. microservices debate — one of the most important architectural decisions in LLM engineering. Get it wrong and your system collapses under real traffic. Get it right and you have a serving system that scales gracefully for years.

First — What is "Model Serving" Anyway?

Model serving is the process of taking a trained AI model and making it available to answer real requests from users — reliably, quickly, and at scale.

💡 Think of it like this:

Your trained model is a brilliant chef who knows how to cook amazing meals. Model serving is the entire restaurant system around that chef — the front desk that takes orders, the kitchen manager who assigns work, the waiters who deliver food, and the billing system that tracks everything. The chef alone is not enough. The whole system matters!

A complete model serving system handles all of these pieces:

  • 📥 Receiving requests → A user sends a question or task to your model
  • 🔐 Auth and rate-limiting → Is this user allowed to call us? How often?
  • ⚙️ Pre-processing → Cleaning and formatting the input before the model sees it
  • 🧠 Running inference → Passing the input through the model to get a result
  • 📤 Post-processing → Formatting and cleaning the output before sending it back
  • 📊 Logging and monitoring → Tracking speed, errors, and usage patterns

Now — where do you put all of these pieces? That is the architectural question we are about to answer! 🗺️

The Restaurant Analogy — Your Core Mental Model 🍽️

Imagine you are opening a restaurant. You have two completely different ways to run it:


  ╔══════════════════════════════════════════════════════════════╗
  ║   Option A: THE MEGA RESTAURANT  (Monolithic)               ║
  ╠══════════════════════════════════════════════════════════════╣
  ║                                                              ║
  ║   One giant building.                                        ║
  ║   One kitchen does everything:                               ║
  ║     appetizers, mains, desserts, drinks — all in one place.  ║
  ║   One team of staff handles every single task.               ║
  ║   One cash register. One menu. One roof.                     ║
  ║                                                              ║
  ║   ✅ Simple to set up. Everything is in one place.           ║
  ║   ❌ If kitchen catches fire → WHOLE restaurant closes.      ║
  ║   ❌ Can't hire extra dessert chefs without hiring everyone. ║
  ╚══════════════════════════════════════════════════════════════╝


  ╔══════════════════════════════════════════════════════════════╗
  ║   Option B: THE FOOD COURT  (Microservices)                 ║
  ╠══════════════════════════════════════════════════════════════╣
  ║                                                              ║
  ║   Many separate stalls inside one building.                  ║
  ║   Stall 1: Pizza only.                                       ║
  ║   Stall 2: Sushi only.                                       ║
  ║   Stall 3: Drinks only.                                      ║
  ║   Each stall has its own team, kitchen, cash register.       ║
  ║                                                              ║
  ║   ✅ If pizza stall closes → sushi and drinks still work!   ║
  ║   ✅ Hire 5 extra pizza chefs without touching sushi stall. ║
  ║   ❌ More complex to coordinate. More moving parts.          ║
  ╚══════════════════════════════════════════════════════════════╝

This is the exact trade-off between monolithic and microservices model serving. Keep this picture in your mind — we will refer back to it throughout! 🧠

Part 1 — Monolithic Model Serving 🏛️

What is a Monolithic Architecture?

In a monolithic serving system, every component lives together in one single unit. One codebase. One running process. One server.

The API handler, the pre-processor, the model engine, the post-processor, the logger, and the auth layer all run inside the same application. They share memory, share the CPU/GPU, and are deployed together as one block.


  ┌─────────────────────────────────────────────────────────────────┐
  │                  MONOLITHIC SERVING APPLICATION                  │
  │                                                                  │
  │   ┌────────────┐    ┌────────────┐    ┌──────────────────────┐  │
  │   │  API Layer │───▶│   Pre-     │───▶│  Model Inference     │  │
  │   │            │    │  Processor │    │  Engine (LLM)        │  │
  │   └────────────┘    └────────────┘    └──────────┬───────────┘  │
  │                                                   │              │
  │   ┌────────────┐    ┌────────────┐    ┌──────────▼───────────┐  │
  │   │  Logger &  │    │  Auth &    │    │  Post-Processor      │  │
  │   │  Monitor   │    │ Rate Limit │    │  & Formatter         │  │
  │   └────────────┘    └────────────┘    └──────────────────────┘  │
  │                                                                  │
  │       All inside ONE process. ONE codebase. ONE deployment.      │
  └─────────────────────────────────────────────────────────────────┘
                                  │
                                  ▼
                           User gets response

How Does a Monolith Handle a Request?

Think of it like a Swiss Army knife — everything is in one tool, and each function is reached in a straight line. Here is the journey of one user request through a monolith:


  User sends: "Summarize this article for me"
        │
        ▼
  ┌─────────────────────────────────────────┐
  │  STEP 1 — API Layer receives request    │
  │  "Got it! Passing to auth checker..."   │
  └───────────────────┬─────────────────────┘
                      │
                      ▼
  ┌─────────────────────────────────────────┐
  │  STEP 2 — Auth Check                   │
  │  "Is this API key valid?               │
  │   Has this user hit their limit?"      │
  └───────────────────┬─────────────────────┘
                      │
                      ▼
  ┌─────────────────────────────────────────┐
  │  STEP 3 — Pre-Processing               │
  │  "Clean and format the input text.     │
  │   Apply the chat template."            │
  └───────────────────┬─────────────────────┘
                      │
                      ▼
  ┌─────────────────────────────────────────┐
  │  STEP 4 — Model Inference (GPU)        │
  │  "Run the LLM! Generate the summary."  │
  └───────────────────┬─────────────────────┘
                      │
                      ▼
  ┌─────────────────────────────────────────┐
  │  STEP 5 — Post-Processing & Logging    │
  │  "Clean output, log the request,       │
  │   return the result to the user."      │
  └───────────────────┬─────────────────────┘
                      │
                      ▼
  User receives: "Here is your summary: ..."

  ✅ All 5 steps happen inside ONE application.
     Fast to build. Simple to debug. Easy to understand.

Monolith Scaling: The Vertical Wall 🧱

When a monolith gets too slow or too busy, you have one option: make the whole machine bigger. This is called vertical scaling.


  MONOLITH SCALING PROBLEM:
  ─────────────────────────────────────────────────────────────────

  Traffic doubles → You must upgrade the ENTIRE server:
    Add more GPU memory  ← because inference is slow
    Add more CPU         ← even though auth barely uses it
    Add more RAM         ← even though logging barely uses it
    Pay for ALL of it    ← even the parts that aren't busy!

  Example situation:
  ┌─────────────────────────────────────────────────────────────┐
  │  Your inference is using 95% of GPU — it's the bottleneck!  │
  │  Your auth uses only 2% of CPU.                             │
  │  Your logger uses only 1% of CPU.                           │
  │                                                             │
  │  To scale inference → you must buy a BIGGER server          │
  │  that scales auth and logging too — even though they        │
  │  don't need it. You're wasting money! 💸                    │
  └─────────────────────────────────────────────────────────────┘

  Also: There is a physical ceiling!
  At some point, no single machine is big enough.
✅ DO — Use Monolithic Architecture When:
  • You are building a prototype, MVP, or internal demo
  • Your team is small — 1 to 3 engineers
  • Traffic is low and predictable — under a few thousand requests per day
  • You are serving one model type with no plans for multi-model routing
  • Speed of shipping matters more than ability to scale infinitely
❌ DON'T — Avoid Monolith When:
  • You need to scale only the GPU inference without scaling auth and logging too
  • Multiple teams each own a different part of the pipeline
  • A bug in one component (like logging) can crash inference for all users
  • You need to update one component without redeploying the entire application
  • Traffic spikes require fast autoscaling of specific parts independently
⚠️ The Monolith Trap:

Teams often start with a monolith — which is smart! But they never migrate away from it. They keep adding features to the same codebase until it becomes an unmaintainable mess where touching one thing breaks another. Engineers call this the "Big Ball of Mud."

The fix: plan your migration path before you desperately need it — not after! 🕐

Part 2 — Microservices Model Serving 🧩

What is Microservices Architecture?

In a microservices serving system, each concern gets its own separate, independently running service.

Auth is one service. Pre-processing is another service. Inference runs in its own service. Logging has its own container. Post-processing runs independently.

They talk to each other through clear, well-defined communication channels. Each service can be scaled independently, updated independently, and fail independently — without bringing down the others.


  ┌────────────┐
  │   Client   │
  └─────┬──────┘
        │  sends request
        ▼
  ┌─────────────────────────────────────────────────────────────┐
  │                    API GATEWAY                              │
  │          (single front door for all requests)               │
  └──────┬──────────────────────────┬──────────────────────────┘
         │                          │
         │ routes to                │ routes to
         ▼                          ▼
  ┌──────────────┐          ┌──────────────────┐
  │ AUTH SERVICE │          │  RATE LIMIT SVC  │
  │              │          │                  │
  │ "Is this     │          │ "Has this user   │
  │  user valid?"│          │  exceeded quota?"│
  └──────┬───────┘          └──────────────────┘
         │ (if valid, continue)
         ▼
  ┌──────────────────────────────────────────┐
  │          PRE-PROCESSING SERVICE          │
  │  Tokenize · Format · Validate input      │
  └──────────────────┬───────────────────────┘
                     │
                     ▼
  ┌──────────────────────────────────────────┐
  │      INFERENCE SERVICE  🖥️ (GPU pods)    │
  │                                          │
  │  Pod 1 │  Pod 2 │  Pod 3  ← auto-scaled  │
  │  (GPU)   (GPU)    (GPU)     by traffic   │
  └──────────────────┬───────────────────────┘
                     │
                     ▼
  ┌──────────────────────────────────────────┐
  │         POST-PROCESSING SERVICE          │
  │  Format · Safety filter · Truncate       │
  └──────────────────┬───────────────────────┘
                     │
                     ▼
              Response to Client ✅

How Does a Microservices System Handle a Request?

The same request that a monolith handles in one place now travels through multiple independent services. Think of it as an assembly line in a factory:


  User sends: "Summarize this article for me"
        │
        ▼
  [API Gateway]          — "I'll route this to the right services."
        │
        ▼
  [Auth Service]         — "Valid user? Yes. Tier: Pro. Proceed!"
        │
        ▼
  [Pre-Processing Svc]   — "Formatted prompt ready. Token count: 450."
        │
        ▼
  [Inference Service]    — "Running on GPU Pod #2. Output generated!"
  (Only this service     — Pod #1 and Pod #3 are serving other users
   touches the GPU)        at the same time → parallel processing!
        │
        ▼
  [Post-Processing Svc]  — "Cleaned up. Safety check passed. Done."
        │
        ▼
  User receives: "Here is your summary: ..."

  Each arrow = a network call between two independent services.
  Each box = its own server, its own code, its own scaling rules.

The Superpower: Independent Scaling ⚡

This is the biggest reason teams choose microservices. Each service scales exactly as much as it needs to — and no more.


  SCENARIO: Black Friday traffic spike — 50x normal load
  ──────────────────────────────────────────────────────────────────

  MONOLITH RESPONSE:
  ┌──────────────────────────────────────────────────────────────┐
  │  Scale the whole server 50x                                  │
  │  50x GPU    ← needed (inference is the bottleneck)           │
  │  50x CPU    ← NOT needed (auth is barely busy)               │
  │  50x RAM    ← NOT needed (logging is barely busy)            │
  │  50x cost   ← you pay for all of it                         │
  └──────────────────────────────────────────────────────────────┘

  MICROSERVICES RESPONSE:
  ┌──────────────────────────────────────────────────────────────┐
  │  Inference Service  → scale from 3 to 150 GPU pods ✅        │
  │  Auth Service       → scale from 3 to 10 tiny CPU pods ✅    │
  │  Pre-Process Svc    → scale from 2 to 8 CPU pods ✅          │
  │  Post-Process Svc   → scale from 2 to 8 CPU pods ✅          │
  │  Logging Service    → stays the same — async, not a blocker  │
  │                                                              │
  │  Result: You pay for exactly what you need, nothing more! 💡 │
  └──────────────────────────────────────────────────────────────┘
✅ DO — Use Microservices When:
  • You need GPU inference to scale independently from cheap CPU-only services
  • Multiple teams each own different parts of the pipeline and deploy separately
  • You need to update the safety filter without touching the model inference code
  • Different parts of the system have different reliability requirements
  • You are building a platform that will serve multiple different models over time
  • Traffic patterns are unpredictable and you need fine-grained autoscaling
❌ DON'T — Common Microservices Mistakes:
  • Don't split into microservices too early — premature microservices are painful
  • Don't forget timeouts on calls between services — one slow service can freeze the whole chain
  • Don't share a single database between services — each service must own its own data
  • Don't skip health checks — every service must know when its neighbors are alive or dead
  • Don't ignore tracing — debugging across 5 services without logs is a nightmare

Part 3 — The Hidden Cost: Network Latency ⏱️

Microservices have one unavoidable downside that monoliths don't have: every service-to-service call adds time.


  MONOLITH: All steps happen inside one process
  ─────────────────────────────────────────────────────────────────
  Auth check:          0.1ms  (in-memory function call)
  Pre-processing:      2ms    (in-memory function call)
  Inference:         500ms    (GPU computation — same for both)
  Post-processing:     1ms    (in-memory function call)
  ─────────────────
  Total extra:         3.1ms  (the non-GPU parts are near-zero cost)


  MICROSERVICES: Each step = a network round trip
  ─────────────────────────────────────────────────────────────────
  Gateway → Auth:        1–3ms    (network call)
  Auth → Pre-process:    1–3ms    (network call)
  Pre-process → Infer:   1–3ms    (network call)
  Inference:           500ms      (GPU computation — same as monolith)
  Infer → Post-process:  1–3ms    (network call)
  ─────────────────
  Total extra:          5–15ms    (added by network hops)

  Verdict: For LLMs where inference takes 200–2000ms,
           the 5–15ms of microservices overhead is acceptable.
           For sub-10ms latency use cases, it matters more.
⚠️ The Latency Rule of Thumb:

If your model takes more than 100ms to generate a response, microservices network overhead is negligible — go for it!

If you need responses in under 20ms (real-time autocomplete, gaming, etc.), every millisecond counts — consider keeping critical hot paths in the monolith.

Part 4 — Key Patterns in Microservices Serving 🔧

These are the five most important patterns you will encounter when building or working with microservices for LLM serving. You don't need to implement them from scratch — but understanding what they do is essential.

Pattern 1: The API Gateway 🚪

The API Gateway is the single front door to your entire system. External users only ever talk to the gateway. It routes requests to the right internal service.


  WITHOUT Gateway:
  ────────────────────────────────────────────────────────────────
  User → Auth Service     (user must know the auth URL)
  User → Inference Svc    (user must know the inference URL)
  User → Logging Svc      (user must know the logging URL)
  Problem: Too many URLs to manage. Security nightmare.


  WITH Gateway:
  ────────────────────────────────────────────────────────────────
  User → Gateway → Auth Service
                → Inference Service
                → Logging Service

  User only ever knows ONE URL.
  Gateway handles all the routing internally.
  Internal services are hidden and protected.

  The gateway also handles:
    🔒 TLS / HTTPS termination
    ⚖️  Load balancing across service instances
    🚦 Rate limiting (before requests even reach services)
    📋 Request logging and metrics at the entry point

Pattern 2: The Circuit Breaker 🔌

What happens if the inference service crashes? Without a circuit breaker, every incoming request waits in a queue, piling up until the whole system runs out of memory and crashes too.

The circuit breaker stops sending requests to a failing service and immediately returns a friendly error — just like a fuse in your home's electrical system that trips before a wire fire can start.


  CIRCUIT BREAKER — Three States:
  ─────────────────────────────────────────────────────────────────

  State 1: CLOSED (normal)
  ┌────────────────────────────────────────────────────────────┐
  │  Requests flow through normally.                           │
  │  Circuit monitors: "Is the inference service responding?"  │
  │  If 5 consecutive failures detected → trip to OPEN state.  │
  └────────────────────────────────────────────────────────────┘

  State 2: OPEN (tripped — failing fast)
  ┌────────────────────────────────────────────────────────────┐
  │  Immediately return error: "Service unavailable."          │
  │  Do NOT send requests to the inference service at all.     │
  │  Wait 30 seconds. Then move to HALF-OPEN state.            │
  └────────────────────────────────────────────────────────────┘

  State 3: HALF-OPEN (testing recovery)
  ┌────────────────────────────────────────────────────────────┐
  │  Let ONE test request through to the inference service.    │
  │  If it succeeds → reset to CLOSED (service recovered!).   │
  │  If it fails    → go back to OPEN (still broken).          │
  └────────────────────────────────────────────────────────────┘

  Result: A single failing service NEVER takes down the whole system!

Pattern 3: The Async Queue 📬

For heavy tasks like generating a long report or processing a large document, you don't want the user to sit with an open connection waiting for 2 minutes. The async queue pattern solves this elegantly.


  SYNCHRONOUS (bad for long tasks):
  ──────────────────────────────────────────────────────────────
  User sends request → waits → waits → waits (2 minutes!) → gets result
  If connection drops in those 2 minutes → result is lost!


  ASYNC QUEUE (great for long tasks):
  ──────────────────────────────────────────────────────────────

  Step 1: User sends request
          │
          ▼
  [Gateway] puts job in queue → returns immediately with:
          "Your job ID is abc-123. Check back in a moment!"
          (User gets a response in under 1 second!)

  Step 2: In the background:
  [Queue Worker] picks up the job
  [Inference Service] runs the model (takes 2 minutes — that's fine now!)
  [Worker] stores result with job ID: "abc-123 → completed"

  Step 3: User polls for result:
  User asks: "What's the status of job abc-123?"
  [Gateway] checks: "Completed! Here is your result."

  Benefits:
  ✅ User gets instant acknowledgement
  ✅ Connection drops don't lose work
  ✅ System can process many jobs in parallel
  ✅ Peak traffic is smoothed out — queue absorbs the spike

Pattern 4: The Sidecar 🏍️

A sidecar is a small helper that runs alongside every service — handling things like logging, security, and metrics collection — so the main service doesn't have to think about those concerns at all.


  WITHOUT Sidecar:
  ──────────────────────────────────────────────────────────────
  Inference Service code must handle:
    → Running the model
    → Collecting metrics
    → TLS encryption
    → Request logging
    → Retry logic
  Result: 500-line code file mixing business logic with infra concerns!


  WITH Sidecar:
  ──────────────────────────────────────────────────────────────
  ┌──────────────────────────────────────────────────────────┐
  │                    ONE DEPLOYMENT UNIT                    │
  │                                                          │
  │  ┌─────────────────────┐  ┌──────────────────────────┐  │
  │  │   Main Container    │  │   Sidecar Container      │  │
  │  │  (Inference Svc)    │◀─▶  (Envoy / Istio proxy)   │  │
  │  │                     │  │                          │  │
  │  │  Just runs the model│  │  Handles:                │  │
  │  │  → Clean & focused  │  │  → TLS encryption        │  │
  │  │  → 50 lines of code │  │  → Metrics collection    │  │
  │  │                     │  │  → Circuit breaking      │  │
  │  │                     │  │  → Request tracing       │  │
  │  └─────────────────────┘  └──────────────────────────┘  │
  └──────────────────────────────────────────────────────────┘

  Result: Your service code focuses ONLY on what it does.
          The sidecar handles ALL the infrastructure concerns!

Pattern 5: The Strangler Fig — Migrating Safely 🌳

Named after the strangler fig tree that slowly wraps around and eventually replaces another tree. This pattern lets you extract services from a monolith one piece at a time — while keeping the monolith running the whole time. No risky big-bang rewrites!


  PHASE 1: Everything in the monolith
  ─────────────────────────────────────────────────────────────────
  All traffic → [ Monolith: Auth + Pre + Inference + Post ]


  PHASE 2: Extract the INFERENCE service first
  (it's the most expensive and most worth isolating)
  ─────────────────────────────────────────────────────────────────
                           ┌── /v1/infer → [New Inference Svc] ──┐
  All traffic → [Proxy] ───┤                                      │
                           └── everything else → [Monolith] ─────┘


  PHASE 3: Extract AUTH and PRE-PROCESSING
  ─────────────────────────────────────────────────────────────────
                    ┌── [Auth Svc]        ──┐
  Traffic → [Proxy]─┤  [Pre-Process Svc]  ──┤→ [Inference Svc] → [Shrinking Monolith]
                    └── [Logging Svc]     ──┘


  PHASE 4: Monolith is fully replaced — all microservices!
  ─────────────────────────────────────────────────────────────────
  Traffic → [Gateway] → [Auth] → [Pre] → [Inference] → [Post] → User

  Key rule: At every phase, the system keeps working for real users!
            Zero downtime. Zero panic. Gradual, safe migration. ✅

Part 5 — Observability: Seeing Inside Your System 🔭

A monolith has one log file. One place to look when something breaks. Microservices have 5, 10, or 20 services — each with their own logs.

Without proper observability, debugging microservices is like trying to find a single faulty wire in a city's entire power grid with no map and no tools. 😱

The three pillars of microservices observability are:


  THE THREE PILLARS OF OBSERVABILITY
  ─────────────────────────────────────────────────────────────────

  PILLAR 1: LOGS 📋
  ────────────────────────────────────────────────────────────────
  What: Every service writes structured events to a central log system.
  Why: "At 14:32:05, auth service rejected API key xyz-789."
  Tools: ELK Stack (Elasticsearch + Logstash + Kibana), Loki + Grafana

  Format every service must use (structured logging):
  {
    "timestamp": "2025-02-25T14:32:05Z",
    "service":   "auth-service",
    "level":     "ERROR",
    "message":   "Invalid API key",
    "request_id": "abc-123",          ← same ID across ALL services!
    "user_id":   "user-456",
    "latency_ms": 2
  }


  PILLAR 2: METRICS 📊
  ────────────────────────────────────────────────────────────────
  What: Numbers that track the health of each service over time.
  Why: "Inference latency jumped from 500ms to 3000ms at 2pm."
  Tools: Prometheus (collects metrics) + Grafana (displays dashboards)

  Key metrics every LLM service should track:
  → Requests per second (throughput)
  → p50 / p95 / p99 latency (how long requests take)
  → Error rate (% of requests that fail)
  → GPU utilization % (is the inference service working hard?)
  → Queue depth (how many requests are waiting?)
  → Tokens generated per second (LLM-specific throughput)


  PILLAR 3: DISTRIBUTED TRACES 🗺️
  ────────────────────────────────────────────────────────────────
  What: A single request is tracked as it flows through ALL services.
  Why: "Request abc-123 took 2 seconds. Auth: 3ms. Pre: 5ms.
        Inference: 1,800ms. Post: 2ms. Bottleneck: clearly inference!"
  Tools: Jaeger, Zipkin, Tempo + Grafana

  A trace looks like this (a waterfall/timeline):
  Request abc-123 (total: 1,810ms)
  │
  ├─ Auth Service          ██  3ms
  ├─ Pre-Processing        ████  5ms
  ├─ Inference Service     ████████████████████████████████  1,800ms  ← bottleneck!
  └─ Post-Processing       ██  2ms
⚠️ The Request ID Rule — Never Skip This!

Every request that enters your system must be assigned a unique ID (called a trace ID or correlation ID). This ID is passed from service to service in every call.

When a user reports "my request was slow at 2:34pm," you search your logs for their trace ID and instantly see exactly what every service did for that specific request. Without this, debugging microservices is nearly impossible! 🔍

Part 6 — Real-World Architecture Examples 🏢

Let's look at how real teams at different stages build their serving systems. These are realistic examples, not theoretical ones.

Example A: The Startup (0 → 1,000 users/day)


  ARCHITECTURE: Simple Monolith
  ─────────────────────────────────────────────────────────────────

  Stack:
  → One FastAPI server (Python)
  → One GPU cloud instance (e.g., AWS g4dn.xlarge, ~$0.50/hr)
  → vLLM for fast inference (still inside the monolith)
  → Simple API key check in the same code

  Deployment:
  → One Docker container
  → One command to deploy

  ┌───────────────────────────────────────────────┐
  │               SINGLE SERVER                   │
  │                                               │
  │  FastAPI app:                                 │
  │    ├── API handler                            │
  │    ├── Auth logic                             │
  │    ├── Prompt formatter                       │
  │    ├── vLLM inference engine (GPU)            │
  │    ├── Output cleaner                         │
  │    └── Request logger                         │
  │                                               │
  │  Cost: ~$360/month for 24/7 uptime            │
  │  Team needed: 1 engineer                      │
  └───────────────────────────────────────────────┘

  💡 This is the RIGHT call at this stage.
     Ship fast. Validate the product. Don't over-engineer!

Example B: The Growing Company (10,000 → 100,000 users/day)


  ARCHITECTURE: Partial Microservices (Strangler Fig in progress)
  ─────────────────────────────────────────────────────────────────

  What changed and why:
  → Inference was the bottleneck — extracted to its own scalable service
  → Auth is hit on EVERY request — separated so it can cache using Redis
  → Logging separated so a logging bug never touches inference
  → Gateway added to route between services

  ┌────────────┐    ┌───────────────────┐
  │  Gateway   │───▶│   Auth Service    │ ← Redis cache, scales to 10 pods
  │            │    │   (CPU only)      │
  │            │    └───────────────────┘
  │            │
  │            │    ┌───────────────────┐    ┌──────────────────────────┐
  │            │───▶│  Pre-Process Svc  │───▶│   Inference Service      │
  │            │    │  (CPU only)       │    │   3 GPU pods, auto-scale  │
  └────────────┘    └───────────────────┘    │   to 10 on high traffic  │
                                             └──────────────────────────┘

  Cost: Inference pods: $1,200/month
        Auth + Pre + Post pods: $80/month (CPU is cheap!)
        Savings vs. scaling the monolith: ~40% cost reduction

  Team needed: 2–4 engineers (one DevOps-focused)

Example C: The Platform (Millions of users/day)


  ARCHITECTURE: Full Microservices on Kubernetes
  ─────────────────────────────────────────────────────────────────

  ┌─────────────────────────────────────────────────────────┐
  │                  Kubernetes Cluster                      │
  │                                                         │
  │  ┌───────────┐  ┌───────────┐  ┌─────────────────────┐ │
  │  │ Gateway   │  │   Auth    │  │  Rate Limiter (Redis)│ │
  │  │ (nginx /  │  │  Service  │  │                     │ │
  │  │  Kong)    │  │ x5 pods   │  └─────────────────────┘ │
  │  └─────┬─────┘  └───────────┘                          │
  │        │                                                │
  │  ┌─────▼─────────────────────────────────────────────┐ │
  │  │           Pre-Processing Service (x8 pods)         │ │
  │  └─────┬─────────────────────────────────────────────┘ │
  │        │                                                │
  │  ┌─────▼─────────────────────────────────────────────┐ │
  │  │     Model Router — picks the right model          │ │
  │  │     (GPT-4 equivalent / 7B model / vision model)  │ │
  │  └──┬─────────────┬───────────────┬──────────────────┘ │
  │     │             │               │                     │
  │  ┌──▼───┐    ┌────▼───┐    ┌──────▼──────────────────┐ │
  │  │LLM-A │    │ LLM-B  │    │  Vision Model Service   │ │
  │  │Svc   │    │  Svc   │    │  (image + text models)  │ │
  │  │x20   │    │  x10   │    │  x5 GPU pods            │ │
  │  │GPU   │    │  GPU   │    └─────────────────────────┘ │
  │  │pods  │    │  pods  │                                 │
  │  └──────┘    └────────┘                                 │
  │                                                         │
  │  ┌───────────────────────────────────────────────────┐ │
  │  │  Post-Processing + Safety Filter Service (x8 pods)│ │
  │  └───────────────────────────────────────────────────┘ │
  │                                                         │
  │  ┌───────────────┐  ┌───────────────┐  ┌────────────┐  │
  │  │ Observability │  │  Async Queue  │  │  Billing   │  │
  │  │ (Prometheus + │  │  (Redis/Kafka)│  │  Service   │  │
  │  │  Grafana +    │  │               │  │            │  │
  │  │  Jaeger)      │  └───────────────┘  └────────────┘  │
  │  └───────────────┘                                      │
  └─────────────────────────────────────────────────────────┘

  Team needed: 10+ engineers with dedicated platform/DevOps team

Part 7 — Managed Platforms: The Sweet Middle Ground 🌟

Not ready for full Kubernetes microservices, but outgrowing a simple monolith? Managed serving platforms give you microservices-like benefits without all the complexity of running your own Kubernetes cluster.


  COMPLEXITY SPECTRUM — Pick your position honestly:
  ─────────────────────────────────────────────────────────────────────────
  Simplest                                                    Most Powerful
  ────────────────────────────────────────────────────────────────────────▶

  │              │               │               │              │
  Plain        vLLM +          Ray Serve /    Docker +       Full
  FastAPI      thin wrapper    BentoML        Compose        Kubernetes
  Monolith                                                   Microservices

  │              │               │               │              │
  Easiest       Great LLM       Structured      Good for       Production
  to start      performance     approach        small teams    at scale
  No scaling    Some auto-      Built-in        Limited        Full
  options       scaling         scaling         scaling        flexibility

Which Managed Platform Should You Choose?

  • 🟡 Ray Serve → Python-native. You define your services in Python code — no Kubernetes YAML required. Ray handles deployment, autoscaling, and traffic routing automatically. Each service (auth, inference, post-processing) scales independently on the same cluster. Best for: Python teams who understand distributed systems but not K8s.
  • 🟢 BentoML → Package your entire pipeline — pre-processing, model, post-processing — as a "Bento." One command builds and deploys the whole thing to any cloud. Handles batching, multi-model routing, and containerization automatically. Best for: Teams who want structure without writing infrastructure code.
  • 🔵 Triton Inference Server (NVIDIA) → Maximum GPU throughput for production workloads. Think of it as a battle-hardened inference-only microservice that is already built for you. Supports PyTorch, TensorRT, ONNX, and more. Best for: When performance is the only goal and you have an ops team.
  • 🟠 vLLM with a thin API wrapper → The most popular path for LLM serving right now. vLLM handles the hard parts — KV caching, continuous batching, paged attention. You just write a thin wrapper on top. Easiest upgrade from a monolith with massive performance gains. Best for: Anyone serving LLMs who wants fast inference without complexity.
⚠️ Practical Advice for Beginners:

If you are just starting out with LLM serving: start with vLLM + a simple FastAPI wrapper. It is free, powerful, and widely used in industry. You can serve a 7B model with streaming, batching, and GPU efficiency in under an hour — without needing to understand Kubernetes at all.

Graduate to Ray Serve or BentoML when you outgrow the single server. Graduate to full Kubernetes only when you have a dedicated infrastructure team.

Part 8 — Head-to-Head Comparison 🥊

Let's compare both architectures across every dimension that matters when making a real architectural decision:


  Dimension              Monolith                  Microservices
  ─────────────────────────────────────────────────────────────────────────
  Setup time             Hours                     Days to weeks
  Team size needed       1–3 engineers             3+ engineers + DevOps
  Deployment             One command               CI/CD pipeline per service
  Scaling approach       Vertical (bigger server)  Horizontal (more pods)
  GPU cost efficiency    Poor (pays for GPU x all) Excellent (GPU pods only)
  Fault isolation        None (crash = all down)   High (services fail alone)
  Updating one part      Redeploy everything        Deploy just that service
  Debugging              Simple (one log file)     Complex (need tracing)
  Network overhead       Zero (in-process calls)   5–15ms per service hop
  Cost at low traffic    Low                       Higher (many containers)
  Cost at high traffic   High (over-provisioned)   Lower (precise scaling)
  Monitoring             Simple                    Requires Prometheus+Grafana
  Security surface       Small (one application)   Larger (many endpoints)
  Right for              MVP · prototype · one team Platform · many teams

Cost Comparison: A Concrete Scenario


  Scenario: 50,000 inference requests per day
  Model: 7B parameter LLM
  Traffic pattern: heavy 9am–9pm, near-zero overnight

  ─────────────────────────────────────────────────────────────────────────
  MONOLITH APPROACH:
  ─────────────────────────────────────────────────────────────────────────
  Need 3 GPU servers to handle peak load.
  Must run ALL 3 servers 24/7 (can't scale down overnight — monolith!).

  3 servers × $4/hr × 24 hours × 30 days = $8,640/month

  But at 2am you have near-zero traffic → those 3 GPU servers idle!
  You're paying $8,640 for ~12 peak hours of real usage.


  MICROSERVICES APPROACH:
  ─────────────────────────────────────────────────────────────────────────
  Inference pods: scale 1 → 10 at 9am, scale back to 1 at 9pm
    → 10 pods × $4/hr × 12 peak hours = $1,440
    →  1 pod  × $4/hr × 12 off hours  = $144
    Total inference: $1,584/month

  Auth + Pre + Post (CPU pods, tiny cost):
    → Always-on, but cheap: ~$100/month

  Total microservices cost: ~$1,684/month

  Savings: $8,640 - $1,684 = $6,956/month saved! 💰
  That's $83,472 saved per year just from architectural choice.
⚠️ The Real Power of Microservices:

The biggest savings don't come from raw compute costs. They come from being able to scale to zero overnight when traffic drops.

A monolith runs 24/7 because you can never know which component will get a request. A microservices system lets GPU pods scale to zero at 2am and spin back up in 60 seconds when the morning traffic arrives. This alone can save thousands of dollars per month at scale.

Part 9 — Decision Framework: What Should YOU Build? 🎯

Use this decision tree every time you start a new model serving project. Answer each question honestly!


  START HERE
      │
      ▼
  Is this a proof-of-concept, demo, or internal tool?
      │
      ├── YES ──▶  MONOLITH ✅
      │            (Fast to build, easy to change, no ops overhead)
      │
      └── NO ──▶  How many users per day do you expect?
                      │
                      ├── Under 1,000/day ──▶  MONOLITH ✅
                      │                         (Still perfectly right!)
                      │
                      └── Over 10,000/day ──▶  Do you have DevOps support?
                                                    │
                                                    ├── NO ──▶  Managed Platform ✅
                                                    │           (vLLM / Ray Serve / BentoML)
                                                    │
                                                    └── YES ──▶  Do multiple teams
                                                                 own different parts?
                                                                     │
                                                                     ├── NO ──▶  Managed Platform
                                                                     │           or Modular Monolith
                                                                     │
                                                                     └── YES ──▶  MICROSERVICES ✅
                                                                                  (Kubernetes)
✅ The Golden Rule of Architecture:

Start with the simplest thing that could possibly work. That is almost always a monolith. Only add complexity when you can clearly measure the specific problem that complexity solves.

The best engineers are not the ones who build the most complex systems. They are the ones who keep systems as simple as possible for as long as possible, and migrate thoughtfully when the evidence demands it.

❌ The Most Common LLM Engineering Mistake:

Jumping straight to microservices because it "sounds more professional." Engineers spend weeks configuring Kubernetes, service meshes, and distributed tracing — for an LLM product that currently serves 50 users per day.

This kills momentum and delays shipping. Build the monolith. Ship it. Validate the idea. Migrate when the data tells you to.

Quick Summary 📝

What we learned today:

  • Model Serving → The full system around your model: API, auth, pre-processing, inference, post-processing, and logging
  • Monolithic Architecture → One codebase, one deployment, one server. Simple to build and debug. Hard to scale individual components independently. Perfect for prototypes, small teams, and early products.
  • Microservices Architecture → Each concern is its own service, deployed and scaled independently. Complex to set up, powerful at scale. Essential when GPU inference must scale separately from cheap CPU services.
  • API Gateway Pattern → Single front door that hides all internal services from the outside world
  • Circuit Breaker Pattern → Stops sending requests to a failing service so the whole system does not freeze
  • Async Queue Pattern → Decouple long-running inference from HTTP — user gets a job ID and polls for results
  • Sidecar Pattern → Helper container handles logging, TLS, metrics — keeps your main service code clean
  • Strangler Fig Migration → Extract services one at a time from a live monolith with zero downtime
  • Observability (Logs + Metrics + Traces) → The three pillars that let you see what every service is doing in production
  • Managed Platforms → vLLM, Ray Serve, BentoML give microservices-level performance without full Kubernetes complexity

Architecture decisions compound over time. A thoughtful decision made early saves thousands of engineering hours later. A rushed decision made under pressure costs millions in rework and downtime. Now you have the knowledge and the framework to choose wisely — every single time! 🏗️✨

Comments