Skip to main content

Docker Image to Inference Endpoint: Build and Deploy ML Models

Calculating read time…

Today we are going to learn something that most tutorials skip entirely —
how to actually deploy your AI app so real users in the world can use it. 🌍

We will build a small, real app together — step by step — and then ship it like a professional engineer.
By the end, you will know how to create a Docker image, generate an inference endpoint, and understand all the modern deployment tricks used.

💡 What are we building today?

A tiny but powerful app called MeetingMuse — it takes your messy Zoom meeting notes and turns them into a clean, professional story using OCI Generative AI. 🎙️✨

We will then learn how to package it, deploy it, and share it with the world.


🗺️ The Big Picture — What are we doing?

Think of it like baking a cake 🎂 and opening a bakery shop.
First, you perfect your recipe (write the app).
Then, you pack it in a nice box (Docker image).
Then, you open a counter where customers can order it (inference endpoint).
Simple, right?

  Your Code (MeetingMuse App)
         │
         ▼
  ┌─────────────────────────┐
  │   Docker Image          │  ◄── Box that holds everything the app needs
  │   (your app + Python    │
  │    + libraries packed   │
  │    together)            │
  └────────────┬────────────┘
               │
               ▼
  ┌─────────────────────────┐
  │   Inference Endpoint    │  ◄── The "front door" users knock on
  │   (a URL like           │
  │    https://api.you.com/ │
  │    summarize)           │
  └────────────┬────────────┘
               │
               ▼
         User sends Zoom notes
         → Gets back a polished story ✨

🏗️ Step 1 — Build the App (MeetingMuse)

Let us first understand what our app will do before we write a single line of code.

The job of MeetingMuse:
A user pastes raw, messy Zoom meeting notes into our app.
Our app sends those notes to OCI Generative AI.
OCI GenAI thinks hard and returns a polished, clear summary written like a story.
We send that story back to the user.

✅ Folder Structure — How we organize our files:

Good organization is like keeping your school bag tidy.
Everything has its right place so you never lose anything.
meetingmuse/
├── app.py              ← The brain of our app (FastAPI server)
├── requirements.txt    ← List of Python libraries we need
├── Dockerfile          ← The recipe to build our Docker box
├── .env                ← Secret keys (never share this!)
└── README.md           ← Instructions for other developers

📄 File 1 — requirements.txt

📝 What does this file do?

Imagine you are moving to a new house and you write a shopping list:
"I need a bed, a fridge, a table..."

requirements.txt is exactly that — a shopping list for Python.
When someone runs your app for the first time, Python reads this list and automatically downloads everything your app needs. Without this list, the app would crash because it would not find the tools it needs to work.
fastapi==0.110.0
uvicorn==0.29.0
oci==2.126.0
python-dotenv==1.0.1
pydantic==2.6.4

fastapi — The web framework that creates our API server (the "waiter" of our app).
uvicorn — The engine that actually runs the FastAPI server.
oci — Oracle's official Python library to talk to OCI GenAI.
python-dotenv — Reads our secret keys from the .env file safely.
pydantic — Makes sure the data coming into our app is in the right shape.


📄 File 2 — app.py (The Brain)

📝 What does this file do?

This is the main file — the actual brain of MeetingMuse.
It does three things:
1. Listens for incoming requests from users (like a receptionist).
2. Takes the user's messy notes and sends them to OCI GenAI with a smart prompt.
3. Gets back the polished story and returns it to the user.
# ── Import the tools we need ──────────────────────────────────────────
import os
from fastapi import FastAPI
from pydantic import BaseModel
from dotenv import load_dotenv
import oci

# Load secret keys from our .env file
load_dotenv()

# ── Create the FastAPI app ─────────────────────────────────────────────
app = FastAPI(title="MeetingMuse", version="1.0")

# ── Define what a user request looks like ─────────────────────────────
# Think of this like a form — the user must fill in the 'notes' field
class MeetingRequest(BaseModel):
    notes: str   # The raw, messy Zoom meeting notes

# ── Set up our connection to OCI GenAI ────────────────────────────────
def get_oci_client():
    config = oci.config.from_file()   # Reads ~/.oci/config on your machine
    return oci.generative_ai_inference.GenerativeAiInferenceClient(config)

# ── The main endpoint — what happens when a user sends notes ──────────
@app.post("/summarize")
def summarize_meeting(request: MeetingRequest):

    client = get_oci_client()

    # This is the instruction we give to the AI model
    prompt = f"""
    You are a professional storyteller and executive assistant.
    Below are raw, unstructured notes from a Zoom meeting.
    Transform them into a clean, professional narrative summary.
    Use clear headings, highlight key decisions, action items,
    and write it in a warm but business-appropriate tone.

    --- RAW NOTES START ---
    {request.notes}
    --- RAW NOTES END ---

    Write the polished summary now:
    """

    # Build the request body for OCI GenAI (Cohere Command R+ model)
    chat_detail = oci.generative_ai_inference.models.CohereChatRequest()
    chat_detail.message = prompt
    chat_detail.max_tokens = 1024
    chat_detail.temperature = 0.5   # 0 = very focused, 1 = very creative

    chat_request = oci.generative_ai_inference.models.ChatDetails()
    chat_request.serving_mode = oci.generative_ai_inference.models.OnDemandServingMode(
        model_id="cohere.command-r-plus"
    )
    chat_request.compartment_id = os.getenv("OCI_COMPARTMENT_ID")
    chat_request.chat_request = chat_detail

    # Send to OCI GenAI and get the response
    response = client.chat(chat_request)
    summary = response.data.chat_response.text

    return {"polished_summary": summary}

# ── Health check endpoint ─────────────────────────────────────────────
@app.get("/health")
def health():
    return {"status": "MeetingMuse is alive and ready! 🎙️"}
✅ Quick Test (on your laptop first):

Before we package anything, always make sure the app works locally.
Run: uvicorn app:app --reload
Then open your browser and go to http://localhost:8000/health
If you see "MeetingMuse is alive" — you are ready to proceed! 🟢

📦 Step 2 — Create a Docker Image (From Scratch)

Now here is where it gets really exciting.
You have a working app on your laptop. But how do you give it to someone else
without worrying about "it works on my machine but not yours"?

The answer is Docker. 🐳

🤔 What is Docker? (Explained Like You're 10)

Imagine you made the world's best sandwich. 🥪
You want to send the exact same sandwich to your friend in another city.
But if you just send the recipe, they might use different bread, different butter... and the taste changes.

Docker is like a sealed lunchbox that contains the sandwich AND
the exact bread, exact butter, exact everything — already assembled.
Your friend opens the box and gets the identical sandwich. Every time.

  WITHOUT Docker:                    WITH Docker:
  ┌──────────────────┐               ┌──────────────────────┐
  │ Your App         │               │ Docker Image (Box)   │
  │ + Python 3.11    │   ─────►      │ ├─ Your App          │
  │ + Libraries      │               │ ├─ Python 3.11 ✓     │
  │ + OCI config     │               │ ├─ All Libraries ✓   │
  │                  │               │ └─ OCI config ✓      │
  │ "Works on MY     │               │                      │
  │  machine only"   │               │ "Works EVERYWHERE"   │
  └──────────────────┘               └──────────────────────┘

📄 The Dockerfile — The Recipe for Your Box

📝 What does the Dockerfile do?

A Dockerfile is like a cooking recipe, but for computers.
It tells Docker: "Start with this base, add these ingredients, cook it this way."
Docker reads this file and automatically builds your sealed lunchbox (the image).
# ── STAGE 1: Choose your base kitchen ────────────────────────────────
# We start with an official Python 3.11 image — a clean, minimal Linux
# environment that already has Python installed. The "slim" version
# means it's lightweight (faster to download and deploy).
FROM python:3.11-slim

# ── STAGE 2: Set the working folder inside the container ─────────────
# Think of this like choosing which room in the kitchen you will work in.
# All future commands will run from this folder.
WORKDIR /app

# ── STAGE 3: Copy the shopping list first (smart trick!) ─────────────
# We copy ONLY requirements.txt first (not the whole app).
# Why? Because Docker is smart — if requirements.txt hasn't changed,
# it skips reinstalling libraries next time. This saves LOTS of time.
COPY requirements.txt .

# ── STAGE 4: Install all Python libraries ────────────────────────────
# --no-cache-dir means don't save the downloaded files after installing.
# This keeps our final image size small and clean.
RUN pip install --no-cache-dir -r requirements.txt

# ── STAGE 5: Copy the rest of our app into the container ─────────────
# Now we bring in everything else — our app.py, .env, etc.
COPY . .

# ── STAGE 6: Tell Docker which port our app uses ─────────────────────
# Our FastAPI app listens on port 8000. This line documents that fact.
# It doesn't actually "open" the port — we do that when we run the image.
EXPOSE 8000

# ── STAGE 7: The startup command ─────────────────────────────────────
# This is what runs when someone starts our Docker container.
# uvicorn = the engine, app:app = our app.py file, 0.0.0.0 means
# "accept connections from any IP address" (important for cloud!).
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]

🔨 Building the Docker Image — Step by Step Commands

📝 What are we doing here?

These are the actual commands you type in your terminal (Command Prompt / VS Code terminal).
Each command does one specific job — like steps in a recipe.
Run them in order, from your project folder.
# ── Step A: Make sure Docker is installed and running ─────────────────
docker --version
# You should see something like: Docker version 25.x.x

# ── Step B: Build the Docker image ───────────────────────────────────
# -t means "tag" = give your image a name
# meetingmuse:v1 = name:version
# The dot (.) at the end means "use the Dockerfile in this folder"
docker build -t meetingmuse:v1 .

# ── Step C: See that your image was created successfully ──────────────
docker images
# You should see "meetingmuse" in the list

# ── Step D: Run it locally to test ───────────────────────────────────
# -p 8000:8000 = connect your laptop's port 8000 to the container's port 8000
# --env-file .env = pass your secret keys into the container
docker run -p 8000:8000 --env-file .env meetingmuse:v1

# ── Step E: Test it! Open a new terminal and run: ────────────────────
curl -X POST http://localhost:8000/summarize \
  -H "Content-Type: application/json" \
  -d '{"notes": "discussed Q3 targets, john said sales are up 20%, action: sarah to send report by friday"}'
✅ What a successful docker build looks like:

You will see Docker printing lines like:
Step 1/7 : FROM python:3.11-slim ✓
Step 2/7 : WORKDIR /app ✓
... and finally ...
Successfully built a1b2c3d4e5f6 🎉
That long random code is your image ID — like a fingerprint for your box.
🚫 Common Beginner Mistakes:

❌ Don't forget the dot (.) at the end of docker build — it tells Docker where to look.
❌ Don't add your .env file to Git — it contains secret keys!
Create a .gitignore file and add .env to it.
❌ Don't use --host localhost inside Docker — use 0.0.0.0 instead, otherwise no one outside the container can reach your app.

☁️ Step 3 — Push to Oracle Container Registry (OCIR)

Right now your Docker image lives only on your laptop.
To deploy it on OCI (the cloud), you need to first upload it to OCIR —
Oracle's Container Image Registry. Think of OCIR as a cloud storage locker for your images.

  Your Laptop              OCIR (Cloud Locker)         OCI Deployment
  ┌──────────┐    push     ┌──────────────┐   pull     ┌────────────┐
  │ Docker   │  ─────────► │ meetingmuse  │ ─────────► │ OKE / OCI  │
  │ Image    │             │ :v1 stored   │            │ Container  │
  └──────────┘             └──────────────┘            └────────────┘
📝 What do these commands actually do?

We are tagging our image with the OCIR address (like writing the delivery address on a package),
then pushing (uploading) it to the Oracle cloud storage locker.
# ── Step 1: Log in to OCIR ────────────────────────────────────────────
# Replace "ap-mumbai-1" with your OCI region
# Replace "your-tenancy" with your actual OCI tenancy namespace
docker login ap-mumbai-1.ocir.io -u your-tenancy/your-email@example.com

# ── Step 2: Tag your image with the OCIR address ──────────────────────
# Format: region.ocir.io/tenancy-namespace/repo-name:tag
docker tag meetingmuse:v1 ap-mumbai-1.ocir.io/your-tenancy/meetingmuse:v1

# ── Step 3: Push (upload) to OCIR ─────────────────────────────────────
docker push ap-mumbai-1.ocir.io/your-tenancy/meetingmuse:v1

# ── Step 4: Verify it is there ────────────────────────────────────────
# Go to OCI Console → Developer Services → Container Registry → see your image 🎉

🔌 Step 4 — Generate an Inference Endpoint

An inference endpoint is simply a URL — a web address —
that your users (or other apps) can send data to and get AI predictions back.

Think of it like a pizza shop telephone number. 📞
You call it (send a request), tell them what you want (your Zoom notes),
and they bring back the pizza (the polished summary).

Option A — Deploy on OCI Container Instances (Easiest for Beginners)

OCI Container Instances let you run your Docker image in the cloud with just a few clicks.
No servers to manage. No complicated setup. Perfect for beginners.

  FLOW: OCI Container Instance

  User sends notes
       │
       ▼
  https://your-ip:8000/summarize   ◄── This is your inference endpoint URL
       │
       ▼
  OCI Container Instance
  (runs your meetingmuse Docker image)
       │
       ▼
  Calls OCI GenAI API
       │
       ▼
  Returns polished story to user ✨
📝 Steps in OCI Console (clicks, not code):

1. Go to OCI Console → Developer Services → Containers → Container Instances
2. Click "Create Container Instance"
3. Choose your region, VCN, and subnet
4. Under "Containers" → add container → paste your OCIR image URL
5. Add environment variables (your OCI_COMPARTMENT_ID, etc.)
6. Set port mapping: 8000
7. Click Create → wait 2-3 minutes → copy the public IP
8. Your endpoint is ready: http://<public-ip>:8000/summarize 🎉

Option B — Deploy on OCI Kubernetes Engine (OKE) — For Scale

When your app grows and thousands of users hit it at the same time,
Container Instances alone may not be enough.
That's when Kubernetes (OKE) comes in.

Think of OKE like a fleet of food trucks 🚚 instead of one pizza shop.
If demand spikes, Kubernetes automatically opens more trucks. When demand drops, it closes them.
You only pay for what you use.

📝 What do these Kubernetes config files do?

A Deployment file tells Kubernetes: "Keep 2 copies of my app running always."
A Service file tells Kubernetes: "Give those 2 copies one single public URL."
Together, they create a resilient, scalable inference endpoint automatically.
# ── deployment.yaml ───────────────────────────────────────────────────
# This tells Kubernetes to always keep 2 copies of MeetingMuse running.
# If one crashes, Kubernetes automatically starts a new one. Zero downtime!

apiVersion: apps/v1
kind: Deployment
metadata:
  name: meetingmuse
spec:
  replicas: 2           # Run 2 copies (for reliability)
  selector:
    matchLabels:
      app: meetingmuse
  template:
    metadata:
      labels:
        app: meetingmuse
    spec:
      containers:
      - name: meetingmuse
        image: ap-mumbai-1.ocir.io/your-tenancy/meetingmuse:v1
        ports:
        - containerPort: 8000
        env:
        - name: OCI_COMPARTMENT_ID
          value: "ocid1.compartment.oc1..xxxxxx"
---
# ── service.yaml ──────────────────────────────────────────────────────
# This creates ONE public-facing URL (Load Balancer) that distributes
# incoming requests evenly across all 2 running copies of our app.

apiVersion: v1
kind: Service
metadata:
  name: meetingmuse-service
spec:
  type: LoadBalancer     # Creates a public IP address automatically
  selector:
    app: meetingmuse
  ports:
  - port: 80             # Users connect on port 80 (standard HTTP)
    targetPort: 8000     # Forwards to our app's port 8000 inside the container
# Apply both files to your OKE cluster
kubectl apply -f deployment.yaml
kubectl apply -f service.yaml

# Check that everything is running
kubectl get pods
kubectl get service meetingmuse-service

# Copy the EXTERNAL-IP from the service output
# Your inference endpoint is: http://EXTERNAL-IP/summarize 🎉

🚀 Step 5 — Other Best Deployment Techniques (Trend)

🔵 Technique 1 — OCI API Gateway (Add a Professional Front Door)

Right now your endpoint is a raw IP address. Not pretty. Not secure.
OCI API Gateway gives you a clean, professional URL with authentication, rate limiting, and HTTPS — all automatically.

  BEFORE API Gateway:
  http://129.213.45.67:8000/summarize   ← Ugly, insecure, hard to remember

  AFTER API Gateway:
  https://api.meetingmuse.com/v1/summarize   ← Clean, HTTPS, professional ✨

  API Gateway also gives you:
  ├── Authentication (only allowed users can call the API)
  ├── Rate Limiting (block users who send too many requests)
  ├── Usage Analytics (see how many calls are made per day)
  └── Versioning (run v1 and v2 of your API side by side)

🟣 Technique 2 — OCI Functions (Serverless — Pay Only When Used)

What if your app is not used every hour of every day?
Running a full container 24/7 wastes money. OCI Functions solves this.

OCI Functions = Serverless.
The app sleeps when no one is using it.
The moment a user sends a request, it wakes up in milliseconds, processes it, and goes back to sleep.
You pay only for the exact milliseconds it runs. Nothing more.

  Traditional Server:             OCI Functions (Serverless):
  ┌────────────────┐              ┌────────────────────────────┐
  │ Server running │              │ App is SLEEPING 😴          │
  │ 24 hours/day   │              │ (costs $0 while sleeping)  │
  │ even when no   │              └────────────────────────────┘
  │ one uses it    │                          │
  │ 💸 costs money │              User sends a request
  │ constantly     │                          │
  └────────────────┘                          ▼
                                 ┌────────────────────────────┐
                                 │ App WAKES UP instantly ⚡   │
                                 │ Processes request          │
                                 │ Goes back to sleep         │
                                 │ 💚 You pay only for this!  │
                                 └────────────────────────────┘
📝 What does this OCI Function code do?

This is how you write your app as an OCI Function instead of a full server.
The handler function is the entry point — OCI calls this every time a request arrives.
Notice it's much simpler than a full FastAPI app — perfect for lightweight tasks.
# func.py — OCI Function version of MeetingMuse
import io
import json
import oci
from fdk import response

def handler(ctx, data: io.BytesIO = None):
    """OCI calls this function every time a user hits the endpoint."""

    # Read the incoming request body (the user's Zoom notes)
    body = json.loads(data.getvalue())
    notes = body.get("notes", "")

    # Call OCI GenAI (same as before)
    config = oci.config.from_file("/function/.oci/config")
    client = oci.generative_ai_inference.GenerativeAiInferenceClient(config)

    # (same prompt and API call as app.py above...)
    # ... (abbreviated for brevity)
    summary = "Polished summary returned here..."

    # Return the result back to the user
    return response.Response(
        ctx,
        response_data=json.dumps({"polished_summary": summary}),
        headers={"Content-Type": "application/json"}
    )

🟠 Technique 3 — Blue/Green Deployment (Zero Downtime Updates)

What if you release a new version of MeetingMuse and it has a bug?
With regular deployment, your users would see errors while you fix it. Not good!

Blue/Green Deployment is like running two restaurants side by side.
🔵 Blue = current version (serving customers right now)
🟢 Green = new version (being tested quietly in the background)
When Green is confirmed working, you flip all traffic to Green instantly.
If anything goes wrong, flip back to Blue in seconds. Zero downtime. 🎯

  STAGE 1: Both versions running, Blue serving traffic
  ┌──────────────────────────────────────────────────────┐
  │  Load Balancer                                       │
  │      │                                               │
  │      ├──── 100% traffic ────► 🔵 Blue (v1, live)    │
  │      └──── 0% traffic  ────► 🟢 Green (v2, testing) │
  └──────────────────────────────────────────────────────┘

  STAGE 2: Green tested and approved — flip the switch!
  ┌──────────────────────────────────────────────────────┐
  │  Load Balancer                                       │
  │      │                                               │
  │      ├──── 0% traffic  ────► 🔵 Blue (v1, standby)  │
  │      └──── 100% traffic ───► 🟢 Green (v2, live!) ✨│
  └──────────────────────────────────────────────────────┘

  If something goes wrong with v2? Flip back to Blue instantly.
  No downtime. No angry users. 😌

🟡 Technique 4 — Health Checks and Auto-Restart

Apps can crash. It happens. But your users should never notice.
Kubernetes (and even Docker alone) can automatically detect when your app crashes
and restart it within seconds. This is called a health check.

📝 What does this Kubernetes health check config do?

livenessProbe — Kubernetes pings your /health endpoint every 10 seconds.
If it gets no response 3 times in a row, it automatically kills and restarts the container.
readinessProbe — Kubernetes waits until your app is truly ready before sending users to it.
Together, they make your app self-healing. 💊
# Add this inside your Kubernetes deployment.yaml, under containers:

livenessProbe:
  httpGet:
    path: /health         # Hits our /health endpoint we built in app.py
    port: 8000
  initialDelaySeconds: 10  # Wait 10 seconds before starting health checks
  periodSeconds: 10         # Check every 10 seconds after that
  failureThreshold: 3       # If it fails 3 times → restart the container

readinessProbe:
  httpGet:
    path: /health
    port: 8000
  initialDelaySeconds: 5   # Start checking after 5 seconds
  periodSeconds: 5          # Check every 5 seconds

🗂️ Full Architecture — Everything Together

                         USER (browser / mobile app)
                                    │
                                    │  HTTPS request
                                    ▼
                         ┌─────────────────────┐
                         │   OCI API Gateway   │  ← Authentication, rate limiting, clean URL
                         └──────────┬──────────┘
                                    │
                                    ▼
                         ┌─────────────────────┐
                         │  OCI Load Balancer  │  ← Distributes traffic evenly
                         └──────────┬──────────┘
                                    │
                     ┌──────────────┼──────────────┐
                     ▼              ▼              ▼
              ┌─────────────┐ ┌─────────────┐ ┌─────────────┐
              │ MeetingMuse │ │ MeetingMuse │ │ MeetingMuse │
              │  Pod 1 🔵   │ │  Pod 2 🟢   │ │  Pod 3 🟡   │
              │ (Docker     │ │ (Docker     │ │ (Docker     │
              │  Container) │ │  Container) │ │  Container) │
              └──────┬──────┘ └──────┬──────┘ └──────┬──────┘
                     └──────────────┬──────────────┘
                                    │
                                    ▼
                         ┌─────────────────────┐
                         │   OCI GenAI API     │  ← Cohere Command R+ model
                         │ (does the AI magic) │
                         └─────────────────────┘
                                    │
                                    ▼
                      Polished meeting summary returned ✨

📋 Comparison — Which Technique Should You Use?

┌─────────────────────────┬─────────────────┬──────────────┬─────────────────┐
│ Technique               │ Best For        │ Complexity   │ Cost            │
├─────────────────────────┼─────────────────┼──────────────┼─────────────────┤
│ Container Instance      │ Quick demos,    │ ⭐ Easy      │ 💲 Low          │
│                         │ beginners       │              │                 │
├─────────────────────────┼─────────────────┼──────────────┼─────────────────┤
│ OCI Functions           │ Low traffic,    │ ⭐⭐ Medium  │ 💲 Very Low     │
│ (Serverless)            │ cost-sensitive  │              │ (pay per call)  │
├─────────────────────────┼─────────────────┼──────────────┼─────────────────┤
│ OKE (Kubernetes)        │ Production,     │ ⭐⭐⭐ Hard  │ 💲💲 Medium     │
│                         │ high traffic    │              │                 │
├─────────────────────────┼─────────────────┼──────────────┼─────────────────┤
│ + OCI API Gateway       │ Any production  │ ⭐⭐ Medium  │ 💲 Low add-on   │
│                         │ API             │              │                 │
├─────────────────────────┼─────────────────┼──────────────┼─────────────────┤
│ Blue/Green              │ Zero-downtime   │ ⭐⭐⭐ Hard  │ 💲💲 (2x infra │
│                         │ releases        │              │ during switch)  │
└─────────────────────────┴─────────────────┴──────────────┴─────────────────┘
🚫 Things to NEVER do in production:

❌ Never hardcode your API keys directly inside app.py — use environment variables.
❌ Never run your Docker container as the root user — create a limited user in the Dockerfile.
❌ Never expose port 8000 directly to the internet without an API Gateway or load balancer.
❌ Never skip the health check endpoints — they are the safety net of your app.

Happy building & deploying! 🐳✨

Comments