Skip to main content

Docker for MLOps: Containerize and Deploy Machine Learning Models

Calculating read time…

Imagine you spend 3 weeks building the most amazing LEGO castle. You photograph every detail. You're so proud of it.

Then your friend says: "Can I build the exact same castle?" You hand over the instructions, but they have different LEGO bricks — different colors, slightly different shapes. Their castle looks completely wrong.

Now imagine a magic box that contains every single LEGO brick you used, in exactly the right type and quantity. Your friend opens the box and builds a perfect replica — every time.




💡 That magic box = a Docker Container!

In Machine Learning, we spend weeks training a model on our laptop. It works perfectly. But when we deploy it to Oracle OCI (Cloud Infrastructure), it crashes — wrong Python version, missing library, different OS.

Docker packages your ML model with everything it needs to run: Python version, libraries, environment variables, and code — all in one portable, isolated box. Push it to Oracle Container Registry. Pull it onto any OCI Compute instance. It always works. 🎯

📚 What We'll Cover

  • 🔹 What is Docker? (The Shipping Container Story)
  • 🔹 VMs vs Containers — What's the Difference?
  • 🔹 Core Docker Concepts: Image, Container, Registry
  • 🔹 The Dockerfile — Line by Line (Reference Image Explained)
  • 🔹 Every Dockerfile Instruction Decoded
  • 🔹 Step-by-Step: Dockerizing a Real ML Model (FastAPI + OCI)
  • 🔹 Essential Docker Commands Cheat Sheet
  • 🔹 Docker Compose — Running Multi-Container ML Systems
  • 🔹 Oracle Container Registry (OCIR) — Push & Pull Images on OCI
  • 🔹 Deploying Docker on OCI Compute & OCI Container Instances
  • 🔹 Docker in CI/CD with OCI DevOps Pipelines
  • 🔹 Multi-Stage Builds — Slim Production Images
  • 🔹 GPU Support in Docker — Running Deep Learning on OCI GPU Shapes
  • 🔹 Common Mistakes & Best Practices
  • 🔹 High-Level Summary & Next Steps

🚢 Section 1: What is Docker? (The Shipping Container Story)

Before Docker, deploying software was like moving house without boxes. You'd grab items one by one — a lamp here, a plate there — and hope everything survived the trip and fit in the new place.

In the 1950s, shipping companies had the exact same problem. Every ship loaded cargo differently. Port workers had to figure out how to fit irregular objects every single time. Slow. Expensive. Chaotic.

Then someone invented the standardized shipping container. Pack everything inside. Seal it. Any crane, any ship, any port handles it identically. It revolutionized global trade overnight.

THE DOCKER ANALOGY (with Oracle OCI):

🏠 Your ML Code = The cargo (fragile, needs specific conditions)
📦 Docker Image = The shipping container (contains everything needed)
🚢 Docker Engine = The crane that moves containers
🏭 OCIR = Oracle Container Registry — the global port
☁️ OCI Compute / K8s = The destination port (container runs identically here)

Result: "Works on my machine" problem — PERMANENTLY ELIMINATED. ✅

🐳 What Docker Actually Does for ML

Docker is a platform that packages an application with all its dependencies into a single unit called a container. That container runs identically on your laptop, your colleague's Windows PC, an Oracle OCI Compute VM (VM.Standard.E4.Flex), or an OCI Kubernetes cluster (OKE).

For ML specifically, Docker eliminates the most painful deployment problem: "My model works in my Jupyter notebook but crashes the moment I deploy it to OCI."

✅ What Docker guarantees for your ML model on Oracle OCI:
  • ✅ Same Python version (3.10.12, not "whatever OCI has installed")
  • ✅ Same library versions (torch==2.3.0, scikit-learn==1.4.0)
  • ✅ Same OS base (Ubuntu 22.04 / Debian Bookworm)
  • ✅ Same file paths and folder structure
  • ✅ Same environment variables (OCI config, model paths)
  • ✅ Runs identically across OCI regions (Frankfurt, Mumbai, Sydney, Phoenix)

🖥️ Section 2: VMs vs Containers — What's the Difference?

Oracle OCI offers both Virtual Machines (VM.Standard shapes) and Container-based deployments (OCI Container Instances). Understanding the difference helps you choose wisely.

🏠 The Apartment Building Analogy

Virtual Machine (OCI VM.Standard.E4.Flex) = A separate house
Each house has its own foundation, plumbing, electricity, and walls — a full copy of the OS. Moving or copying takes hours. Very isolated, but very heavy and expensive to run in parallel.

Docker Container (OCI Container Instance) = An apartment in a building
Everyone shares the same foundation and infrastructure (OS kernel). Each apartment has its own furniture (libraries, app code). Starts in seconds. Extremely lightweight. Still nicely isolated.

OCI VIRTUAL MACHINE: OCI CONTAINER INSTANCE:

┌─────────────────────────┐ ┌─────────────────────────┐
│ ML App A │ │ ML App A │
│ Python 3.10 + libs │ │ Python 3.10 + libs │
│ Guest OS (2GB RAM min) │ └─────────────────────────┘
├─────────────────────────┤ ┌─────────────────────────┐
│ ML App B │ │ ML App B │
│ Python 3.11 + libs │ │ Python 3.11 + libs │
│ Guest OS (2GB RAM min) │ └─────────────────────────┘
├─────────────────────────┤ ┌─────────────────────────┐
│ OCI Hypervisor (KVM) │ │ Docker Engine │
│ Host OS │ │ OCI Host OS (shared) │
└─────────────────────────┘ └─────────────────────────┘

VM startup: 1–5 minutes 🐢 Container startup: 1–3 secs 🚀
VM size: 4–20 GB disk Container size: 200 MB – 3 GB
OCI cost: Full VM shape billing OCI cost: Per-second, per-OCPU
Aspect OCI Virtual Machine OCI Container / Docker
Startup time 1–5 minutes 1–3 seconds ⚡
Image size 40–80 GB (full OS disk) 200 MB – 3 GB
Resource use High (full OS per VM) Low (shared kernel)
OCI billing Per shape, even when idle Per second, per OCPU used
Portability OCI-region limited Any OCI region, any cloud ✅
Best use case Long-running dedicated servers ML APIs, batch jobs, microservices

🧱 Section 3: Core Docker Concepts — Image, Container, Registry

Docker has three fundamental building blocks. Let's use a cooking analogy to make them completely clear. 🍕

1️⃣ Docker Image = A Recipe Card

A Docker Image is a read-only blueprint — a set of layered instructions for creating a container. It contains everything: the OS base, installed libraries, your model files, your API code, and the startup command.

Just like a pizza recipe doesn't make the pizza — the image is the recipe. The container is the actual pizza. 🍕

2️⃣ Docker Container = A Running Pizza

A Docker Container is a live, running instance of an image. You can create 50 identical containers from one image — all independent, all running in parallel, all serving ML predictions simultaneously. This is how ML APIs handle high traffic on OCI.

3️⃣ Oracle Container Registry (OCIR) = The Recipe Library

OCIR (Oracle Cloud Infrastructure Registry) is Oracle's managed container registry. It's where you push your Docker images so they can be pulled by OCI Compute instances, OCI Kubernetes Engine (OKE), or OCI Container Instances.

OCIR is region-specific but images can be replicated across OCI regions. It supports both public and private repositories and integrates natively with OCI IAM policies for access control. 🔐

THE COMPLETE DOCKER + OCI LIFECYCLE:

1. Write Dockerfile → recipe instructions on your laptop
↓
2. docker build → Docker Image built locally
↓
3. docker tag + docker push → Image uploaded to OCIR
(ocir.io/<tenancy>/<repo>:<tag>)
↓
4. OCI Compute / OCI CI → docker pull from OCIR
↓
5. docker run → Container runs your ML model! 🎉

🧅 Images Are Layered — Why This Matters

Every Docker image is built in layers. Each Dockerfile instruction creates a new layer. Unchanged layers are cached by Docker — meaning if only your ML code changes but your pip packages didn't, Docker skips the slow pip install layer and rebuilds in seconds.

Docker Image Layers (bottom to top):

┌──────────────────────────────────────────────┐
│ Layer 6: ENTRYPOINT — startup command │ ← thin, always rebuilt
├──────────────────────────────────────────────┤
│ Layer 5: COPY . /app — your ML code │ ← rebuilt on code change
├──────────────────────────────────────────────┤
│ Layer 4: RUN pip install -r requirements │ ← cached if reqs unchanged
├──────────────────────────────────────────────┤
│ Layer 3: COPY requirements.txt │ ← trigger for layer 4
├──────────────────────────────────────────────┤
│ Layer 2: RUN apt-get install build-essential│ ← almost never changes
├──────────────────────────────────────────────┤
│ Layer 1: FROM python:3.10-slim-buster │ ← base, never changes
└──────────────────────────────────────────────┘

💡 Rule: Put STABLE instructions first. CHANGING instructions last.

📝 Section 4: The Dockerfile — Line by Line (Reference Image Decoded)

Below is the exact Dockerfile from our reference image — a real production example. We'll explain every single line as if you're 10 years old. No jargon left unexplained.

# ── The base image (your starting blank canvas) ─────────────────
FROM python:3.10-slim-buster

# ── Set working directory inside container ───────────────────────
WORKDIR /app

# ── Install OS-level packages ────────────────────────────────────
RUN apt-get update && \
    apt-get install --no-install-recommends -y \
            build-essential \
            && apt-get clean && rm -rf /tmp/* /var/tmp/*

# ── Copy requirements FIRST (cache optimization trick!) ──────────
COPY requirements.txt /app/requirements.txt

# ── Upgrade pip and install Python dependencies ──────────────────
RUN pip3 install --upgrade pip
RUN pip3 install --no-cache-dir -r requirements.txt

# ── Copy ALL application code into the container ─────────────────
COPY . /app

# ── Document which port the app listens on ───────────────────────
EXPOSE 8083

# ── Set runtime environment variables ────────────────────────────
ENV PYTHONPATH="/app"

# ── Accept build-time argument (e.g. dev / staging / prod) ───────
ARG environment
ENV ENVIRONMENT $environment

# ── Make the entrypoint script executable ────────────────────────
RUN chmod +x /app/entrypoint.sh

# ── Set the default command that starts when container launches ───
ENTRYPOINT ["/app/entrypoint.sh"]

🔍 Section 5: Every Dockerfile Instruction Decoded

Let's go through every instruction step by step — plain English, no assumed knowledge.

1️⃣ FROM — Choose Your Starting Point

📌 What this instruction does: Selects the base image — the pre-built foundation that everything else builds on. Every Dockerfile must start with FROM. Think of it as choosing which kitchen you'll cook in before starting the recipe.
FROM python:3.10-slim-buster

Breaking this down word by word:

  • python — the official Python Docker image from Docker Hub. Python is already installed and ready. You don't have to install it yourself.
  • 3.10 — Python version 3.10 specifically. Not 3.9, not 3.11 — exactly 3.10. This is why environments are reproducible!
  • slim — a minimal version of the image with fewer pre-installed tools. Results in a much smaller final image (100 MB vs 900 MB for the full version).
  • buster — based on Debian 10 "Buster" Linux operating system. Buster is well-tested and stable.
💡 Popular Base Images for ML on Oracle OCI in 2025:

For CPU-only models (most common):
python:3.11-slim-bookworm — lightest, fastest cold start on OCI Container Instances

For GPU / Deep Learning models (OCI GPU shapes: BM.GPU.A10.4, VM.GPU3.1):
pytorch/pytorch:2.3.0-cuda12.1-cudnn8-runtime
nvcr.io/nvidia/pytorch:24.03-py3 — NVIDIA NGC image, optimized for OCI A100/A10 GPUs

For TensorFlow models on OCI:
tensorflow/tensorflow:2.16.1-gpu

Oracle's own base images (pre-optimized for OCI):
container-registry.oracle.com/os/oraclelinux:8-slim
container-registry.oracle.com/os/oraclelinux:9

2️⃣ WORKDIR — Set Your Home Inside the Container

📌 What this instruction does: Sets the current working directory inside the container for all subsequent instructions. It's exactly like typing cd /app in a terminal and staying there. If the folder doesn't exist, Docker creates it automatically.
WORKDIR /app

All COPY, RUN, and CMD instructions that come after this will operate inside /app.

✅ Always define WORKDIR! Without it, Docker places files in the root directory (/), which is messy and can conflict with OS files. /app is the universal industry convention for application containers.

3️⃣ RUN — Execute Commands During the Build

📌 What this instruction does: Executes a shell command during the image build process. Whatever the command produces becomes a permanent new layer in your image. This particular RUN installs low-level OS packages needed by Python numerical libraries.
RUN apt-get update && \
    apt-get install --no-install-recommends -y \
                build-essential \
                && apt-get clean && rm -rf /tmp/* /var/tmp/*

Decoding every part:

  • apt-get update — refreshes the list of available packages. Like clicking "Check for Updates" on your laptop before installing anything.
  • apt-get install --no-install-recommends -y — installs only essential packages, skips optional extras. The -y flag auto-answers "yes" to any installation prompt (containers can't respond to interactive prompts during builds).
  • build-essential — installs GCC C/C++ compilers. Required to compile some Python packages like numpy, scipy, and tokenizers from source when pre-built wheels aren't available.
  • apt-get clean && rm -rf /tmp/* — deletes all installation cache files immediately after installing. Critical for keeping image size small. Doing this in a separate RUN command wouldn't help — the cache would already be baked into the previous layer.
❌ DON'T do this — common beginner mistake:
RUN apt-get update && apt-get install -y build-essential
RUN apt-get clean    # ← WRONG! This cleanup is in a separate layer.
                     #   The cache from the install layer still exists!
                     #   Image is still bloated.
✅ Always clean up in the SAME RUN instruction as the install.

4️⃣ COPY requirements.txt — The Famous Cache Trick

📌 What this instruction does: Copies only the requirements.txt file from your local machine into the container — before copying any other code. This is a deliberate ordering decision that dramatically speeds up rebuilds.
COPY requirements.txt /app/requirements.txt

Why copy this file alone before everything else? Because of Docker's layer caching system.

The next instruction (pip install) is the slowest step — it can take 3–10 minutes to download and install all ML packages. If you copy all your code first, then pip install, Docker will re-run pip install every single time you change one line of code.

But if you copy requirements.txt first, then pip install, then copy your code — Docker only re-runs pip install when requirements.txt changes. Code-only changes skip straight to the COPY step. Rebuilds go from 10 minutes to 5 seconds. ⚡

WRONG ORDER (slow every time): CORRECT ORDER (fast rebuilds):

COPY . /app ← code first COPY requirements.txt /app/
RUN pip install -r req ← always slow RUN pip install -r req ← cached! ✅
(rebuilds pip on every code change) COPY . /app ← only this reruns

5️⃣ & 6️⃣ RUN pip install — Install All Python Packages

📌 What this instruction does: First upgrades pip itself to the latest version, then reads requirements.txt and installs every Python library listed — your ML framework, web server, data processing tools, and all their dependencies.
RUN pip3 install --upgrade pip
RUN pip3 install --no-cache-dir -r requirements.txt

Breaking down the flags:

  • --upgrade pip — ensures pip is the latest version. Older pip versions sometimes fail to resolve complex dependency trees (common with torch + cuda packages).
  • --no-cache-dir — tells pip NOT to store downloaded packages locally. Since we're building a container (not a reusable environment), the cache would just waste space in the image. This flag alone can reduce image size by 200–400 MB for large ML projects.
  • -r requirements.txt — installs every package listed in the file.
💡 What goes in requirements.txt for a typical ML API?
# ML frameworks
scikit-learn==1.4.2
torch==2.3.0
xgboost==2.0.3

# API server (FastAPI is standard in 2025)
fastapi==0.111.0
uvicorn[standard]==0.30.0

# Data processing
pandas==2.2.2
numpy==1.26.4

# OCI SDK (for Oracle Cloud integration)
oci==2.127.0

# Model serialization
joblib==1.4.2
  

7️⃣ COPY . /app — Copy All Your Application Code

📌 What this instruction does: Copies everything from your current local directory (the dot .) into the container's /app folder. This includes your trained model files, prediction scripts, API code, configuration files, and the entrypoint script. This is intentionally placed AFTER pip install for the caching benefit described above.
COPY . /app
💡 Always use a .dockerignore file! Just like .gitignore tells Git what to skip, .dockerignore tells Docker what NOT to copy. This keeps your image lean and prevents secrets from leaking into the image.
# .dockerignore
.git/
.gitignore
__pycache__/
*.pyc
*.pyo
.env                 # ← NEVER copy secrets into image!
*.log
notebooks/           # Jupyter notebooks not needed in prod
data/raw/            # Raw data (too large, use OCI Object Storage)
.venv/               # Local virtual environment
tests/               # Unit tests not needed in container
*.md                 # Documentation files
.DS_Store            # macOS metadata files
  

8️⃣ EXPOSE — Document the Network Port

📌 What this instruction does: Documents which network port your application listens on inside the container. This is primarily for documentation and tooling — it does NOT automatically open the port on the host machine. Port mapping still happens at docker run -p time.
EXPOSE 8083

Port 8083 is where the FastAPI/uvicorn web server will listen for HTTP requests inside the container. When running on OCI, you'll map this to an external port.

💡 OCI Security Note: When deploying on Oracle OCI, you also need to open the port in:
  • OCI Security List (on the VCN subnet) — allows traffic at network level
  • OCI Network Security Group (NSG) — fine-grained ingress/egress rules
  • OS-level firewall on the Compute instance (iptables/firewalld)
EXPOSE in the Dockerfile alone is not enough for OCI deployments.

9️⃣ ENV — Set Environment Variables

📌 What this instruction does: Sets a permanent environment variable inside the container. These variables are available to your application code at runtime, just like environment variables in a regular Linux terminal.
ENV PYTHONPATH="/app"

Setting PYTHONPATH="/app" tells Python to look for modules inside /app. This means you can use clean imports like from src.model import FraudDetector instead of messy relative paths.

✅ Good use of ENV for OCI ML deployments:
ENV PYTHONPATH="/app"
ENV MODEL_PATH="/app/models/fraud_model.pkl"
ENV LOG_LEVEL="INFO"
ENV OCI_REGION="ap-mumbai-1"
ENV OBJECT_STORAGE_BUCKET="ml-model-artifacts"
  
❌ NEVER put secrets inside ENV in Dockerfile!

These are baked permanently into the image and visible to anyone who inspects it:
# ❌ WRONG — secret is now inside the image forever!
ENV OCI_API_KEY="my-super-secret-key"
ENV DB_PASSWORD="prod-database-password"
  
✅ Use OCI Vault (secrets management service) and inject secrets at container runtime using --env-file or OCI Container Instance environment variables — never hardcode them in the Dockerfile.

🔟 ARG — Build-Time Arguments

📌 What this instruction does: Declares a build-time variable — a value you pass in when running docker build. Unlike ENV, ARG variables are only available during the build process, not at container runtime.
ARG environment
ENV ENVIRONMENT $environment

The ARG environment line declares a variable called environment. Then ENV ENVIRONMENT $environment converts it into a runtime environment variable so your app can read it.

This lets you build different images for different OCI environments using the same Dockerfile:

# Build for OCI development environment
docker build --build-arg environment=development -t my-model:dev .

# Build for OCI staging environment
docker build --build-arg environment=staging -t my-model:staging .

# Build for OCI production environment
docker build --build-arg environment=production -t my-model:prod .

1️⃣1️⃣ & 1️⃣2️⃣ RUN chmod + ENTRYPOINT — The Startup Command

📌 What these instructions do: chmod +x makes the shell script executable (gives it "run permission"). ENTRYPOINT defines the command that runs automatically when the container starts — it's the "on switch" for your ML application.
RUN chmod +x /app/entrypoint.sh
ENTRYPOINT ["/app/entrypoint.sh"]

The entrypoint.sh script typically starts your ML API server. Here's what a typical entrypoint script looks like for a FastAPI ML model:

#!/bin/bash
# entrypoint.sh — Starts the ML prediction API server

set -e   # Exit immediately if any command fails

echo "🚀 Starting ML Fraud Detection API..."
echo "   Environment: $ENVIRONMENT"
echo "   OCI Region:  $OCI_REGION"
echo "   Model path:  $MODEL_PATH"

# Run database migrations if needed (for apps with a DB)
# python -m alembic upgrade head

# Start the FastAPI server with uvicorn
exec uvicorn src.api.main:app \
    --host 0.0.0.0 \
    --port 8083 \
    --workers 2 \
    --log-level $LOG_LEVEL
💡 ENTRYPOINT vs CMD — What's the Difference?

ENTRYPOINT = the fixed command that always runs. Cannot be overridden easily.
CMD = default arguments to ENTRYPOINT. Can be overridden at docker run time.

For production ML APIs on OCI: use ENTRYPOINT.
For development/research containers: use CMD (more flexible for testing).

🏗️ Section 6: Step-by-Step — Dockerizing a Real ML Model (FastAPI + OCI)

Now let's build a complete, production-ready Docker setup for a Fraud Detection ML model served via FastAPI, deployed on Oracle OCI.

Project Structure

fraud-detection-api/
├── Dockerfile                  ← our container recipe
├── .dockerignore               ← what to exclude
├── requirements.txt            ← Python dependencies
├── entrypoint.sh               ← container startup script
├── src/
│   ├── __init__.py
│   ├── api/
│   │   ├── __init__.py
│   │   └── main.py             ← FastAPI application
│   └── model/
│       ├── __init__.py
│       └── predictor.py        ← model loading + prediction
├── models/
│   └── fraud_model.pkl         ← trained scikit-learn model
└── configs/
    └── config.yaml             ← app configuration

Step 1: The ML Prediction API (src/api/main.py)

📌 Code Purpose — FastAPI ML Prediction Server

What this code does: Creates a production-ready REST API with three endpoints: a health check (for OCI Load Balancer health probes), a single-prediction endpoint, and a batch-prediction endpoint. The model loads once at startup and stays in memory for fast predictions.

Why it matters: OCI Load Balancer and OCI Container Instances use the /health endpoint to decide if the container is healthy and ready to receive traffic. If the health check fails, OCI automatically replaces the container.
"""
src/api/main.py
FastAPI ML Prediction API — production-ready for OCI deployment
"""

from fastapi import FastAPI, HTTPException, Request
from pydantic import BaseModel, Field
from typing import List
import time
import os
import logging

from src.model.predictor import FraudPredictor

# ── Setup logging ─────────────────────────────────────────────────────
logging.basicConfig(level=os.getenv("LOG_LEVEL", "INFO"))
logger = logging.getLogger(__name__)

# ── Create FastAPI app ────────────────────────────────────────────────
app = FastAPI(
    title       = "Fraud Detection API",
    description = "ML-powered fraud scoring service — deployed on Oracle OCI",
    version     = "2.1.0",
)

# ── Load model once at startup (not on every request!) ───────────────
MODEL_PATH = os.getenv("MODEL_PATH", "/app/models/fraud_model.pkl")
predictor  = None

@app.on_event("startup")
async def load_model():
    global predictor
    logger.info(f"Loading model from: {MODEL_PATH}")
    predictor = FraudPredictor(model_path=MODEL_PATH)
    logger.info("✅ Model loaded successfully")


# ── Request / Response schemas ────────────────────────────────────────
class TransactionRequest(BaseModel):
    user_id:            int   = Field(..., example=12345)
    amount:             float = Field(..., example=450.00, gt=0)
    merchant_country:   str   = Field(..., example="NG")
    hour_of_day:        int   = Field(..., example=2, ge=0, le=23)
    transactions_7d:    int   = Field(..., example=8, ge=0)
    spend_velocity:     float = Field(..., example=2.5, ge=0)

class PredictionResponse(BaseModel):
    user_id:      int
    fraud_score:  float
    decision:     str   # "ALLOW" or "BLOCK"
    model_version: str
    latency_ms:   float


# ── Endpoints ──────────────────────────────────────────────────────────

@app.get("/health")
async def health_check():
    """
    OCI Load Balancer health probe endpoint.
    Returns 200 OK when the model is loaded and ready.
    Returns 503 if the model isn't loaded yet.
    """
    if predictor is None:
        raise HTTPException(status_code=503, detail="Model not loaded yet")
    return {
        "status":      "healthy",
        "model":       "fraud_detection_v2",
        "environment": os.getenv("ENVIRONMENT", "unknown"),
        "oci_region":  os.getenv("OCI_REGION", "unknown"),
    }


@app.post("/predict", response_model=PredictionResponse)
async def predict(transaction: TransactionRequest):
    """
    Single transaction fraud prediction.
    Used by real-time payment processing on OCI.
    """
    if predictor is None:
        raise HTTPException(status_code=503, detail="Model not ready")

    t0 = time.perf_counter()

    fraud_score = predictor.predict_single(transaction.dict())
    decision    = "BLOCK" if fraud_score > 0.85 else "ALLOW"

    latency_ms = (time.perf_counter() - t0) * 1000

    logger.info(
        f"user={transaction.user_id} "
        f"score={fraud_score:.4f} "
        f"decision={decision} "
        f"latency={latency_ms:.2f}ms"
    )

    return PredictionResponse(
        user_id       = transaction.user_id,
        fraud_score   = round(fraud_score, 4),
        decision      = decision,
        model_version = "2.1.0",
        latency_ms    = round(latency_ms, 2),
    )


@app.post("/predict/batch")
async def predict_batch(transactions: List[TransactionRequest]):
    """
    Batch fraud prediction for multiple transactions.
    Used by OCI Data Flow batch jobs and nightly scoring pipelines.
    """
    if predictor is None:
        raise HTTPException(status_code=503, detail="Model not ready")

    if len(transactions) > 1000:
        raise HTTPException(
            status_code=400,
            detail="Batch size too large. Maximum 1000 transactions per request."
        )

    t0      = time.perf_counter()
    results = []

    for txn in transactions:
        score    = predictor.predict_single(txn.dict())
        decision = "BLOCK" if score > 0.85 else "ALLOW"
        results.append({
            "user_id":     txn.user_id,
            "fraud_score": round(score, 4),
            "decision":    decision,
        })

    total_ms = (time.perf_counter() - t0) * 1000

    return {
        "predictions":  results,
        "count":        len(results),
        "total_ms":     round(total_ms, 2),
        "avg_ms":       round(total_ms / len(results), 2),
    }

Step 2: The Complete Dockerfile

📌 Code Purpose — Production Dockerfile for OCI ML Deployment

What this Dockerfile does: Builds a lean, production-ready Docker image for the fraud detection API. It follows the exact same structure as the reference image from the tutorial — base image → OS packages → requirements (cached) → app code → startup command — optimized specifically for Oracle OCI Container Instance deployment.

Why this order matters: The requirements.txt is copied and installed BEFORE the app code. This means every code-only change rebuilds in ~5 seconds instead of 8+ minutes, because Docker reuses the cached pip install layer.
# ─────────────────────────────────────────────────────────────────────
# Dockerfile — Fraud Detection ML API
# Deploy target: Oracle OCI Container Instance / OKE
# ─────────────────────────────────────────────────────────────────────

# 1️⃣ Base image: slim Python on Debian Buster
#    slim = minimal OS, ~100MB vs ~900MB for full image
FROM python:3.10-slim-buster

# 2️⃣ Set working directory inside the container
WORKDIR /app

# 3️⃣ Install OS-level C compiler tools required by some Python packages
#    Chain everything in ONE RUN + clean up in SAME command to minimize layer size
RUN apt-get update && \
    apt-get install --no-install-recommends -y \
                build-essential \
                curl \
                && apt-get clean && rm -rf /tmp/* /var/tmp/*

# 4️⃣ Copy requirements file FIRST — enables Docker layer caching
#    If requirements.txt doesn't change, layers 5+6 are reused from cache
COPY requirements.txt /app/requirements.txt

# 5️⃣ Upgrade pip to latest version
RUN pip3 install --upgrade pip

# 6️⃣ Install all Python packages (--no-cache-dir reduces image size)
RUN pip3 install --no-cache-dir -r requirements.txt

# 7️⃣ Copy application code AFTER dependencies (cache stays intact for code changes)
COPY . /app

# 8️⃣ Document which port the FastAPI server listens on
EXPOSE 8083

# 9️⃣ Set Python module search path so imports work from /app
ENV PYTHONPATH="/app"

# OCI-specific environment variables
ENV OCI_REGION="ap-mumbai-1"
ENV LOG_LEVEL="INFO"

# 🔟 Accept build-time argument for environment (dev / staging / prod)
ARG environment=production
ENV ENVIRONMENT=$environment

# 1️⃣1️⃣ Make the startup script executable
RUN chmod +x /app/entrypoint.sh

# 1️⃣2️⃣ Set the container startup command
ENTRYPOINT ["/app/entrypoint.sh"]

Step 3: Build the Docker Image

📌 Code Purpose — Building the Docker Image Locally

What this code does: Runs docker build to execute all Dockerfile instructions and produce a local Docker image tagged with a name and version. The --build-arg flag passes the environment variable at build time. The final lines verify the image was created and check its size.
# ── Navigate to your project folder ──────────────────────────────────
cd fraud-detection-api/

# ── Build the Docker image ────────────────────────────────────────────
# -t = tag (name:version)
# --build-arg = pass ARG values
# . = use current directory as build context (where Dockerfile is)

docker build \
    -t fraud-detection-api:2.1.0 \
    --build-arg environment=production \
    .

# ── Watch the build output ────────────────────────────────────────────
# [+] Building 47.3s (12/12) FINISHED
#  => [1/8] FROM docker.io/library/python:3.10-slim-buster     12.1s
#  => [2/8] WORKDIR /app                                         0.1s
#  => [3/8] RUN apt-get update && apt-get install ...           18.4s
#  => [4/8] COPY requirements.txt /app/requirements.txt          0.1s
#  => [5/8] RUN pip3 install --upgrade pip                       2.1s
#  => [6/8] RUN pip3 install --no-cache-dir -r requirements.txt 14.2s
#  => [7/8] COPY . /app                                          0.3s
#  => [8/8] RUN chmod +x /app/entrypoint.sh                      0.1s
#  => exporting to image                                          0.4s
#  ✅ Successfully built a9f3b2c1d8e4
#  ✅ Successfully tagged fraud-detection-api:2.1.0


# ── List your Docker images ───────────────────────────────────────────
docker images

# REPOSITORY               TAG       IMAGE ID       SIZE
# fraud-detection-api      2.1.0     a9f3b2c1d8e4   612MB


# ── Second build (code changed, requirements unchanged) ───────────────
# Edit src/api/main.py → add a log line → rebuild:

docker build -t fraud-detection-api:2.1.1 --build-arg environment=production .

# [+] Building 3.2s (12/12) FINISHED  ← notice: 47s became 3s!
#  => CACHED [3/8] RUN apt-get install ...               ← CACHED ✅
#  => CACHED [5/8] RUN pip3 install --upgrade pip        ← CACHED ✅
#  => CACHED [6/8] RUN pip3 install -r requirements.txt  ← CACHED ✅
#  => [7/8] COPY . /app                                  ← only this ran
#  ✅ Built in 3.2 seconds instead of 47 seconds!


# ── Test the container locally before pushing to OCI ─────────────────
# -d = run in background (detached mode)
# -p = map host port 8083 to container port 8083
# --name = give it a friendly name

docker run -d \
    --name fraud-api-test \
    -p 8083:8083 \
    -e ENVIRONMENT=development \
    -e LOG_LEVEL=DEBUG \
    fraud-detection-api:2.1.0

# Check it's running
docker ps

# CONTAINER ID  IMAGE                    STATUS         PORTS
# c4f8a2b3d9e1  fraud-detection-api:2.1.0  Up 3 seconds  0.0.0.0:8083->8083/tcp


# Test the health endpoint
curl http://localhost:8083/health

# {"status":"healthy","model":"fraud_detection_v2",
#  "environment":"development","oci_region":"ap-mumbai-1"}


# Test a prediction
curl -X POST http://localhost:8083/predict \
    -H "Content-Type: application/json" \
    -d '{
      "user_id": 12345,
      "amount": 890.00,
      "merchant_country": "NG",
      "hour_of_day": 2,
      "transactions_7d": 15,
      "spend_velocity": 3.2
    }'

# {"user_id":12345,"fraud_score":0.9234,"decision":"BLOCK",
#  "model_version":"2.1.0","latency_ms":1.83}  ✅


# Stop and remove the test container
docker stop fraud-api-test
docker rm fraud-api-test

📋 Section 7: Essential Docker Commands Cheat Sheet

Here are the commands every ML engineer uses daily. Bookmark this section! 🔖

📌 Code Purpose — Daily Docker Commands Reference

What this covers: All the Docker commands you need for the complete build → run → debug → cleanup lifecycle. Each command includes the most useful flags with plain-English explanations.
# ════════════════════════════════════════════════════════════════
#  BUILDING IMAGES
# ════════════════════════════════════════════════════════════════

# Build an image from current directory
docker build -t my-model:v1.0 .

# Build with specific Dockerfile path
docker build -t my-model:v1.0 -f docker/Dockerfile.prod .

# Build with build arguments
docker build -t my-model:v1.0 --build-arg environment=staging .

# Build without using cache (force full rebuild)
docker build -t my-model:v1.0 --no-cache .


# ════════════════════════════════════════════════════════════════
#  RUNNING CONTAINERS
# ════════════════════════════════════════════════════════════════

# Run in foreground (logs visible, Ctrl+C to stop)
docker run -p 8083:8083 my-model:v1.0

# Run in background (detached mode)
docker run -d -p 8083:8083 --name my-api my-model:v1.0

# Run with environment variables
docker run -d -p 8083:8083 \
    -e ENVIRONMENT=production \
    -e OCI_REGION=ap-mumbai-1 \
    my-model:v1.0

# Run with environment variables from a file
docker run -d -p 8083:8083 --env-file .env.prod my-model:v1.0

# Run and mount a local folder into the container (for dev)
docker run -d -p 8083:8083 \
    -v $(pwd)/models:/app/models \
    my-model:v1.0

# Run and get an interactive shell (great for debugging!)
docker run -it my-model:v1.0 /bin/bash
# Now you're INSIDE the container — explore files, test commands


# ════════════════════════════════════════════════════════════════
#  MONITORING CONTAINERS
# ════════════════════════════════════════════════════════════════

# List running containers
docker ps

# List ALL containers (including stopped)
docker ps -a

# View container logs
docker logs my-api

# Follow logs in real-time (like tail -f)
docker logs -f my-api

# Show container resource usage (CPU, RAM)
docker stats my-api

# Execute a command inside a running container
docker exec -it my-api bash       # Open shell in running container
docker exec my-api python -c "import torch; print(torch.__version__)"


# ════════════════════════════════════════════════════════════════
#  STOPPING AND CLEANING UP
# ════════════════════════════════════════════════════════════════

# Stop a running container gracefully
docker stop my-api

# Force stop (if graceful stop hangs)
docker kill my-api

# Remove a stopped container
docker rm my-api

# Stop AND remove in one command
docker rm -f my-api

# Remove an image
docker rmi my-model:v1.0

# NUCLEAR OPTION: Remove ALL stopped containers, unused images, networks
# (run this monthly to free up disk space)
docker system prune -a --volumes

# ════════════════════════════════════════════════════════════════
#  INSPECTING IMAGES
# ════════════════════════════════════════════════════════════════

# List all local images with sizes
docker images

# Show image history (all layers + sizes)
docker history my-model:v1.0

# Inspect full image metadata (JSON)
docker inspect my-model:v1.0

# Show image layer details (useful for size optimization)
docker history --no-trunc my-model:v1.0

🐙 Section 8: Docker Compose — Running Multi-Container ML Systems

A real ML system isn't just one container. It's a collection of services working together: the ML prediction API, a Redis cache for speed, a PostgreSQL database for logging predictions, and an Nginx reverse proxy for routing traffic.

Docker Compose lets you define and run all these containers together using a single docker-compose.yaml file. One command starts everything. One command stops everything.

📌 Code Purpose — Docker Compose for Complete ML System

What this code does: Defines a 4-service ML system: the fraud detection API, a Redis cache (for storing recent predictions and preventing duplicate processing), a PostgreSQL database (for logging all predictions with audit trail), and an Nginx reverse proxy (for SSL termination and load balancing). All services communicate on a private Docker network.

Why it matters on OCI: This exact Compose file can be used locally for development, then deployed directly to an OCI Compute instance with docker compose up -d — no configuration changes needed. The same file also serves as the template for OCI Kubernetes (OKE) deployments.
 # docker-compose.yaml
# Complete ML system: API + Cache + Database + Reverse Proxy
# Run with: docker compose up -d
# Stop with: docker compose down

version: "3.9"

services:

  # ── Service 1: Fraud Detection ML API ────────────────────────────
  fraud-api:
    build:
      context: .
      dockerfile: Dockerfile
      args:
        environment: ${ENVIRONMENT:-development}   # from .env file
    image: fraud-detection-api:${APP_VERSION:-latest}
    container_name: fraud-api
    restart: unless-stopped   # auto-restart if container crashes

    ports:
      - "8083:8083"   # expose API port to host

    environment:
      - ENVIRONMENT=${ENVIRONMENT:-development}
      - OCI_REGION=${OCI_REGION:-ap-mumbai-1}
      - MODEL_PATH=/app/models/fraud_model.pkl
      - LOG_LEVEL=${LOG_LEVEL:-INFO}
      - REDIS_URL=redis://redis-cache:6379/0
      - DATABASE_URL=postgresql://mluser:${DB_PASSWORD}@postgres-db:5432/predictions

    volumes:
      - ./models:/app/models:ro   # read-only model files (hot reload without rebuild)

    depends_on:
      redis-cache:
        condition: service_healthy
      postgres-db:
        condition: service_healthy

    networks:
      - ml-network

    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8083/health"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 15s   # give model time to load before first check


  # ── Service 2: Redis Cache ────────────────────────────────────────
  redis-cache:
    image: redis:7.2-alpine        # official Redis, alpine = minimal
    container_name: redis-cache
    restart: unless-stopped

    command: redis-server --maxmemory 512mb --maxmemory-policy allkeys-lru
    # maxmemory: cap Redis at 512MB
    # allkeys-lru: evict least recently used keys when full

    ports:
      - "6379:6379"   # expose for debugging (remove in production!)

    volumes:
      - redis-data:/data   # persist cache across container restarts

    networks:
      - ml-network

    healthcheck:
      test: ["CMD", "redis-cli", "ping"]
      interval: 10s
      timeout: 5s
      retries: 3


  # ── Service 3: PostgreSQL — Prediction Audit Log ──────────────────
  postgres-db:
    image: postgres:16-alpine
    container_name: postgres-db
    restart: unless-stopped

    environment:
      - POSTGRES_DB=predictions
      - POSTGRES_USER=mluser
      - POSTGRES_PASSWORD=${DB_PASSWORD}   # from .env file — never hardcode!

    volumes:
      - postgres-data:/var/lib/postgresql/data   # persistent storage
      - ./sql/init.sql:/docker-entrypoint-initdb.d/init.sql   # schema setup

    networks:
      - ml-network

    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U mluser -d predictions"]
      interval: 10s
      timeout: 5s
      retries: 5


  # ── Service 4: Nginx Reverse Proxy ───────────────────────────────
  nginx:
    image: nginx:1.25-alpine
    container_name: nginx-proxy
    restart: unless-stopped

    ports:
      - "80:80"
      - "443:443"

    volumes:
      - ./nginx/nginx.conf:/etc/nginx/nginx.conf:ro
      - ./nginx/ssl:/etc/nginx/ssl:ro   # TLS certificates from OCI Certificates Service

    depends_on:
      - fraud-api

    networks:
      - ml-network


# ── Shared network (containers communicate by service name) ─────────
networks:
  ml-network:
    driver: bridge
    name: ml-internal-network


# ── Persistent data volumes ──────────────────────────────────────────
volumes:
  redis-data:
    driver: local
  postgres-data:
    driver: local
# ── Run the complete ML system ────────────────────────────────────────
# Create .env file first (never commit this to Git!)
cat > .env << EOF
ENVIRONMENT=development
OCI_REGION=ap-mumbai-1
APP_VERSION=2.1.0
LOG_LEVEL=INFO
DB_PASSWORD=your-secure-password-here
EOF

# Start ALL services (builds images if not already built)
docker compose up -d

# Starting ml-network
# Starting postgres-db  ... done
# Starting redis-cache  ... done
# Starting fraud-api    ... done
# Starting nginx-proxy  ... done

# Check all services are healthy
docker compose ps

# NAME           SERVICE      STATUS    PORTS
# fraud-api      fraud-api    running   0.0.0.0:8083->8083/tcp
# redis-cache    redis-cache  running   0.0.0.0:6379->6379/tcp
# postgres-db    postgres-db  running   5432/tcp
# nginx-proxy    nginx        running   0.0.0.0:80->80, 0.0.0.0:443->443

# View logs from all services at once
docker compose logs -f

# View logs from just the API service
docker compose logs -f fraud-api

# Scale the API to 3 replicas (for load testing)
docker compose up -d --scale fraud-api=3

# Stop everything
docker compose down

# Stop everything AND delete volumes (CAREFUL: deletes all data!)
docker compose down -v

☁️ Section 9: Oracle Container Registry (OCIR) — Push & Pull on OCI

Once your Docker image is built and tested locally, the next step is pushing it to OCIR (Oracle Container Registry) so it can be pulled by OCI Compute instances, OKE clusters, or Container Instances.

🏗️ OCIR Structure

OCIR Image Path Format:

<region-key>.ocir.io / <tenancy-namespace> / <repository> : <tag>

Examples:
ap-mumbai-1.ocir.io/mytenancy/fraud-detection-api:2.1.0
eu-frankfurt-1.ocir.io/mytenancy/ml-models/fraud-api:latest
us-phoenix-1.ocir.io/mytenancy/fraud-api:prod-2025-03-10

OCI Region Keys:
ap-mumbai-1 → bom.ocir.io
eu-frankfurt-1 → fra.ocir.io
us-phoenix-1 → phx.ocir.io
ap-sydney-1 → syd.ocir.io
uk-london-1 → lhr.ocir.io
📌 Code Purpose — Push to Oracle Container Registry (OCIR)

What this code does: Authenticates with OCIR using an OCI Auth Token (not your OCI password), tags the local Docker image with the full OCIR path, pushes it to OCIR, and verifies the push was successful. Also shows how to pull the image from OCIR onto any OCI instance.

Why OCI Auth Token instead of password: OCIR uses Auth Tokens for Docker registry authentication. These are separate from your OCI Console password and can be revoked independently — a security best practice. Auth Tokens are generated in OCI Console → User Settings → Auth Tokens.
# ════════════════════════════════════════════════════════════════
#  STEP 1: SET YOUR OCI VARIABLES
# ════════════════════════════════════════════════════════════════

export OCI_REGION="ap-mumbai-1"
export OCI_TENANCY_NAMESPACE="mytenancynamespace"   # found in OCI Console → Tenancy
export OCI_USERNAME="oracleidentitycloudservice/alice@company.com"
export OCI_AUTH_TOKEN="xxxxxxxxxxx"   # from OCI Console → User Settings → Auth Tokens
export IMAGE_NAME="fraud-detection-api"
export IMAGE_TAG="2.1.0"

# Construct the full OCIR image path
export OCIR_IMAGE="${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${IMAGE_NAME}:${IMAGE_TAG}"
echo "OCIR image path: $OCIR_IMAGE"
# ap-mumbai-1.ocir.io/mytenancynamespace/fraud-detection-api:2.1.0


# ════════════════════════════════════════════════════════════════
#  STEP 2: LOGIN TO OCIR
# ════════════════════════════════════════════════════════════════

# Docker login to OCIR
# Username format: /  OR
#                  oracleidentitycloudservice/  (for federated users)

docker login ${OCI_REGION}.ocir.io \
    --username "${OCI_TENANCY_NAMESPACE}/${OCI_USERNAME}" \
    --password "${OCI_AUTH_TOKEN}"

# Login Succeeded ✅


# ════════════════════════════════════════════════════════════════
#  STEP 3: TAG THE LOCAL IMAGE WITH OCIR PATH
# ════════════════════════════════════════════════════════════════

# Your local image:
docker images | grep fraud-detection-api
# fraud-detection-api   2.1.0   a9f3b2c1d8e4   612MB

# Tag it with full OCIR path
docker tag fraud-detection-api:${IMAGE_TAG} ${OCIR_IMAGE}

# Verify both tags exist locally
docker images | grep fraud
# fraud-detection-api                                2.1.0   a9f3b2c1d8e4   612MB
# ap-mumbai-1.ocir.io/mytenancy/fraud-detection-api  2.1.0   a9f3b2c1d8e4   612MB
# (same image ID — no data was duplicated, just a new tag pointer)


# ════════════════════════════════════════════════════════════════
#  STEP 4: PUSH TO OCIR
# ════════════════════════════════════════════════════════════════

docker push ${OCIR_IMAGE}

# The push refers to repository [ap-mumbai-1.ocir.io/mytenancy/fraud-detection-api]
# layer 1: Pushed  — python:3.10-slim-buster base
# layer 2: Pushed  — apt-get install build-essential
# layer 3: Pushed  — pip install -r requirements.txt
# layer 4: Pushed  — application code
# 2.1.0: digest: sha256:f4b8c2e1d9a7... size: 4152
# ✅ Successfully pushed to OCIR!


# Also push a 'latest' tag for convenience
docker tag ${OCIR_IMAGE} ${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${IMAGE_NAME}:latest
docker push ${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${IMAGE_NAME}:latest


# ════════════════════════════════════════════════════════════════
#  STEP 5: PULL AND RUN ON OCI COMPUTE INSTANCE
# ════════════════════════════════════════════════════════════════

# SSH into your OCI Compute instance (Oracle Linux 8 or Ubuntu 22.04)
# ssh -i ~/.ssh/oci_rsa opc@

# On the OCI Compute instance, login to OCIR
docker login ${OCI_REGION}.ocir.io \
    --username "${OCI_TENANCY_NAMESPACE}/${OCI_USERNAME}" \
    --password "${OCI_AUTH_TOKEN}"

# Pull the image from OCIR
docker pull ${OCIR_IMAGE}

# Run the container on OCI
docker run -d \
    --name fraud-api-prod \
    --restart unless-stopped \
    -p 8083:8083 \
    -e ENVIRONMENT=production \
    -e OCI_REGION=ap-mumbai-1 \
    -e LOG_LEVEL=INFO \
    ${OCIR_IMAGE}

# Verify it's running
docker ps
curl http://localhost:8083/health
# {"status":"healthy","environment":"production","oci_region":"ap-mumbai-1"}  ✅

⚡ Section 10: Deploying on OCI Container Instances

OCI Container Instances is Oracle's serverless container service. You provide a Docker image from OCIR, specify the shape (vCPU + RAM), and OCI handles all the underlying infrastructure. No VM to manage. No OS patching. No SSH required. Billed per second of actual use. Perfect for ML APIs! 🎯

📌 Code Purpose — Deploy ML Container to OCI Container Instance via CLI

What this code does: Uses the OCI CLI to create a Container Instance — Oracle's serverless container runtime. The CLI call specifies the compute shape, networking, OCIR image path, environment variables, and health check configuration. After creation, the command waits for the container to become ACTIVE.

Why OCI Container Instance for ML APIs: Unlike a full OCI Compute VM (which runs 24/7 and costs money even when idle), OCI Container Instances scale to zero when not in use, start in seconds when a request arrives, and bill per-second. Ideal for ML APIs with variable traffic patterns.
# ── Deploy to OCI Container Instance using OCI CLI ───────────────────

# Ensure OCI CLI is installed and configured
oci --version
# Oracle Cloud Infrastructure CLI 3.x.x

# ── Get required OCIDs ───────────────────────────────────────────────
# You need these from OCI Console:
export OCI_COMPARTMENT_ID="ocid1.compartment.oc1..aaaaa..."
export OCI_SUBNET_ID="ocid1.subnet.oc1.ap-mumbai-1.aaaaa..."
export OCI_IMAGE_URL="ap-mumbai-1.ocir.io/mytenancy/fraud-detection-api:2.1.0"

# ── Create the Container Instance ────────────────────────────────────
oci container-instances container-instance create \
    --compartment-id   $OCI_COMPARTMENT_ID \
    --display-name     "fraud-detection-api-prod" \
    --shape            "CI.Standard.E4.Flex" \
    --shape-config     '{"ocpus": 1, "memoryInGBs": 4}' \
    --availability-domain "Uocm:AP-MUMBAI-1-AD-1" \
    --vnics            "[{
                          \"subnetId\": \"$OCI_SUBNET_ID\",
                          \"displayName\": \"fraud-api-vnic\",
                          \"isPublicIpAssigned\": true
                        }]" \
    --containers       "[{
                          \"displayName\": \"fraud-api-container\",
                          \"imageUrl\": \"$OCI_IMAGE_URL\",
                          \"resourceConfig\": {
                            \"vcpusLimit\": 1.0,
                            \"memoryLimitInGBs\": 3.0
                          },
                          \"environmentVariables\": {
                            \"ENVIRONMENT\": \"production\",
                            \"OCI_REGION\": \"ap-mumbai-1\",
                            \"LOG_LEVEL\": \"INFO\"
                          },
                          \"healthChecks\": [{
                            \"healthCheckType\": \"HTTP\",
                            \"port\": 8083,
                            \"path\": \"/health\",
                            \"intervalInSeconds\": 30,
                            \"timeoutInSeconds\": 10,
                            \"failureThreshold\": 3,
                            \"successThreshold\": 1
                          }]
                        }]" \
    --wait-for-state ACTIVE

# Output:
# Action completed. Waiting until the resource has entered state: ('ACTIVE',)
# {
#   "data": {
#     "display-name": "fraud-detection-api-prod",
#     "lifecycle-state": "ACTIVE",
#     "vnics": [{"public-ip": "140.238.x.x"}]
#   }
# }
# ✅ Container Instance is ACTIVE!

# Test the deployed API
curl http://140.238.x.x:8083/health
# {"status":"healthy","environment":"production","oci_region":"ap-mumbai-1"}

🔄 Section 11: Docker in CI/CD with OCI DevOps Pipelines

In production, you never build and push Docker images manually. The entire process is automated through a CI/CD pipeline. Every time a developer pushes code to Git, the pipeline automatically: builds the Docker image, tests it, pushes to OCIR, and deploys to OCI.

📌 Code Purpose — OCI DevOps Build Pipeline (build_spec.yaml)

What this code does: Defines an OCI DevOps build pipeline specification that automatically: runs unit tests, builds the Docker image, pushes it to OCIR, and passes the image path to a deployment pipeline. This entire process triggers on every git push to the main branch.

Why this matters: Manual deployments are error-prone and don't scale. With this pipeline, every code change produces a tested, tagged, and deployed container within minutes — with full audit trail in OCI DevOps.
# build_spec.yaml — OCI DevOps Build Pipeline Specification
# Place this file at the root of your repository
# Triggers automatically on: git push to main/release branches

version: 0.1
component: build
timeoutInSeconds: 1800   # 30 minute max build time

env:
  # Build-time variables (injected by OCI DevOps)
  variables:
    APP_NAME:    "fraud-detection-api"
    OCI_REGION:  "ap-mumbai-1"

  # References to OCI Vault secrets (never hardcode credentials!)
  vaultVariables:
    OCI_AUTH_TOKEN: "ocid1.vaultsecret.oc1.ap-mumbai-1.aaaaa..."
    DB_PASSWORD:    "ocid1.vaultsecret.oc1.ap-mumbai-1.bbbbb..."

  exportedVariables:
    - DOCKER_IMAGE_PATH   # passed to deployment pipeline
    - APP_VERSION

steps:

  # ── Step 1: Install dependencies for testing ───────────────────────
  - type: Command
    name: "Install test dependencies"
    command: |
      pip install pytest pytest-asyncio httpx
    onFailure:
      - type: Command
        command: echo "❌ Dependency install failed"

  # ── Step 2: Run unit tests ─────────────────────────────────────────
  - type: Command
    name: "Run unit tests"
    command: |
      echo "🧪 Running unit tests..."
      pytest tests/ -v --tb=short
      echo "✅ All tests passed!"

  # ── Step 3: Set version tag from git commit ────────────────────────
  - type: Command
    name: "Set build version"
    command: |
      export APP_VERSION=$(git rev-parse --short HEAD)
      echo "Build version: $APP_VERSION"

  # ── Step 4: Login to OCIR ──────────────────────────────────────────
  - type: Command
    name: "Login to Oracle Container Registry"
    command: |
      echo "$OCI_AUTH_TOKEN" | docker login \
          ${OCI_REGION}.ocir.io \
          --username "${OCI_TENANCY_NAMESPACE}/${OCI_BUILD_RUNNER_USER}" \
          --password-stdin
      echo "✅ Logged into OCIR"

  # ── Step 5: Build Docker image ────────────────────────────────────
  - type: Command
    name: "Build Docker image"
    command: |
      export OCIR_PATH="${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${APP_NAME}"

      docker build \
          -t ${OCIR_PATH}:${APP_VERSION} \
          -t ${OCIR_PATH}:latest \
          --build-arg environment=production \
          --no-cache \
          .

      echo "✅ Docker image built: ${OCIR_PATH}:${APP_VERSION}"
      export DOCKER_IMAGE_PATH="${OCIR_PATH}:${APP_VERSION}"

  # ── Step 6: Security scan (Trivy vulnerability scanner) ───────────
  - type: Command
    name: "Scan image for vulnerabilities"
    command: |
      # Install Trivy (Docker image scanner)
      curl -sfL https://raw.githubusercontent.com/aquasecurity/trivy/main/contrib/install.sh | sh
      ./bin/trivy image \
          --severity HIGH,CRITICAL \
          --exit-code 1 \
          ${DOCKER_IMAGE_PATH}
      echo "✅ No HIGH/CRITICAL vulnerabilities found"

  # ── Step 7: Push to OCIR ──────────────────────────────────────────
  - type: Command
    name: "Push to Oracle Container Registry"
    command: |
      docker push ${DOCKER_IMAGE_PATH}
      docker push ${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${APP_NAME}:latest
      echo "✅ Image pushed to OCIR: ${DOCKER_IMAGE_PATH}"

outputArtifacts:
  - name: docker_image
    type: DOCKER_IMAGE
    location: ${DOCKER_IMAGE_PATH}

🏗️ Section 12: Multi-Stage Builds — Slim Production Images

The single biggest way to reduce Docker image size for ML models is to use Multi-Stage Builds. This technique uses multiple FROM statements in one Dockerfile — one stage for building/compiling, a final stage that only copies what's needed to run.

📌 Code Purpose — Multi-Stage Dockerfile for Lean OCI Images

What this code does: Uses two build stages: a "builder" stage that installs all compilation tools and builds Python wheels, and a lean "production" stage that only copies the final installed packages — no compilers, no build tools, no cache.

Result: Single-stage image: ~900 MB. Multi-stage image: ~380 MB. Faster pull from OCIR. Smaller attack surface. Lower OCI storage costs.
# ─────────────────────────────────────────────────────────────────────
# Dockerfile.multistage — Two-stage build for minimal production image
# ─────────────────────────────────────────────────────────────────────

# ══════════════════════════════════════════════════
# STAGE 1: Builder — install everything, compile all wheels
# This stage is NOT included in the final image!
# ══════════════════════════════════════════════════
FROM python:3.10-slim-buster AS builder

WORKDIR /build

# Install build tools (compiler for C extensions in numpy, scipy etc.)
RUN apt-get update && \
    apt-get install --no-install-recommends -y \
        build-essential gcc g++ \
        && apt-get clean && rm -rf /tmp/*

# Copy and install requirements — build wheels here
COPY requirements.txt .

# Build wheels into /build/wheels directory
# --wheel-dir stores compiled .whl files
RUN pip install --upgrade pip && \
    pip wheel --no-cache-dir \
              --no-deps \
              --wheel-dir /build/wheels \
              -r requirements.txt

# ══════════════════════════════════════════════════
# STAGE 2: Production — lean runtime image
# ONLY this stage ends up in the final Docker image
# ══════════════════════════════════════════════════
FROM python:3.10-slim-buster AS production

WORKDIR /app

# Install ONLY runtime OS dependencies (no compilers!)
# libgomp1 = required by XGBoost / LightGBM for parallel processing
RUN apt-get update && \
    apt-get install --no-install-recommends -y \
        libgomp1 curl \
        && apt-get clean && rm -rf /tmp/*

# Copy pre-built wheels from builder stage
COPY --from=builder /build/wheels /wheels

# Install from local wheels (fast! no internet download needed)
RUN pip install --no-cache-dir --no-index \
                --find-links=/wheels \
                /wheels/*.whl && \
    rm -rf /wheels   # clean up wheels after install

# Copy application code
COPY . /app

EXPOSE 8083

ENV PYTHONPATH="/app"
ENV OCI_REGION="ap-mumbai-1"

ARG environment=production
ENV ENVIRONMENT=$environment

RUN chmod +x /app/entrypoint.sh
ENTRYPOINT ["/app/entrypoint.sh"]

# ─────────────────────────────────────────────────────────────────────
# Build: docker build -f Dockerfile.multistage -t fraud-api:slim .
# Result: ~900MB → ~380MB  (58% smaller!) ✅
# ─────────────────────────────────────────────────────────────────────

🔥 Section 13: GPU Support in Docker — Deep Learning on OCI GPU Shapes

Oracle OCI offers powerful GPU compute shapes for deep learning:

  • VM.GPU3.1 — 1× NVIDIA V100 (16GB HBM2)
  • VM.GPU.A10.1 — 1× NVIDIA A10 (24GB GDDR6)
  • BM.GPU.A100-v2.8 — 8× NVIDIA A100 (80GB HBM2e each)
  • BM.GPU4.8 — 8× NVIDIA A100 (40GB HBM2)
📌 Code Purpose — GPU-Enabled Dockerfile for OCI Deep Learning

What this code does: Builds a Docker image using NVIDIA's CUDA base image instead of plain Python. The CUDA base gives direct access to GPU hardware. Installs GPU-enabled PyTorch and sets the NVIDIA runtime flag. Includes a model inference test at build time to verify GPU access.

OCI requirement: To use GPU in Docker on OCI, the Compute instance must have the NVIDIA Container Toolkit installed, and Docker must be run with --gpus all or --runtime=nvidia.
# Dockerfile.gpu — For OCI VM.GPU.A10.1 and BM.GPU.A100-v2.8 shapes
# Uses NVIDIA CUDA base image instead of plain Python

# Base image with CUDA 12.1 + cuDNN 8 (matches OCI GPU driver versions)
FROM nvcr.io/nvidia/pytorch:24.03-py3

# OR use official CUDA image and install PyTorch manually:
# FROM nvidia/cuda:12.1.0-cudnn8-runtime-ubuntu22.04

WORKDIR /app

# Install Python dependencies (PyTorch already in base image)
COPY requirements-gpu.txt /app/requirements.txt
RUN pip install --upgrade pip && \
    pip install --no-cache-dir -r requirements.txt

COPY . /app

EXPOSE 8083

ENV PYTHONPATH="/app"
ENV OCI_REGION="ap-mumbai-1"
ENV CUDA_VISIBLE_DEVICES="0"   # use first GPU

RUN chmod +x /app/entrypoint.sh
ENTRYPOINT ["/app/entrypoint.sh"]
# ── Install NVIDIA Container Toolkit on OCI GPU instance ─────────────
# (Run once on the OCI Compute instance — Oracle Linux 8)

# Add NVIDIA Container Toolkit repository
curl -s -L https://nvidia.github.io/libnvidia-container/stable/rpm/nvidia-container-toolkit.repo \
    | sudo tee /etc/yum.repos.d/nvidia-container-toolkit.repo

# Install the toolkit
sudo dnf install -y nvidia-container-toolkit

# Configure Docker to use NVIDIA runtime
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

# Verify GPU is accessible in Docker
docker run --rm --gpus all nvidia/cuda:12.1.0-base-ubuntu22.04 nvidia-smi

# +-----------------------------------------------------------------------------+
# | NVIDIA-SMI 535.54.03    Driver Version: 535.54.03    CUDA Version: 12.2     |
# |-------------------------------+----------------------+----------------------+
# | GPU  Name         Persistence-M | Bus-Id        Disp.A | Volatile Uncorr. ECC |
# |   0  NVIDIA A10             Off  | 00000000:00:04.0 Off |                    0 |
# +-------------------------------+----------------------+----------------------+
# ✅ GPU accessible in Docker container!


# ── Build and run the GPU-enabled ML container ───────────────────────
docker build -f Dockerfile.gpu -t fraud-dl-api:gpu-v1.0 .

docker run -d \
    --gpus all \                    # Enable GPU access
    --runtime=nvidia \              # Use NVIDIA runtime
    -p 8083:8083 \
    -e ENVIRONMENT=production \
    -e CUDA_VISIBLE_DEVICES=0 \
    fraud-dl-api:gpu-v1.0

🚫 Section 14: Common Mistakes & Best Practices

✅ DOs — Docker Best Practices for OCI ML Deployments:
  • ✅ Always use specific version tags: python:3.10-slim-buster not python:latest
  • ✅ Copy requirements.txt BEFORE other code — maximize layer caching
  • ✅ Use .dockerignore to exclude .git/, .env, notebooks, and large raw data files
  • ✅ Chain RUN commands and clean up in the SAME instruction to minimize layer size
  • ✅ Use multi-stage builds for production — reduces image size by 40–60%
  • ✅ Always expose the /health endpoint for OCI Load Balancer health probes
  • ✅ Store secrets in OCI Vault — never in ENV variables in the Dockerfile
  • ✅ Tag images with both version (2.1.0) AND commit hash (a9f3b2c) for traceability
  • ✅ Scan images for vulnerabilities with Trivy before pushing to OCIR
  • ✅ Use --restart unless-stopped for production containers on OCI Compute
❌ DON'Ts — Mistakes That Hurt Production ML Systems:
  • ❌ Don't use :latest tags in production — always pin to specific versions
  • ❌ Don't copy secrets into Docker images — they're visible in docker inspect and image layers
  • ❌ Don't run containers as root — use USER to create a non-root user
  • ❌ Don't install dev tools (git, vim, wget) in production images — unnecessary attack surface
  • ❌ Don't run apt-get clean in a separate RUN from the install — it won't reduce image size
  • ❌ Don't ignore the health check — OCI Container Instances and OKE rely on it for traffic routing
  • ❌ Don't load ML models inside API endpoint functions — load ONCE at startup, reuse for all requests
  • ❌ Don't push 5 GB model files inside the Docker image — store models in OCI Object Storage and download at startup
  • ❌ Don't skip .dockerignore — accidentally copying .git/ adds hundreds of MB to the image
💡 Pro Tip — Non-Root User for Production Security:
# Add this to your Dockerfile before ENTRYPOINT:

# Create a non-root user for security
RUN addgroup --gid 1001 mlgroup && \
    adduser --uid 1001 --gid 1001 --disabled-password mluser

# Set file ownership
RUN chown -R mluser:mlgroup /app

# Switch to non-root user
USER mluser

ENTRYPOINT ["/app/entrypoint.sh"]
  
Running as non-root is required by many OCI security policies and prevents privilege escalation attacks if the container is compromised.

🏆 High-Level Summary

  • 🔹 Docker = the standardized shipping container for software. Eliminates "works on my machine" forever.
  • 🔹 Image = read-only blueprint (recipe). Container = running instance (the pizza). OCIR = Oracle's container registry (the library).
  • 🔹 FROM picks your base. WORKDIR sets your home. RUN executes build-time commands. COPY brings in files.
  • 🔹 ENV sets runtime variables. ARG sets build-time variables. EXPOSE documents ports. ENTRYPOINT is the on switch.
  • 🔹 Layer caching trick: Copy requirements.txt BEFORE your code. pip install runs from cache on code-only changes. Rebuild goes from 10 min → 5 sec.
  • 🔹 Docker Compose = run multiple containers as one system. API + Redis + PostgreSQL + Nginx with one command.
  • 🔹 OCIR (Oracle Container Registry) = push images from local → pull on any OCI region. Authenticated via OCI Auth Token.
  • 🔹 OCI Container Instances = serverless containers. No VM management. Per-second billing. Health check required.
  • 🔹 Multi-stage builds = separate build and runtime stages. 900 MB → 380 MB images. Faster OCIR pulls.
  • 🔹 GPU Docker on OCI = NVIDIA CUDA base image + NVIDIA Container Toolkit + --gpus all flag.
  • 🔹 OCI DevOps pipeline = auto build → test → scan → push to OCIR → deploy on every git push.
  • 🔹 Security rule #1: NEVER put secrets in Dockerfile ENV. Use OCI Vault + runtime injection.

Keep building, keep containerizing, keep shipping! 🐼✨

Comments