Imagine you spend 3 weeks building the most amazing LEGO castle. You photograph every detail. You're so proud of it.
Then your friend says: "Can I build the exact same castle?" You hand over the instructions, but they have different LEGO bricks — different colors, slightly different shapes. Their castle looks completely wrong.
Now imagine a magic box that contains every single LEGO brick you used, in exactly the right type and quantity. Your friend opens the box and builds a perfect replica — every time.
In Machine Learning, we spend weeks training a model on our laptop. It works perfectly. But when we deploy it to Oracle OCI (Cloud Infrastructure), it crashes — wrong Python version, missing library, different OS.
Docker packages your ML model with everything it needs to run: Python version, libraries, environment variables, and code — all in one portable, isolated box. Push it to Oracle Container Registry. Pull it onto any OCI Compute instance. It always works. 🎯
📚 What We'll Cover
- 🔹 What is Docker? (The Shipping Container Story)
- 🔹 VMs vs Containers — What's the Difference?
- 🔹 Core Docker Concepts: Image, Container, Registry
- 🔹 The Dockerfile — Line by Line (Reference Image Explained)
- 🔹 Every Dockerfile Instruction Decoded
- 🔹 Step-by-Step: Dockerizing a Real ML Model (FastAPI + OCI)
- 🔹 Essential Docker Commands Cheat Sheet
- 🔹 Docker Compose — Running Multi-Container ML Systems
- 🔹 Oracle Container Registry (OCIR) — Push & Pull Images on OCI
- 🔹 Deploying Docker on OCI Compute & OCI Container Instances
- 🔹 Docker in CI/CD with OCI DevOps Pipelines
- 🔹 Multi-Stage Builds — Slim Production Images
- 🔹 GPU Support in Docker — Running Deep Learning on OCI GPU Shapes
- 🔹 Common Mistakes & Best Practices
- 🔹 High-Level Summary & Next Steps
🚢 Section 1: What is Docker? (The Shipping Container Story)
Before Docker, deploying software was like moving house without boxes. You'd grab items one by one — a lamp here, a plate there — and hope everything survived the trip and fit in the new place.
In the 1950s, shipping companies had the exact same problem. Every ship loaded cargo differently. Port workers had to figure out how to fit irregular objects every single time. Slow. Expensive. Chaotic.
Then someone invented the standardized shipping container. Pack everything inside. Seal it. Any crane, any ship, any port handles it identically. It revolutionized global trade overnight.
🏠 Your ML Code = The cargo (fragile, needs specific conditions)
📦 Docker Image = The shipping container (contains everything needed)
🚢 Docker Engine = The crane that moves containers
🏭 OCIR = Oracle Container Registry — the global port
☁️ OCI Compute / K8s = The destination port (container runs identically here)
Result: "Works on my machine" problem — PERMANENTLY ELIMINATED. ✅
🐳 What Docker Actually Does for ML
Docker is a platform that packages an application with all its dependencies into a single unit called a container. That container runs identically on your laptop, your colleague's Windows PC, an Oracle OCI Compute VM (VM.Standard.E4.Flex), or an OCI Kubernetes cluster (OKE).
For ML specifically, Docker eliminates the most painful deployment problem: "My model works in my Jupyter notebook but crashes the moment I deploy it to OCI."
- ✅ Same Python version (3.10.12, not "whatever OCI has installed")
- ✅ Same library versions (torch==2.3.0, scikit-learn==1.4.0)
- ✅ Same OS base (Ubuntu 22.04 / Debian Bookworm)
- ✅ Same file paths and folder structure
- ✅ Same environment variables (OCI config, model paths)
- ✅ Runs identically across OCI regions (Frankfurt, Mumbai, Sydney, Phoenix)
🖥️ Section 2: VMs vs Containers — What's the Difference?
Oracle OCI offers both Virtual Machines (VM.Standard shapes) and Container-based deployments (OCI Container Instances). Understanding the difference helps you choose wisely.
🏠 The Apartment Building Analogy
Virtual Machine (OCI VM.Standard.E4.Flex) = A separate house
Each house has its own foundation, plumbing, electricity, and walls —
a full copy of the OS. Moving or copying takes hours.
Very isolated, but very heavy and expensive to run in parallel.
Docker Container (OCI Container Instance) = An apartment in a building
Everyone shares the same foundation and infrastructure (OS kernel).
Each apartment has its own furniture (libraries, app code).
Starts in seconds. Extremely lightweight. Still nicely isolated.
┌─────────────────────────┐ ┌─────────────────────────┐
│ ML App A │ │ ML App A │
│ Python 3.10 + libs │ │ Python 3.10 + libs │
│ Guest OS (2GB RAM min) │ └─────────────────────────┘
├─────────────────────────┤ ┌─────────────────────────┐
│ ML App B │ │ ML App B │
│ Python 3.11 + libs │ │ Python 3.11 + libs │
│ Guest OS (2GB RAM min) │ └─────────────────────────┘
├─────────────────────────┤ ┌─────────────────────────┐
│ OCI Hypervisor (KVM) │ │ Docker Engine │
│ Host OS │ │ OCI Host OS (shared) │
└─────────────────────────┘ └─────────────────────────┘
VM startup: 1–5 minutes 🐢 Container startup: 1–3 secs 🚀
VM size: 4–20 GB disk Container size: 200 MB – 3 GB
OCI cost: Full VM shape billing OCI cost: Per-second, per-OCPU
| Aspect | OCI Virtual Machine | OCI Container / Docker |
|---|---|---|
| Startup time | 1–5 minutes | 1–3 seconds ⚡ |
| Image size | 40–80 GB (full OS disk) | 200 MB – 3 GB |
| Resource use | High (full OS per VM) | Low (shared kernel) |
| OCI billing | Per shape, even when idle | Per second, per OCPU used |
| Portability | OCI-region limited | Any OCI region, any cloud ✅ |
| Best use case | Long-running dedicated servers | ML APIs, batch jobs, microservices |
🧱 Section 3: Core Docker Concepts — Image, Container, Registry
Docker has three fundamental building blocks. Let's use a cooking analogy to make them completely clear. 🍕
1️⃣ Docker Image = A Recipe Card
A Docker Image is a read-only blueprint — a set of layered instructions for creating a container. It contains everything: the OS base, installed libraries, your model files, your API code, and the startup command.
Just like a pizza recipe doesn't make the pizza — the image is the recipe. The container is the actual pizza. 🍕
2️⃣ Docker Container = A Running Pizza
A Docker Container is a live, running instance of an image. You can create 50 identical containers from one image — all independent, all running in parallel, all serving ML predictions simultaneously. This is how ML APIs handle high traffic on OCI.
3️⃣ Oracle Container Registry (OCIR) = The Recipe Library
OCIR (Oracle Cloud Infrastructure Registry) is Oracle's managed container registry. It's where you push your Docker images so they can be pulled by OCI Compute instances, OCI Kubernetes Engine (OKE), or OCI Container Instances.
OCIR is region-specific but images can be replicated across OCI regions. It supports both public and private repositories and integrates natively with OCI IAM policies for access control. 🔐
1. Write Dockerfile → recipe instructions on your laptop
↓
2. docker build → Docker Image built locally
↓
3. docker tag + docker push → Image uploaded to OCIR
(ocir.io/<tenancy>/<repo>:<tag>)
↓
4. OCI Compute / OCI CI → docker pull from OCIR
↓
5. docker run → Container runs your ML model! 🎉
🧅 Images Are Layered — Why This Matters
Every Docker image is built in layers.
Each Dockerfile instruction creates a new layer.
Unchanged layers are cached by Docker — meaning if only
your ML code changes but your pip packages didn't,
Docker skips the slow pip install layer and rebuilds in seconds.
┌──────────────────────────────────────────────┐
│ Layer 6: ENTRYPOINT — startup command │ ← thin, always rebuilt
├──────────────────────────────────────────────┤
│ Layer 5: COPY . /app — your ML code │ ← rebuilt on code change
├──────────────────────────────────────────────┤
│ Layer 4: RUN pip install -r requirements │ ← cached if reqs unchanged
├──────────────────────────────────────────────┤
│ Layer 3: COPY requirements.txt │ ← trigger for layer 4
├──────────────────────────────────────────────┤
│ Layer 2: RUN apt-get install build-essential│ ← almost never changes
├──────────────────────────────────────────────┤
│ Layer 1: FROM python:3.10-slim-buster │ ← base, never changes
└──────────────────────────────────────────────┘
💡 Rule: Put STABLE instructions first. CHANGING instructions last.
📝 Section 4: The Dockerfile — Line by Line (Reference Image Decoded)
Below is the exact Dockerfile from our reference image — a real production example. We'll explain every single line as if you're 10 years old. No jargon left unexplained.
FROM python:3.10-slim-buster
# ── Set working directory inside container ───────────────────────
WORKDIR /app
# ── Install OS-level packages ────────────────────────────────────
RUN apt-get update && \
apt-get install --no-install-recommends -y \
build-essential \
&& apt-get clean && rm -rf /tmp/* /var/tmp/*
# ── Copy requirements FIRST (cache optimization trick!) ──────────
COPY requirements.txt /app/requirements.txt
# ── Upgrade pip and install Python dependencies ──────────────────
RUN pip3 install --upgrade pip
RUN pip3 install --no-cache-dir -r requirements.txt
# ── Copy ALL application code into the container ─────────────────
COPY . /app
# ── Document which port the app listens on ───────────────────────
EXPOSE 8083
# ── Set runtime environment variables ────────────────────────────
ENV PYTHONPATH="/app"
# ── Accept build-time argument (e.g. dev / staging / prod) ───────
ARG environment
ENV ENVIRONMENT $environment
# ── Make the entrypoint script executable ────────────────────────
RUN chmod +x /app/entrypoint.sh
# ── Set the default command that starts when container launches ───
ENTRYPOINT ["/app/entrypoint.sh"]
🔍 Section 5: Every Dockerfile Instruction Decoded
Let's go through every instruction step by step — plain English, no assumed knowledge.
1️⃣ FROM — Choose Your Starting Point
FROM python:3.10-slim-buster
Breaking this down word by word:
- python — the official Python Docker image from Docker Hub. Python is already installed and ready. You don't have to install it yourself.
- 3.10 — Python version 3.10 specifically. Not 3.9, not 3.11 — exactly 3.10. This is why environments are reproducible!
- slim — a minimal version of the image with fewer pre-installed tools. Results in a much smaller final image (100 MB vs 900 MB for the full version).
- buster — based on Debian 10 "Buster" Linux operating system. Buster is well-tested and stable.
For CPU-only models (most common):
python:3.11-slim-bookworm — lightest, fastest cold start on OCI Container InstancesFor GPU / Deep Learning models (OCI GPU shapes: BM.GPU.A10.4, VM.GPU3.1):
pytorch/pytorch:2.3.0-cuda12.1-cudnn8-runtimenvcr.io/nvidia/pytorch:24.03-py3 — NVIDIA NGC image, optimized for OCI A100/A10 GPUsFor TensorFlow models on OCI:
tensorflow/tensorflow:2.16.1-gpuOracle's own base images (pre-optimized for OCI):
container-registry.oracle.com/os/oraclelinux:8-slimcontainer-registry.oracle.com/os/oraclelinux:9
2️⃣ WORKDIR — Set Your Home Inside the Container
cd /app in a terminal and staying there.
If the folder doesn't exist, Docker creates it automatically.
WORKDIR /app
All COPY, RUN, and CMD instructions
that come after this will operate inside /app.
/),
which is messy and can conflict with OS files.
/app is the universal industry convention for application containers.
3️⃣ RUN — Execute Commands During the Build
RUN apt-get update && \
apt-get install --no-install-recommends -y \
build-essential \
&& apt-get clean && rm -rf /tmp/* /var/tmp/*
Decoding every part:
-
apt-get update— refreshes the list of available packages. Like clicking "Check for Updates" on your laptop before installing anything. -
apt-get install --no-install-recommends -y— installs only essential packages, skips optional extras. The-yflag auto-answers "yes" to any installation prompt (containers can't respond to interactive prompts during builds). -
build-essential— installs GCC C/C++ compilers. Required to compile some Python packages like numpy, scipy, and tokenizers from source when pre-built wheels aren't available. -
apt-get clean && rm -rf /tmp/*— deletes all installation cache files immediately after installing. Critical for keeping image size small. Doing this in a separate RUN command wouldn't help — the cache would already be baked into the previous layer.
RUN apt-get update && apt-get install -y build-essential
RUN apt-get clean # ← WRONG! This cleanup is in a separate layer.
# The cache from the install layer still exists!
# Image is still bloated.
✅ Always clean up in the SAME RUN instruction as the install.
4️⃣ COPY requirements.txt — The Famous Cache Trick
requirements.txt file from your local machine
into the container — before copying any other code.
This is a deliberate ordering decision that dramatically speeds up rebuilds.
COPY requirements.txt /app/requirements.txt
Why copy this file alone before everything else? Because of Docker's layer caching system.
The next instruction (pip install) is the slowest step —
it can take 3–10 minutes to download and install all ML packages.
If you copy all your code first, then pip install,
Docker will re-run pip install every single time you change one line of code.
But if you copy requirements.txt first,
then pip install, then copy your code —
Docker only re-runs pip install when requirements.txt changes.
Code-only changes skip straight to the COPY step. Rebuilds go from 10 minutes to 5 seconds. ⚡
COPY . /app ← code first COPY requirements.txt /app/
RUN pip install -r req ← always slow RUN pip install -r req ← cached! ✅
(rebuilds pip on every code change) COPY . /app ← only this reruns
5️⃣ & 6️⃣ RUN pip install — Install All Python Packages
requirements.txt and installs every Python library listed —
your ML framework, web server, data processing tools, and all their dependencies.
RUN pip3 install --upgrade pip
RUN pip3 install --no-cache-dir -r requirements.txt
Breaking down the flags:
-
--upgrade pip— ensures pip is the latest version. Older pip versions sometimes fail to resolve complex dependency trees (common with torch + cuda packages). -
--no-cache-dir— tells pip NOT to store downloaded packages locally. Since we're building a container (not a reusable environment), the cache would just waste space in the image. This flag alone can reduce image size by 200–400 MB for large ML projects. -
-r requirements.txt— installs every package listed in the file.
# ML frameworks scikit-learn==1.4.2 torch==2.3.0 xgboost==2.0.3 # API server (FastAPI is standard in 2025) fastapi==0.111.0 uvicorn[standard]==0.30.0 # Data processing pandas==2.2.2 numpy==1.26.4 # OCI SDK (for Oracle Cloud integration) oci==2.127.0 # Model serialization joblib==1.4.2
7️⃣ COPY . /app — Copy All Your Application Code
.)
into the container's /app folder.
This includes your trained model files, prediction scripts, API code,
configuration files, and the entrypoint script.
This is intentionally placed AFTER pip install for the caching benefit described above.
COPY . /app
.gitignore tells Git what to skip,
.dockerignore tells Docker what NOT to copy.
This keeps your image lean and prevents secrets from leaking into the image.
# .dockerignore .git/ .gitignore __pycache__/ *.pyc *.pyo .env # ← NEVER copy secrets into image! *.log notebooks/ # Jupyter notebooks not needed in prod data/raw/ # Raw data (too large, use OCI Object Storage) .venv/ # Local virtual environment tests/ # Unit tests not needed in container *.md # Documentation files .DS_Store # macOS metadata files
8️⃣ EXPOSE — Document the Network Port
docker run -p time.
EXPOSE 8083
Port 8083 is where the FastAPI/uvicorn web server will listen for HTTP requests inside the container. When running on OCI, you'll map this to an external port.
- OCI Security List (on the VCN subnet) — allows traffic at network level
- OCI Network Security Group (NSG) — fine-grained ingress/egress rules
- OS-level firewall on the Compute instance (iptables/firewalld)
9️⃣ ENV — Set Environment Variables
ENV PYTHONPATH="/app"
Setting PYTHONPATH="/app" tells Python to look for modules
inside /app. This means you can use clean imports like
from src.model import FraudDetector instead of messy relative paths.
ENV PYTHONPATH="/app" ENV MODEL_PATH="/app/models/fraud_model.pkl" ENV LOG_LEVEL="INFO" ENV OCI_REGION="ap-mumbai-1" ENV OBJECT_STORAGE_BUCKET="ml-model-artifacts"
These are baked permanently into the image and visible to anyone who inspects it:
# ❌ WRONG — secret is now inside the image forever! ENV OCI_API_KEY="my-super-secret-key" ENV DB_PASSWORD="prod-database-password"✅ Use OCI Vault (secrets management service) and inject secrets at container runtime using
--env-file or OCI Container Instance
environment variables — never hardcode them in the Dockerfile.
🔟 ARG — Build-Time Arguments
docker build. Unlike ENV, ARG variables are only available
during the build process, not at container runtime.
ARG environment
ENV ENVIRONMENT $environment
The ARG environment line declares a variable called environment.
Then ENV ENVIRONMENT $environment converts it into a runtime environment
variable so your app can read it.
This lets you build different images for different OCI environments using the same Dockerfile:
# Build for OCI development environment
docker build --build-arg environment=development -t my-model:dev .
# Build for OCI staging environment
docker build --build-arg environment=staging -t my-model:staging .
# Build for OCI production environment
docker build --build-arg environment=production -t my-model:prod .
1️⃣1️⃣ & 1️⃣2️⃣ RUN chmod + ENTRYPOINT — The Startup Command
chmod +x makes the shell script executable (gives it "run permission").
ENTRYPOINT defines the command that runs automatically
when the container starts — it's the "on switch" for your ML application.
RUN chmod +x /app/entrypoint.sh
ENTRYPOINT ["/app/entrypoint.sh"]
The entrypoint.sh script typically starts your ML API server.
Here's what a typical entrypoint script looks like for a FastAPI ML model:
#!/bin/bash
# entrypoint.sh — Starts the ML prediction API server
set -e # Exit immediately if any command fails
echo "🚀 Starting ML Fraud Detection API..."
echo " Environment: $ENVIRONMENT"
echo " OCI Region: $OCI_REGION"
echo " Model path: $MODEL_PATH"
# Run database migrations if needed (for apps with a DB)
# python -m alembic upgrade head
# Start the FastAPI server with uvicorn
exec uvicorn src.api.main:app \
--host 0.0.0.0 \
--port 8083 \
--workers 2 \
--log-level $LOG_LEVEL
ENTRYPOINT = the fixed command that always runs. Cannot be overridden easily.
CMD = default arguments to ENTRYPOINT. Can be overridden at
docker run time.For production ML APIs on OCI: use ENTRYPOINT.
For development/research containers: use CMD (more flexible for testing).
🏗️ Section 6: Step-by-Step — Dockerizing a Real ML Model (FastAPI + OCI)
Now let's build a complete, production-ready Docker setup for a Fraud Detection ML model served via FastAPI, deployed on Oracle OCI.
Project Structure
fraud-detection-api/
├── Dockerfile ← our container recipe
├── .dockerignore ← what to exclude
├── requirements.txt ← Python dependencies
├── entrypoint.sh ← container startup script
├── src/
│ ├── __init__.py
│ ├── api/
│ │ ├── __init__.py
│ │ └── main.py ← FastAPI application
│ └── model/
│ ├── __init__.py
│ └── predictor.py ← model loading + prediction
├── models/
│ └── fraud_model.pkl ← trained scikit-learn model
└── configs/
└── config.yaml ← app configuration
Step 1: The ML Prediction API (src/api/main.py)
What this code does: Creates a production-ready REST API with three endpoints: a health check (for OCI Load Balancer health probes), a single-prediction endpoint, and a batch-prediction endpoint. The model loads once at startup and stays in memory for fast predictions.
Why it matters: OCI Load Balancer and OCI Container Instances use the
/health
endpoint to decide if the container is healthy and ready to receive traffic.
If the health check fails, OCI automatically replaces the container.
"""
src/api/main.py
FastAPI ML Prediction API — production-ready for OCI deployment
"""
from fastapi import FastAPI, HTTPException, Request
from pydantic import BaseModel, Field
from typing import List
import time
import os
import logging
from src.model.predictor import FraudPredictor
# ── Setup logging ─────────────────────────────────────────────────────
logging.basicConfig(level=os.getenv("LOG_LEVEL", "INFO"))
logger = logging.getLogger(__name__)
# ── Create FastAPI app ────────────────────────────────────────────────
app = FastAPI(
title = "Fraud Detection API",
description = "ML-powered fraud scoring service — deployed on Oracle OCI",
version = "2.1.0",
)
# ── Load model once at startup (not on every request!) ───────────────
MODEL_PATH = os.getenv("MODEL_PATH", "/app/models/fraud_model.pkl")
predictor = None
@app.on_event("startup")
async def load_model():
global predictor
logger.info(f"Loading model from: {MODEL_PATH}")
predictor = FraudPredictor(model_path=MODEL_PATH)
logger.info("✅ Model loaded successfully")
# ── Request / Response schemas ────────────────────────────────────────
class TransactionRequest(BaseModel):
user_id: int = Field(..., example=12345)
amount: float = Field(..., example=450.00, gt=0)
merchant_country: str = Field(..., example="NG")
hour_of_day: int = Field(..., example=2, ge=0, le=23)
transactions_7d: int = Field(..., example=8, ge=0)
spend_velocity: float = Field(..., example=2.5, ge=0)
class PredictionResponse(BaseModel):
user_id: int
fraud_score: float
decision: str # "ALLOW" or "BLOCK"
model_version: str
latency_ms: float
# ── Endpoints ──────────────────────────────────────────────────────────
@app.get("/health")
async def health_check():
"""
OCI Load Balancer health probe endpoint.
Returns 200 OK when the model is loaded and ready.
Returns 503 if the model isn't loaded yet.
"""
if predictor is None:
raise HTTPException(status_code=503, detail="Model not loaded yet")
return {
"status": "healthy",
"model": "fraud_detection_v2",
"environment": os.getenv("ENVIRONMENT", "unknown"),
"oci_region": os.getenv("OCI_REGION", "unknown"),
}
@app.post("/predict", response_model=PredictionResponse)
async def predict(transaction: TransactionRequest):
"""
Single transaction fraud prediction.
Used by real-time payment processing on OCI.
"""
if predictor is None:
raise HTTPException(status_code=503, detail="Model not ready")
t0 = time.perf_counter()
fraud_score = predictor.predict_single(transaction.dict())
decision = "BLOCK" if fraud_score > 0.85 else "ALLOW"
latency_ms = (time.perf_counter() - t0) * 1000
logger.info(
f"user={transaction.user_id} "
f"score={fraud_score:.4f} "
f"decision={decision} "
f"latency={latency_ms:.2f}ms"
)
return PredictionResponse(
user_id = transaction.user_id,
fraud_score = round(fraud_score, 4),
decision = decision,
model_version = "2.1.0",
latency_ms = round(latency_ms, 2),
)
@app.post("/predict/batch")
async def predict_batch(transactions: List[TransactionRequest]):
"""
Batch fraud prediction for multiple transactions.
Used by OCI Data Flow batch jobs and nightly scoring pipelines.
"""
if predictor is None:
raise HTTPException(status_code=503, detail="Model not ready")
if len(transactions) > 1000:
raise HTTPException(
status_code=400,
detail="Batch size too large. Maximum 1000 transactions per request."
)
t0 = time.perf_counter()
results = []
for txn in transactions:
score = predictor.predict_single(txn.dict())
decision = "BLOCK" if score > 0.85 else "ALLOW"
results.append({
"user_id": txn.user_id,
"fraud_score": round(score, 4),
"decision": decision,
})
total_ms = (time.perf_counter() - t0) * 1000
return {
"predictions": results,
"count": len(results),
"total_ms": round(total_ms, 2),
"avg_ms": round(total_ms / len(results), 2),
}
Step 2: The Complete Dockerfile
What this Dockerfile does: Builds a lean, production-ready Docker image for the fraud detection API. It follows the exact same structure as the reference image from the tutorial — base image → OS packages → requirements (cached) → app code → startup command — optimized specifically for Oracle OCI Container Instance deployment.
Why this order matters: The
requirements.txt is copied and installed BEFORE the app code.
This means every code-only change rebuilds in ~5 seconds instead of 8+ minutes,
because Docker reuses the cached pip install layer.
# ─────────────────────────────────────────────────────────────────────
# Dockerfile — Fraud Detection ML API
# Deploy target: Oracle OCI Container Instance / OKE
# ─────────────────────────────────────────────────────────────────────
# 1️⃣ Base image: slim Python on Debian Buster
# slim = minimal OS, ~100MB vs ~900MB for full image
FROM python:3.10-slim-buster
# 2️⃣ Set working directory inside the container
WORKDIR /app
# 3️⃣ Install OS-level C compiler tools required by some Python packages
# Chain everything in ONE RUN + clean up in SAME command to minimize layer size
RUN apt-get update && \
apt-get install --no-install-recommends -y \
build-essential \
curl \
&& apt-get clean && rm -rf /tmp/* /var/tmp/*
# 4️⃣ Copy requirements file FIRST — enables Docker layer caching
# If requirements.txt doesn't change, layers 5+6 are reused from cache
COPY requirements.txt /app/requirements.txt
# 5️⃣ Upgrade pip to latest version
RUN pip3 install --upgrade pip
# 6️⃣ Install all Python packages (--no-cache-dir reduces image size)
RUN pip3 install --no-cache-dir -r requirements.txt
# 7️⃣ Copy application code AFTER dependencies (cache stays intact for code changes)
COPY . /app
# 8️⃣ Document which port the FastAPI server listens on
EXPOSE 8083
# 9️⃣ Set Python module search path so imports work from /app
ENV PYTHONPATH="/app"
# OCI-specific environment variables
ENV OCI_REGION="ap-mumbai-1"
ENV LOG_LEVEL="INFO"
# 🔟 Accept build-time argument for environment (dev / staging / prod)
ARG environment=production
ENV ENVIRONMENT=$environment
# 1️⃣1️⃣ Make the startup script executable
RUN chmod +x /app/entrypoint.sh
# 1️⃣2️⃣ Set the container startup command
ENTRYPOINT ["/app/entrypoint.sh"]
Step 3: Build the Docker Image
What this code does: Runs
docker build to execute all Dockerfile instructions
and produce a local Docker image tagged with a name and version.
The --build-arg flag passes the environment variable at build time.
The final lines verify the image was created and check its size.
# ── Navigate to your project folder ──────────────────────────────────
cd fraud-detection-api/
# ── Build the Docker image ────────────────────────────────────────────
# -t = tag (name:version)
# --build-arg = pass ARG values
# . = use current directory as build context (where Dockerfile is)
docker build \
-t fraud-detection-api:2.1.0 \
--build-arg environment=production \
.
# ── Watch the build output ────────────────────────────────────────────
# [+] Building 47.3s (12/12) FINISHED
# => [1/8] FROM docker.io/library/python:3.10-slim-buster 12.1s
# => [2/8] WORKDIR /app 0.1s
# => [3/8] RUN apt-get update && apt-get install ... 18.4s
# => [4/8] COPY requirements.txt /app/requirements.txt 0.1s
# => [5/8] RUN pip3 install --upgrade pip 2.1s
# => [6/8] RUN pip3 install --no-cache-dir -r requirements.txt 14.2s
# => [7/8] COPY . /app 0.3s
# => [8/8] RUN chmod +x /app/entrypoint.sh 0.1s
# => exporting to image 0.4s
# ✅ Successfully built a9f3b2c1d8e4
# ✅ Successfully tagged fraud-detection-api:2.1.0
# ── List your Docker images ───────────────────────────────────────────
docker images
# REPOSITORY TAG IMAGE ID SIZE
# fraud-detection-api 2.1.0 a9f3b2c1d8e4 612MB
# ── Second build (code changed, requirements unchanged) ───────────────
# Edit src/api/main.py → add a log line → rebuild:
docker build -t fraud-detection-api:2.1.1 --build-arg environment=production .
# [+] Building 3.2s (12/12) FINISHED ← notice: 47s became 3s!
# => CACHED [3/8] RUN apt-get install ... ← CACHED ✅
# => CACHED [5/8] RUN pip3 install --upgrade pip ← CACHED ✅
# => CACHED [6/8] RUN pip3 install -r requirements.txt ← CACHED ✅
# => [7/8] COPY . /app ← only this ran
# ✅ Built in 3.2 seconds instead of 47 seconds!
# ── Test the container locally before pushing to OCI ─────────────────
# -d = run in background (detached mode)
# -p = map host port 8083 to container port 8083
# --name = give it a friendly name
docker run -d \
--name fraud-api-test \
-p 8083:8083 \
-e ENVIRONMENT=development \
-e LOG_LEVEL=DEBUG \
fraud-detection-api:2.1.0
# Check it's running
docker ps
# CONTAINER ID IMAGE STATUS PORTS
# c4f8a2b3d9e1 fraud-detection-api:2.1.0 Up 3 seconds 0.0.0.0:8083->8083/tcp
# Test the health endpoint
curl http://localhost:8083/health
# {"status":"healthy","model":"fraud_detection_v2",
# "environment":"development","oci_region":"ap-mumbai-1"}
# Test a prediction
curl -X POST http://localhost:8083/predict \
-H "Content-Type: application/json" \
-d '{
"user_id": 12345,
"amount": 890.00,
"merchant_country": "NG",
"hour_of_day": 2,
"transactions_7d": 15,
"spend_velocity": 3.2
}'
# {"user_id":12345,"fraud_score":0.9234,"decision":"BLOCK",
# "model_version":"2.1.0","latency_ms":1.83} ✅
# Stop and remove the test container
docker stop fraud-api-test
docker rm fraud-api-test
📋 Section 7: Essential Docker Commands Cheat Sheet
Here are the commands every ML engineer uses daily. Bookmark this section! 🔖
What this covers: All the Docker commands you need for the complete build → run → debug → cleanup lifecycle. Each command includes the most useful flags with plain-English explanations.
# ════════════════════════════════════════════════════════════════
# BUILDING IMAGES
# ════════════════════════════════════════════════════════════════
# Build an image from current directory
docker build -t my-model:v1.0 .
# Build with specific Dockerfile path
docker build -t my-model:v1.0 -f docker/Dockerfile.prod .
# Build with build arguments
docker build -t my-model:v1.0 --build-arg environment=staging .
# Build without using cache (force full rebuild)
docker build -t my-model:v1.0 --no-cache .
# ════════════════════════════════════════════════════════════════
# RUNNING CONTAINERS
# ════════════════════════════════════════════════════════════════
# Run in foreground (logs visible, Ctrl+C to stop)
docker run -p 8083:8083 my-model:v1.0
# Run in background (detached mode)
docker run -d -p 8083:8083 --name my-api my-model:v1.0
# Run with environment variables
docker run -d -p 8083:8083 \
-e ENVIRONMENT=production \
-e OCI_REGION=ap-mumbai-1 \
my-model:v1.0
# Run with environment variables from a file
docker run -d -p 8083:8083 --env-file .env.prod my-model:v1.0
# Run and mount a local folder into the container (for dev)
docker run -d -p 8083:8083 \
-v $(pwd)/models:/app/models \
my-model:v1.0
# Run and get an interactive shell (great for debugging!)
docker run -it my-model:v1.0 /bin/bash
# Now you're INSIDE the container — explore files, test commands
# ════════════════════════════════════════════════════════════════
# MONITORING CONTAINERS
# ════════════════════════════════════════════════════════════════
# List running containers
docker ps
# List ALL containers (including stopped)
docker ps -a
# View container logs
docker logs my-api
# Follow logs in real-time (like tail -f)
docker logs -f my-api
# Show container resource usage (CPU, RAM)
docker stats my-api
# Execute a command inside a running container
docker exec -it my-api bash # Open shell in running container
docker exec my-api python -c "import torch; print(torch.__version__)"
# ════════════════════════════════════════════════════════════════
# STOPPING AND CLEANING UP
# ════════════════════════════════════════════════════════════════
# Stop a running container gracefully
docker stop my-api
# Force stop (if graceful stop hangs)
docker kill my-api
# Remove a stopped container
docker rm my-api
# Stop AND remove in one command
docker rm -f my-api
# Remove an image
docker rmi my-model:v1.0
# NUCLEAR OPTION: Remove ALL stopped containers, unused images, networks
# (run this monthly to free up disk space)
docker system prune -a --volumes
# ════════════════════════════════════════════════════════════════
# INSPECTING IMAGES
# ════════════════════════════════════════════════════════════════
# List all local images with sizes
docker images
# Show image history (all layers + sizes)
docker history my-model:v1.0
# Inspect full image metadata (JSON)
docker inspect my-model:v1.0
# Show image layer details (useful for size optimization)
docker history --no-trunc my-model:v1.0
🐙 Section 8: Docker Compose — Running Multi-Container ML Systems
A real ML system isn't just one container. It's a collection of services working together: the ML prediction API, a Redis cache for speed, a PostgreSQL database for logging predictions, and an Nginx reverse proxy for routing traffic.
Docker Compose lets you define and run all these
containers together using a single docker-compose.yaml file.
One command starts everything. One command stops everything.
What this code does: Defines a 4-service ML system: the fraud detection API, a Redis cache (for storing recent predictions and preventing duplicate processing), a PostgreSQL database (for logging all predictions with audit trail), and an Nginx reverse proxy (for SSL termination and load balancing). All services communicate on a private Docker network.
Why it matters on OCI: This exact Compose file can be used locally for development, then deployed directly to an OCI Compute instance with
docker compose up -d — no configuration changes needed.
The same file also serves as the template for OCI Kubernetes (OKE) deployments.
# docker-compose.yaml
# Complete ML system: API + Cache + Database + Reverse Proxy
# Run with: docker compose up -d
# Stop with: docker compose down
version: "3.9"
services:
# ── Service 1: Fraud Detection ML API ────────────────────────────
fraud-api:
build:
context: .
dockerfile: Dockerfile
args:
environment: ${ENVIRONMENT:-development} # from .env file
image: fraud-detection-api:${APP_VERSION:-latest}
container_name: fraud-api
restart: unless-stopped # auto-restart if container crashes
ports:
- "8083:8083" # expose API port to host
environment:
- ENVIRONMENT=${ENVIRONMENT:-development}
- OCI_REGION=${OCI_REGION:-ap-mumbai-1}
- MODEL_PATH=/app/models/fraud_model.pkl
- LOG_LEVEL=${LOG_LEVEL:-INFO}
- REDIS_URL=redis://redis-cache:6379/0
- DATABASE_URL=postgresql://mluser:${DB_PASSWORD}@postgres-db:5432/predictions
volumes:
- ./models:/app/models:ro # read-only model files (hot reload without rebuild)
depends_on:
redis-cache:
condition: service_healthy
postgres-db:
condition: service_healthy
networks:
- ml-network
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8083/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 15s # give model time to load before first check
# ── Service 2: Redis Cache ────────────────────────────────────────
redis-cache:
image: redis:7.2-alpine # official Redis, alpine = minimal
container_name: redis-cache
restart: unless-stopped
command: redis-server --maxmemory 512mb --maxmemory-policy allkeys-lru
# maxmemory: cap Redis at 512MB
# allkeys-lru: evict least recently used keys when full
ports:
- "6379:6379" # expose for debugging (remove in production!)
volumes:
- redis-data:/data # persist cache across container restarts
networks:
- ml-network
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 10s
timeout: 5s
retries: 3
# ── Service 3: PostgreSQL — Prediction Audit Log ──────────────────
postgres-db:
image: postgres:16-alpine
container_name: postgres-db
restart: unless-stopped
environment:
- POSTGRES_DB=predictions
- POSTGRES_USER=mluser
- POSTGRES_PASSWORD=${DB_PASSWORD} # from .env file — never hardcode!
volumes:
- postgres-data:/var/lib/postgresql/data # persistent storage
- ./sql/init.sql:/docker-entrypoint-initdb.d/init.sql # schema setup
networks:
- ml-network
healthcheck:
test: ["CMD-SHELL", "pg_isready -U mluser -d predictions"]
interval: 10s
timeout: 5s
retries: 5
# ── Service 4: Nginx Reverse Proxy ───────────────────────────────
nginx:
image: nginx:1.25-alpine
container_name: nginx-proxy
restart: unless-stopped
ports:
- "80:80"
- "443:443"
volumes:
- ./nginx/nginx.conf:/etc/nginx/nginx.conf:ro
- ./nginx/ssl:/etc/nginx/ssl:ro # TLS certificates from OCI Certificates Service
depends_on:
- fraud-api
networks:
- ml-network
# ── Shared network (containers communicate by service name) ─────────
networks:
ml-network:
driver: bridge
name: ml-internal-network
# ── Persistent data volumes ──────────────────────────────────────────
volumes:
redis-data:
driver: local
postgres-data:
driver: local
# ── Run the complete ML system ────────────────────────────────────────
# Create .env file first (never commit this to Git!)
cat > .env << EOF
ENVIRONMENT=development
OCI_REGION=ap-mumbai-1
APP_VERSION=2.1.0
LOG_LEVEL=INFO
DB_PASSWORD=your-secure-password-here
EOF
# Start ALL services (builds images if not already built)
docker compose up -d
# Starting ml-network
# Starting postgres-db ... done
# Starting redis-cache ... done
# Starting fraud-api ... done
# Starting nginx-proxy ... done
# Check all services are healthy
docker compose ps
# NAME SERVICE STATUS PORTS
# fraud-api fraud-api running 0.0.0.0:8083->8083/tcp
# redis-cache redis-cache running 0.0.0.0:6379->6379/tcp
# postgres-db postgres-db running 5432/tcp
# nginx-proxy nginx running 0.0.0.0:80->80, 0.0.0.0:443->443
# View logs from all services at once
docker compose logs -f
# View logs from just the API service
docker compose logs -f fraud-api
# Scale the API to 3 replicas (for load testing)
docker compose up -d --scale fraud-api=3
# Stop everything
docker compose down
# Stop everything AND delete volumes (CAREFUL: deletes all data!)
docker compose down -v
☁️ Section 9: Oracle Container Registry (OCIR) — Push & Pull on OCI
Once your Docker image is built and tested locally, the next step is pushing it to OCIR (Oracle Container Registry) so it can be pulled by OCI Compute instances, OKE clusters, or Container Instances.
🏗️ OCIR Structure
<region-key>.ocir.io / <tenancy-namespace> / <repository> : <tag>
Examples:
ap-mumbai-1.ocir.io/mytenancy/fraud-detection-api:2.1.0
eu-frankfurt-1.ocir.io/mytenancy/ml-models/fraud-api:latest
us-phoenix-1.ocir.io/mytenancy/fraud-api:prod-2025-03-10
OCI Region Keys:
ap-mumbai-1 → bom.ocir.io
eu-frankfurt-1 → fra.ocir.io
us-phoenix-1 → phx.ocir.io
ap-sydney-1 → syd.ocir.io
uk-london-1 → lhr.ocir.io
What this code does: Authenticates with OCIR using an OCI Auth Token (not your OCI password), tags the local Docker image with the full OCIR path, pushes it to OCIR, and verifies the push was successful. Also shows how to pull the image from OCIR onto any OCI instance.
Why OCI Auth Token instead of password: OCIR uses Auth Tokens for Docker registry authentication. These are separate from your OCI Console password and can be revoked independently — a security best practice. Auth Tokens are generated in OCI Console → User Settings → Auth Tokens.
# ════════════════════════════════════════════════════════════════
# STEP 1: SET YOUR OCI VARIABLES
# ════════════════════════════════════════════════════════════════
export OCI_REGION="ap-mumbai-1"
export OCI_TENANCY_NAMESPACE="mytenancynamespace" # found in OCI Console → Tenancy
export OCI_USERNAME="oracleidentitycloudservice/alice@company.com"
export OCI_AUTH_TOKEN="xxxxxxxxxxx" # from OCI Console → User Settings → Auth Tokens
export IMAGE_NAME="fraud-detection-api"
export IMAGE_TAG="2.1.0"
# Construct the full OCIR image path
export OCIR_IMAGE="${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${IMAGE_NAME}:${IMAGE_TAG}"
echo "OCIR image path: $OCIR_IMAGE"
# ap-mumbai-1.ocir.io/mytenancynamespace/fraud-detection-api:2.1.0
# ════════════════════════════════════════════════════════════════
# STEP 2: LOGIN TO OCIR
# ════════════════════════════════════════════════════════════════
# Docker login to OCIR
# Username format: / OR
# oracleidentitycloudservice/ (for federated users)
docker login ${OCI_REGION}.ocir.io \
--username "${OCI_TENANCY_NAMESPACE}/${OCI_USERNAME}" \
--password "${OCI_AUTH_TOKEN}"
# Login Succeeded ✅
# ════════════════════════════════════════════════════════════════
# STEP 3: TAG THE LOCAL IMAGE WITH OCIR PATH
# ════════════════════════════════════════════════════════════════
# Your local image:
docker images | grep fraud-detection-api
# fraud-detection-api 2.1.0 a9f3b2c1d8e4 612MB
# Tag it with full OCIR path
docker tag fraud-detection-api:${IMAGE_TAG} ${OCIR_IMAGE}
# Verify both tags exist locally
docker images | grep fraud
# fraud-detection-api 2.1.0 a9f3b2c1d8e4 612MB
# ap-mumbai-1.ocir.io/mytenancy/fraud-detection-api 2.1.0 a9f3b2c1d8e4 612MB
# (same image ID — no data was duplicated, just a new tag pointer)
# ════════════════════════════════════════════════════════════════
# STEP 4: PUSH TO OCIR
# ════════════════════════════════════════════════════════════════
docker push ${OCIR_IMAGE}
# The push refers to repository [ap-mumbai-1.ocir.io/mytenancy/fraud-detection-api]
# layer 1: Pushed — python:3.10-slim-buster base
# layer 2: Pushed — apt-get install build-essential
# layer 3: Pushed — pip install -r requirements.txt
# layer 4: Pushed — application code
# 2.1.0: digest: sha256:f4b8c2e1d9a7... size: 4152
# ✅ Successfully pushed to OCIR!
# Also push a 'latest' tag for convenience
docker tag ${OCIR_IMAGE} ${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${IMAGE_NAME}:latest
docker push ${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${IMAGE_NAME}:latest
# ════════════════════════════════════════════════════════════════
# STEP 5: PULL AND RUN ON OCI COMPUTE INSTANCE
# ════════════════════════════════════════════════════════════════
# SSH into your OCI Compute instance (Oracle Linux 8 or Ubuntu 22.04)
# ssh -i ~/.ssh/oci_rsa opc@
# On the OCI Compute instance, login to OCIR
docker login ${OCI_REGION}.ocir.io \
--username "${OCI_TENANCY_NAMESPACE}/${OCI_USERNAME}" \
--password "${OCI_AUTH_TOKEN}"
# Pull the image from OCIR
docker pull ${OCIR_IMAGE}
# Run the container on OCI
docker run -d \
--name fraud-api-prod \
--restart unless-stopped \
-p 8083:8083 \
-e ENVIRONMENT=production \
-e OCI_REGION=ap-mumbai-1 \
-e LOG_LEVEL=INFO \
${OCIR_IMAGE}
# Verify it's running
docker ps
curl http://localhost:8083/health
# {"status":"healthy","environment":"production","oci_region":"ap-mumbai-1"} ✅
⚡ Section 10: Deploying on OCI Container Instances
OCI Container Instances is Oracle's serverless container service. You provide a Docker image from OCIR, specify the shape (vCPU + RAM), and OCI handles all the underlying infrastructure. No VM to manage. No OS patching. No SSH required. Billed per second of actual use. Perfect for ML APIs! 🎯
What this code does: Uses the OCI CLI to create a Container Instance — Oracle's serverless container runtime. The CLI call specifies the compute shape, networking, OCIR image path, environment variables, and health check configuration. After creation, the command waits for the container to become ACTIVE.
Why OCI Container Instance for ML APIs: Unlike a full OCI Compute VM (which runs 24/7 and costs money even when idle), OCI Container Instances scale to zero when not in use, start in seconds when a request arrives, and bill per-second. Ideal for ML APIs with variable traffic patterns.
# ── Deploy to OCI Container Instance using OCI CLI ───────────────────
# Ensure OCI CLI is installed and configured
oci --version
# Oracle Cloud Infrastructure CLI 3.x.x
# ── Get required OCIDs ───────────────────────────────────────────────
# You need these from OCI Console:
export OCI_COMPARTMENT_ID="ocid1.compartment.oc1..aaaaa..."
export OCI_SUBNET_ID="ocid1.subnet.oc1.ap-mumbai-1.aaaaa..."
export OCI_IMAGE_URL="ap-mumbai-1.ocir.io/mytenancy/fraud-detection-api:2.1.0"
# ── Create the Container Instance ────────────────────────────────────
oci container-instances container-instance create \
--compartment-id $OCI_COMPARTMENT_ID \
--display-name "fraud-detection-api-prod" \
--shape "CI.Standard.E4.Flex" \
--shape-config '{"ocpus": 1, "memoryInGBs": 4}' \
--availability-domain "Uocm:AP-MUMBAI-1-AD-1" \
--vnics "[{
\"subnetId\": \"$OCI_SUBNET_ID\",
\"displayName\": \"fraud-api-vnic\",
\"isPublicIpAssigned\": true
}]" \
--containers "[{
\"displayName\": \"fraud-api-container\",
\"imageUrl\": \"$OCI_IMAGE_URL\",
\"resourceConfig\": {
\"vcpusLimit\": 1.0,
\"memoryLimitInGBs\": 3.0
},
\"environmentVariables\": {
\"ENVIRONMENT\": \"production\",
\"OCI_REGION\": \"ap-mumbai-1\",
\"LOG_LEVEL\": \"INFO\"
},
\"healthChecks\": [{
\"healthCheckType\": \"HTTP\",
\"port\": 8083,
\"path\": \"/health\",
\"intervalInSeconds\": 30,
\"timeoutInSeconds\": 10,
\"failureThreshold\": 3,
\"successThreshold\": 1
}]
}]" \
--wait-for-state ACTIVE
# Output:
# Action completed. Waiting until the resource has entered state: ('ACTIVE',)
# {
# "data": {
# "display-name": "fraud-detection-api-prod",
# "lifecycle-state": "ACTIVE",
# "vnics": [{"public-ip": "140.238.x.x"}]
# }
# }
# ✅ Container Instance is ACTIVE!
# Test the deployed API
curl http://140.238.x.x:8083/health
# {"status":"healthy","environment":"production","oci_region":"ap-mumbai-1"}
🔄 Section 11: Docker in CI/CD with OCI DevOps Pipelines
In production, you never build and push Docker images manually. The entire process is automated through a CI/CD pipeline. Every time a developer pushes code to Git, the pipeline automatically: builds the Docker image, tests it, pushes to OCIR, and deploys to OCI.
What this code does: Defines an OCI DevOps build pipeline specification that automatically: runs unit tests, builds the Docker image, pushes it to OCIR, and passes the image path to a deployment pipeline. This entire process triggers on every git push to the main branch.
Why this matters: Manual deployments are error-prone and don't scale. With this pipeline, every code change produces a tested, tagged, and deployed container within minutes — with full audit trail in OCI DevOps.
# build_spec.yaml — OCI DevOps Build Pipeline Specification
# Place this file at the root of your repository
# Triggers automatically on: git push to main/release branches
version: 0.1
component: build
timeoutInSeconds: 1800 # 30 minute max build time
env:
# Build-time variables (injected by OCI DevOps)
variables:
APP_NAME: "fraud-detection-api"
OCI_REGION: "ap-mumbai-1"
# References to OCI Vault secrets (never hardcode credentials!)
vaultVariables:
OCI_AUTH_TOKEN: "ocid1.vaultsecret.oc1.ap-mumbai-1.aaaaa..."
DB_PASSWORD: "ocid1.vaultsecret.oc1.ap-mumbai-1.bbbbb..."
exportedVariables:
- DOCKER_IMAGE_PATH # passed to deployment pipeline
- APP_VERSION
steps:
# ── Step 1: Install dependencies for testing ───────────────────────
- type: Command
name: "Install test dependencies"
command: |
pip install pytest pytest-asyncio httpx
onFailure:
- type: Command
command: echo "❌ Dependency install failed"
# ── Step 2: Run unit tests ─────────────────────────────────────────
- type: Command
name: "Run unit tests"
command: |
echo "🧪 Running unit tests..."
pytest tests/ -v --tb=short
echo "✅ All tests passed!"
# ── Step 3: Set version tag from git commit ────────────────────────
- type: Command
name: "Set build version"
command: |
export APP_VERSION=$(git rev-parse --short HEAD)
echo "Build version: $APP_VERSION"
# ── Step 4: Login to OCIR ──────────────────────────────────────────
- type: Command
name: "Login to Oracle Container Registry"
command: |
echo "$OCI_AUTH_TOKEN" | docker login \
${OCI_REGION}.ocir.io \
--username "${OCI_TENANCY_NAMESPACE}/${OCI_BUILD_RUNNER_USER}" \
--password-stdin
echo "✅ Logged into OCIR"
# ── Step 5: Build Docker image ────────────────────────────────────
- type: Command
name: "Build Docker image"
command: |
export OCIR_PATH="${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${APP_NAME}"
docker build \
-t ${OCIR_PATH}:${APP_VERSION} \
-t ${OCIR_PATH}:latest \
--build-arg environment=production \
--no-cache \
.
echo "✅ Docker image built: ${OCIR_PATH}:${APP_VERSION}"
export DOCKER_IMAGE_PATH="${OCIR_PATH}:${APP_VERSION}"
# ── Step 6: Security scan (Trivy vulnerability scanner) ───────────
- type: Command
name: "Scan image for vulnerabilities"
command: |
# Install Trivy (Docker image scanner)
curl -sfL https://raw.githubusercontent.com/aquasecurity/trivy/main/contrib/install.sh | sh
./bin/trivy image \
--severity HIGH,CRITICAL \
--exit-code 1 \
${DOCKER_IMAGE_PATH}
echo "✅ No HIGH/CRITICAL vulnerabilities found"
# ── Step 7: Push to OCIR ──────────────────────────────────────────
- type: Command
name: "Push to Oracle Container Registry"
command: |
docker push ${DOCKER_IMAGE_PATH}
docker push ${OCI_REGION}.ocir.io/${OCI_TENANCY_NAMESPACE}/${APP_NAME}:latest
echo "✅ Image pushed to OCIR: ${DOCKER_IMAGE_PATH}"
outputArtifacts:
- name: docker_image
type: DOCKER_IMAGE
location: ${DOCKER_IMAGE_PATH}
🏗️ Section 12: Multi-Stage Builds — Slim Production Images
The single biggest way to reduce Docker image size for ML models
is to use Multi-Stage Builds.
This technique uses multiple FROM statements in one Dockerfile —
one stage for building/compiling, a final stage that only copies
what's needed to run.
What this code does: Uses two build stages: a "builder" stage that installs all compilation tools and builds Python wheels, and a lean "production" stage that only copies the final installed packages — no compilers, no build tools, no cache.
Result: Single-stage image: ~900 MB. Multi-stage image: ~380 MB. Faster pull from OCIR. Smaller attack surface. Lower OCI storage costs.
# ─────────────────────────────────────────────────────────────────────
# Dockerfile.multistage — Two-stage build for minimal production image
# ─────────────────────────────────────────────────────────────────────
# ══════════════════════════════════════════════════
# STAGE 1: Builder — install everything, compile all wheels
# This stage is NOT included in the final image!
# ══════════════════════════════════════════════════
FROM python:3.10-slim-buster AS builder
WORKDIR /build
# Install build tools (compiler for C extensions in numpy, scipy etc.)
RUN apt-get update && \
apt-get install --no-install-recommends -y \
build-essential gcc g++ \
&& apt-get clean && rm -rf /tmp/*
# Copy and install requirements — build wheels here
COPY requirements.txt .
# Build wheels into /build/wheels directory
# --wheel-dir stores compiled .whl files
RUN pip install --upgrade pip && \
pip wheel --no-cache-dir \
--no-deps \
--wheel-dir /build/wheels \
-r requirements.txt
# ══════════════════════════════════════════════════
# STAGE 2: Production — lean runtime image
# ONLY this stage ends up in the final Docker image
# ══════════════════════════════════════════════════
FROM python:3.10-slim-buster AS production
WORKDIR /app
# Install ONLY runtime OS dependencies (no compilers!)
# libgomp1 = required by XGBoost / LightGBM for parallel processing
RUN apt-get update && \
apt-get install --no-install-recommends -y \
libgomp1 curl \
&& apt-get clean && rm -rf /tmp/*
# Copy pre-built wheels from builder stage
COPY --from=builder /build/wheels /wheels
# Install from local wheels (fast! no internet download needed)
RUN pip install --no-cache-dir --no-index \
--find-links=/wheels \
/wheels/*.whl && \
rm -rf /wheels # clean up wheels after install
# Copy application code
COPY . /app
EXPOSE 8083
ENV PYTHONPATH="/app"
ENV OCI_REGION="ap-mumbai-1"
ARG environment=production
ENV ENVIRONMENT=$environment
RUN chmod +x /app/entrypoint.sh
ENTRYPOINT ["/app/entrypoint.sh"]
# ─────────────────────────────────────────────────────────────────────
# Build: docker build -f Dockerfile.multistage -t fraud-api:slim .
# Result: ~900MB → ~380MB (58% smaller!) ✅
# ─────────────────────────────────────────────────────────────────────
🔥 Section 13: GPU Support in Docker — Deep Learning on OCI GPU Shapes
Oracle OCI offers powerful GPU compute shapes for deep learning:
- VM.GPU3.1 — 1× NVIDIA V100 (16GB HBM2)
- VM.GPU.A10.1 — 1× NVIDIA A10 (24GB GDDR6)
- BM.GPU.A100-v2.8 — 8× NVIDIA A100 (80GB HBM2e each)
- BM.GPU4.8 — 8× NVIDIA A100 (40GB HBM2)
What this code does: Builds a Docker image using NVIDIA's CUDA base image instead of plain Python. The CUDA base gives direct access to GPU hardware. Installs GPU-enabled PyTorch and sets the NVIDIA runtime flag. Includes a model inference test at build time to verify GPU access.
OCI requirement: To use GPU in Docker on OCI, the Compute instance must have the NVIDIA Container Toolkit installed, and Docker must be run with
--gpus all or --runtime=nvidia.
# Dockerfile.gpu — For OCI VM.GPU.A10.1 and BM.GPU.A100-v2.8 shapes
# Uses NVIDIA CUDA base image instead of plain Python
# Base image with CUDA 12.1 + cuDNN 8 (matches OCI GPU driver versions)
FROM nvcr.io/nvidia/pytorch:24.03-py3
# OR use official CUDA image and install PyTorch manually:
# FROM nvidia/cuda:12.1.0-cudnn8-runtime-ubuntu22.04
WORKDIR /app
# Install Python dependencies (PyTorch already in base image)
COPY requirements-gpu.txt /app/requirements.txt
RUN pip install --upgrade pip && \
pip install --no-cache-dir -r requirements.txt
COPY . /app
EXPOSE 8083
ENV PYTHONPATH="/app"
ENV OCI_REGION="ap-mumbai-1"
ENV CUDA_VISIBLE_DEVICES="0" # use first GPU
RUN chmod +x /app/entrypoint.sh
ENTRYPOINT ["/app/entrypoint.sh"]
# ── Install NVIDIA Container Toolkit on OCI GPU instance ─────────────
# (Run once on the OCI Compute instance — Oracle Linux 8)
# Add NVIDIA Container Toolkit repository
curl -s -L https://nvidia.github.io/libnvidia-container/stable/rpm/nvidia-container-toolkit.repo \
| sudo tee /etc/yum.repos.d/nvidia-container-toolkit.repo
# Install the toolkit
sudo dnf install -y nvidia-container-toolkit
# Configure Docker to use NVIDIA runtime
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
# Verify GPU is accessible in Docker
docker run --rm --gpus all nvidia/cuda:12.1.0-base-ubuntu22.04 nvidia-smi
# +-----------------------------------------------------------------------------+
# | NVIDIA-SMI 535.54.03 Driver Version: 535.54.03 CUDA Version: 12.2 |
# |-------------------------------+----------------------+----------------------+
# | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
# | 0 NVIDIA A10 Off | 00000000:00:04.0 Off | 0 |
# +-------------------------------+----------------------+----------------------+
# ✅ GPU accessible in Docker container!
# ── Build and run the GPU-enabled ML container ───────────────────────
docker build -f Dockerfile.gpu -t fraud-dl-api:gpu-v1.0 .
docker run -d \
--gpus all \ # Enable GPU access
--runtime=nvidia \ # Use NVIDIA runtime
-p 8083:8083 \
-e ENVIRONMENT=production \
-e CUDA_VISIBLE_DEVICES=0 \
fraud-dl-api:gpu-v1.0
🚫 Section 14: Common Mistakes & Best Practices
- ✅ Always use specific version tags:
python:3.10-slim-busternotpython:latest - ✅ Copy
requirements.txtBEFORE other code — maximize layer caching - ✅ Use
.dockerignoreto exclude.git/,.env, notebooks, and large raw data files - ✅ Chain RUN commands and clean up in the SAME instruction to minimize layer size
- ✅ Use multi-stage builds for production — reduces image size by 40–60%
- ✅ Always expose the
/healthendpoint for OCI Load Balancer health probes - ✅ Store secrets in OCI Vault — never in ENV variables in the Dockerfile
- ✅ Tag images with both version (2.1.0) AND commit hash (a9f3b2c) for traceability
- ✅ Scan images for vulnerabilities with Trivy before pushing to OCIR
- ✅ Use
--restart unless-stoppedfor production containers on OCI Compute
- ❌ Don't use
:latesttags in production — always pin to specific versions - ❌ Don't copy secrets into Docker images — they're visible in
docker inspectand image layers - ❌ Don't run containers as root — use
USERto create a non-root user - ❌ Don't install dev tools (git, vim, wget) in production images — unnecessary attack surface
- ❌ Don't run
apt-get cleanin a separate RUN from the install — it won't reduce image size - ❌ Don't ignore the health check — OCI Container Instances and OKE rely on it for traffic routing
- ❌ Don't load ML models inside API endpoint functions — load ONCE at startup, reuse for all requests
- ❌ Don't push 5 GB model files inside the Docker image — store models in OCI Object Storage and download at startup
- ❌ Don't skip
.dockerignore— accidentally copying.git/adds hundreds of MB to the image
# Add this to your Dockerfile before ENTRYPOINT:
# Create a non-root user for security
RUN addgroup --gid 1001 mlgroup && \
adduser --uid 1001 --gid 1001 --disabled-password mluser
# Set file ownership
RUN chown -R mluser:mlgroup /app
# Switch to non-root user
USER mluser
ENTRYPOINT ["/app/entrypoint.sh"]
Running as non-root is required by many OCI security policies and
prevents privilege escalation attacks if the container is compromised.
🏆 High-Level Summary
- 🔹 Docker = the standardized shipping container for software. Eliminates "works on my machine" forever.
- 🔹 Image = read-only blueprint (recipe). Container = running instance (the pizza). OCIR = Oracle's container registry (the library).
- 🔹 FROM picks your base. WORKDIR sets your home. RUN executes build-time commands. COPY brings in files.
- 🔹 ENV sets runtime variables. ARG sets build-time variables. EXPOSE documents ports. ENTRYPOINT is the on switch.
- 🔹 Layer caching trick: Copy requirements.txt BEFORE your code. pip install runs from cache on code-only changes. Rebuild goes from 10 min → 5 sec.
- 🔹 Docker Compose = run multiple containers as one system. API + Redis + PostgreSQL + Nginx with one command.
- 🔹 OCIR (Oracle Container Registry) = push images from local → pull on any OCI region. Authenticated via OCI Auth Token.
- 🔹 OCI Container Instances = serverless containers. No VM management. Per-second billing. Health check required.
- 🔹 Multi-stage builds = separate build and runtime stages. 900 MB → 380 MB images. Faster OCIR pulls.
- 🔹 GPU Docker on OCI = NVIDIA CUDA base image + NVIDIA Container Toolkit +
--gpus allflag. - 🔹 OCI DevOps pipeline = auto build → test → scan → push to OCIR → deploy on every git push.
- 🔹 Security rule #1: NEVER put secrets in Dockerfile ENV. Use OCI Vault + runtime injection.
Keep building, keep containerizing, keep shipping! 🐼✨
Comments
Post a Comment