Skip to main content

Deploy Fine-Tuned AI Models on OCI: A Practical Guide to Model Deployment

Calculating read time…

Imagine you just baked the most amazing cake. 🎂 But baking the cake is only the beginning!

You still need to: put it in a nice box 📦, check it hasn't gone bad ✅, label it with ingredients 🏷️, ship it to the shop 🚚, and finally open the shop counter so customers can buy it! 🏪

In OCI AI, turning a trained Python model into a live, working REST API follows the exact same journey — and there are 6 key methods that handle each step for you!





🗺️ The Big Picture — The 6 Steps of Model Deployment

🔄 Complete OCI Model Deployment Journey

Step 1
📦 .prepare()
Package the model
→
Step 2
📝 score.py
Write the menu
→
Step 3
✅ .verify()
Test locally first
→
Step 4
🏛️ .save()
Upload to catalog
→
Step 5
🚀 .deploy()
Go live!
→
Step 6
📊 .summary_status()
Health check

⬆️ Every deployed OCI model goes through ALL 6 steps in this exact order!

⭐ Before We Start — Setup:
All examples below assume you are working inside an OCI Data Science Notebook with ADS installed. Authentication is set via ads.set_auth("resource_principal"). We'll use a customer churn classifier as our example model throughout!

🏋️ First Things First — Train a Model to Work With

Before using any of the 6 methods, you need a trained model. Let's quickly train one so we have something real to deploy!

📝 What the code below does:
This trains a simple Random Forest model to predict whether a telecom customer will cancel their service (churn) or not. Think of it as teaching an AI student 🎓 by showing it 10,000 past examples. We'll use this trained model for ALL the deployment steps that follow!
import ads
import pandas as pd
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder

ads.set_auth("resource_principal")

# ---- Create sample churn dataset ----
np.random.seed(42)
n = 2000

data = pd.DataFrame({
    'tenure'          : np.random.randint(1, 72, n),
    'monthly_charges' : np.round(np.random.uniform(20, 120, n), 2),
    'num_support_calls': np.random.randint(0, 10, n),
    'contract_type'   : np.random.choice(['month-to-month', 'one-year', 'two-year'], n),
    'has_internet'    : np.random.choice([0, 1], n),
    'churn'           : np.random.choice([0, 1], n, p=[0.75, 0.25])
})

# Encode the contract_type text column to numbers
le = LabelEncoder()
data['contract_type'] = le.fit_transform(data['contract_type'])

# Split features (X) and label (y)
X = data.drop('churn', axis=1)
y = data['churn']

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# Train the model
clf = RandomForestClassifier(n_estimators=100, random_state=42)
clf.fit(X_train, y_train)

accuracy = clf.score(X_test, y_test)
print(f"✅ Model trained!")
print(f"   Algorithm : Random Forest")
print(f"   Test Accuracy : {accuracy:.4f}")
print(f"   Features  : {list(X.columns)}")
print(f"\n   Ready for the 6-step deployment journey! 🚀")

Output:

✅ Model trained!
   Algorithm : Random Forest
   Test Accuracy : 0.7625
   Features  : ['tenure', 'monthly_charges', 'num_support_calls', 'contract_type', 'has_internet']

   Ready for the 6-step deployment journey! 🚀

📦 Step 1: .prepare() — Pack Your Model into a Box

🧠 What is .prepare()?

Imagine you baked a cake at home. 🎂 To ship it to a bakery, you can't just hand over the raw cake. You need to: put it in a proper bakery box, add a label, include instructions for serving it, and list all ingredients.

.prepare() does exactly this for your AI model! It creates a complete deployment package — a folder with everything OCI needs to serve your model as an API.

📂 What Does .prepare() Create?

🗂️ The Artifact Folder Created by .prepare()

📁 churn_model_artifact/
    ├── 🐍 score.py            ← The inference brain (YOU customize this!)
    ├── 🤖 model.pkl           ← Your trained model (auto-saved)
    ├── 📋 runtime.yaml        ← Python environment definition
    ├── 📄 input_schema.json   ← What inputs the model expects
    └── 📄 output_schema.json ← What the model outputs
📝 What the code below does:
This wraps your trained sklearn model inside an ADS SklearnModel wrapper and runs .prepare() to create the deployment package folder. It automatically saves your model file, generates the score.py template, detects your Python environment, and creates schema files — all automatically! You just tell it which folder to use. 📁
from ads.model.framework.sklearn_model import SklearnModel
from ads.common.object_storage_details import ObjectStorageDetails
import tempfile, os

# Step 1: Define where to store the artifact folder
artifact_dir = "./churn_model_artifact"

# Step 2: Wrap your trained model inside ADS's SklearnModel
model_artifact = SklearnModel(
    estimator    = clf,        # your trained RandomForestClassifier
    artifact_dir = artifact_dir
)

# Step 3: Run .prepare() — this does all the heavy lifting!
model_artifact.prepare(
    inference_conda_env = "generalml_p38_cpu_v1",
    # ↑ This is an OCI-managed conda environment with scikit-learn, pandas, etc.
    # OCI has many ready-made environments so you don't build one from scratch!

    training_dataset    = X_train,  # Optional: helps ADS auto-detect input schema
    label_column        = "churn",  # The column we're predicting
    force_overwrite     = True      # Overwrite if folder already exists

    # Optional: provide sample input so OCI can validate data shape later
    # X_sample = X_test.iloc[:5]
)

# Step 4: Check what was created
print("✅ .prepare() completed! Here's what was created:\n")
for f in os.listdir(artifact_dir):
    size = os.path.getsize(os.path.join(artifact_dir, f))
    print(f"   📄 {f:<30 bytes="" code="" size:="">

Output:

✅ .prepare() completed! Here's what was created:

   📄 score.py                       (1,842 bytes)
   📄 model.pkl                      (2,341,120 bytes)
   📄 runtime.yaml                   (312 bytes)
   📄 input_schema.json              (687 bytes)
   📄 output_schema.json             (223 bytes)

🔧 Key Parameters of .prepare()

Parameter What it does Required?
inference_conda_env Which Python environment to use when serving predictions ✅ Yes
training_dataset Helps ADS auto-generate input schema from your feature columns ⚪ Optional
X_sample Sample input rows for schema detection and local testing ⚪ Optional but recommended
force_overwrite If True, deletes and recreates the artifact folder if it exists ⚪ Optional
use_case_type Hints to OCI what kind of ML task it is (classification, regression, etc.) ⚪ Optional but helpful
✅ DO: Always pass X_sample or training_dataset so ADS can auto-generate accurate input/output schemas. These schemas act as a contract between your model and any application calling it!
❌ DON'T: Skip .prepare() and try to call .save() directly. Without the artifact folder, there's nothing to upload! OCI will throw an error. 🚫

🧰 Framework-Specific Model Wrappers

ADS has different model wrapper classes depending on what library you used to train:

  • Scikit-learn → SklearnModel
  • TensorFlow / Keras → TensorFlowModel
  • PyTorch → PyTorchModel
  • XGBoost → XGBoostModel
  • LightGBM → LightGBMModel
  • HuggingFace Transformers → HuggingFacePipelineModel
  • Any other framework → GenericModel (universal fallback)
⭐ Tip:
HuggingFacePipelineModel is now the most popular wrapper! It lets you deploy fine-tuned LLMs (like Llama 3, Mistral) with the same prepare → verify → save → deploy workflow. Same 6 steps, any model! 🤖

📝 Step 2: score.py — The Brain of Your Deployment

🧠 What is score.py?

When OCI receives a prediction request from an application, it doesn't know what to DO with the incoming data. Should it preprocess it? Which model file to load? What format to return?

score.py is the instruction manual you write that answers all these questions! It's the heart and brain of your deployment.

Think of it like this: your trained model is a chef 👨‍🍳. score.py is the waiter 🧑‍🍽️ who: takes the customer's order → brings it to the chef → takes the food back → and serves it properly!

📋 The 3 Mandatory Functions in score.py

🔑 The 3 Functions Every score.py Must Have

1. load_model()
Runs ONCE when the server starts.
Loads the model file into memory.
2. predict()
Runs for EVERY request.
Takes input → returns prediction.
3. pre_inference() / post_inference()
Optional hooks.
Clean input before / format output after.
📝 What the code below does:
This is a complete, production-ready score.py file for our churn model. It defines how to load the model when OCI starts the server, how to clean incoming data, how to run the prediction, and how to format the response. This file runs inside OCI's servers — so write it carefully! ⚠️
# ============================================================
# score.py — The Inference Brain for Customer Churn Model
# This file lives inside your artifact folder.
# OCI calls these functions automatically when predictions are requested.
# ============================================================

import json
import os
import pickle
import pandas as pd
import numpy as np
import logging

# Set up logging so you can debug issues in OCI Log Explorer
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

# ----------------------------------------------------------------
# FUNCTION 1: load_model()
# Called ONCE when the deployment server starts.
# Loads the model from disk into memory.
# Think of this as the chef arriving at the kitchen and setting up! 👨‍🍳
# ----------------------------------------------------------------
def load_model(model_file_name="model.pkl"):
    """
    Loads the saved model file from the artifact directory.
    Returns the model object ready for predictions.
    """
    model_dir = os.path.dirname(os.path.realpath(__file__))
    model_path = os.path.join(model_dir, model_file_name)

    logger.info(f"Loading model from: {model_path}")

    with open(model_path, "rb") as f:
        model = pickle.load(f)

    logger.info(f"✅ Model loaded successfully! Type: {type(model).__name__}")
    return model


# ----------------------------------------------------------------
# FUNCTION 2: predict()
# Called for EVERY prediction request.
# Receives the raw input data → returns predictions.
# Think of this as the waiter taking an order and delivering food! 🧑‍🍽️
# ----------------------------------------------------------------
def predict(data, model=load_model()):
    """
    Main prediction function.

    data  = the incoming JSON payload from the API caller
    model = the loaded model object from load_model()

    Returns a dict with predictions and probabilities.
    """
    logger.info(f"Received prediction request. Data keys: {list(data.keys())}")

    # Step A: Parse and validate incoming data
    # Callers can send data as a list of dicts or as a dict of lists
    try:
        if isinstance(data, dict) and "data" in data:
            # Format: {"data": [{"tenure": 24, "monthly_charges": 89.5, ...}]}
            input_df = pd.DataFrame(data["data"])
        elif isinstance(data, list):
            # Format: [{"tenure": 24, "monthly_charges": 89.5, ...}]
            input_df = pd.DataFrame(data)
        else:
            input_df = pd.DataFrame([data])

        logger.info(f"Parsed input shape: {input_df.shape}")

    except Exception as e:
        logger.error(f"❌ Failed to parse input: {e}")
        return {"error": f"Invalid input format: {str(e)}"}

    # Step B: Ensure columns are in the correct order
    expected_columns = [
        'tenure', 'monthly_charges', 'num_support_calls',
        'contract_type', 'has_internet'
    ]

    missing = set(expected_columns) - set(input_df.columns)
    if missing:
        return {"error": f"Missing required columns: {missing}"}

    input_df = input_df[expected_columns]  # Reorder to match training order

    # Step C: Run predictions
    predictions  = model.predict(input_df).tolist()
    probabilities = model.predict_proba(input_df).tolist()

    # Step D: Format a clear, readable response
    results = []
    for i, (pred, prob) in enumerate(zip(predictions, probabilities)):
        results.append({
            "row_index"         : i,
            "prediction"        : int(pred),
            "prediction_label"  : "Will Churn ⚠️" if pred == 1 else "Will Stay ✅",
            "probability_churn" : round(prob[1], 4),
            "probability_stay"  : round(prob[0], 4),
            "confidence_pct"    : f"{max(prob)*100:.1f}%"
        })

    logger.info(f"✅ Predictions complete for {len(results)} row(s).")
    return {"predictions": results, "model_version": "churn_v1", "total_rows": len(results)}
✅ DO: Always add logging to score.py! When something breaks in production, the logs in OCI Log Explorer will be your only window into what went wrong. logger.info() is your best debugging friend! 🔍
❌ DON'T: Hardcode file paths or model names in score.py. Always use os.path.dirname(os.path.realpath(__file__)) to get the correct directory — the file runs from a different location in OCI! 📁

⚡ Advanced score.py Patterns

Many teams use advanced patterns inside score.py:

📝 What the code below does:
This shows THREE advanced patterns you might add to your score.py: (1) Batch size limiting to prevent overload, (2) Input type casting to fix data type mismatches automatically, (3) Adding a drift warning flag when incoming data looks suspicious. 🚨
# ============================================================
# ADVANCED score.py PATTERNS — Add these to your predict() function
# ============================================================

import numpy as np

# Pattern 1: Limit batch size to prevent server overload
MAX_BATCH_SIZE = 500

def predict_with_batch_limit(data, model=load_model()):
    if len(data.get("data", [])) > MAX_BATCH_SIZE:
        return {"error": f"Batch size exceeds limit of {MAX_BATCH_SIZE} rows."}
    return predict(data, model)


# Pattern 2: Auto-cast incoming data types to prevent type errors
def safe_cast_input(input_df):
    """
    Sometimes API callers send numbers as strings ('89.5' instead of 89.5).
    This function automatically converts them to the right types.
    """
    type_map = {
        'tenure'           : int,
        'monthly_charges'  : float,
        'num_support_calls': int,
        'contract_type'    : int,
        'has_internet'     : int
    }
    for col, dtype in type_map.items():
        if col in input_df.columns:
            input_df[col] = input_df[col].astype(dtype)
    return input_df


# Pattern 3: Add a drift warning to the response
def check_input_drift(input_df):
    """
    Quick sanity check — if monthly_charges is way outside the expected range,
    warn the caller that the input may have drifted from training distribution.
    """
    EXPECTED_CHARGE_MIN = 20
    EXPECTED_CHARGE_MAX = 120

    unusual = input_df[
        (input_df['monthly_charges'] < EXPECTED_CHARGE_MIN) |
        (input_df['monthly_charges'] > EXPECTED_CHARGE_MAX)
    ]

    if len(unusual) > 0:
        return {
            "drift_warning"  : True,
            "warning_message": f"{len(unusual)} row(s) have monthly_charges outside training range ({EXPECTED_CHARGE_MIN}-{EXPECTED_CHARGE_MAX}). Predictions may be less reliable.",
            "affected_rows"  : unusual.index.tolist()
        }
    return {"drift_warning": False}

✅ Step 3: .verify() — Test Before You Ship!

🧠 What is .verify()?

Imagine you're about to ship thousands of boxes of your product worldwide. 📦 Before loading the trucks, a quality inspector opens a few boxes and checks: Does it open properly? Is everything inside? Does it work?

.verify() is that quality inspector! ✅ It runs your score.py locally on your machine before uploading anything to OCI. This catches bugs early — when it's cheap to fix them!

🔄 What .verify() Actually Does Step by Step

🔍 Inside .verify() — The Checklist

  1. 🔄 Simulates the exact OCI runtime environment locally
  2. 📂 Checks that the artifact folder has all required files
  3. 🐍 Verifies score.py has load_model() and predict() functions
  4. 📥 Passes your sample input through predict()
  5. 📤 Validates the output can be serialized to JSON
  6. 🧐 Reports any errors with helpful messages
📝 What the code below does:
This runs .verify() with a few sample rows of input data. It simulates what will happen when your live API receives a request — without actually deploying anything to OCI yet. Think of it as a dress rehearsal 🎭 before the big performance!
# Step: Create sample input data — exactly what a real API caller would send
sample_input = {
    "data": [
        {
            "tenure"            : 6,
            "monthly_charges"   : 89.50,
            "num_support_calls" : 4,
            "contract_type"     : 0,       # 0 = month-to-month
            "has_internet"      : 1
        },
        {
            "tenure"            : 48,
            "monthly_charges"   : 45.00,
            "num_support_calls" : 0,
            "contract_type"     : 2,       # 2 = two-year
            "has_internet"      : 1
        },
        {
            "tenure"            : 24,
            "monthly_charges"   : 110.00,
            "num_support_calls" : 2,
            "contract_type"     : 1,       # 1 = one-year
            "has_internet"      : 0
        }
    ]
}

# Run the local verification!
# This is totally local — nothing is sent to OCI yet.
print("🔍 Running .verify() — local pre-flight check...\n")

prediction = model_artifact.verify(sample_input)

print("\n✅ .verify() passed! Here are the results:")
print(f"   Keys in response : {list(prediction.keys())}")
print(f"   Total rows       : {prediction.get('total_rows', '?')}")
print(f"\n📊 Predictions:")

for result in prediction.get("predictions", []):
    print(f"   Row {result['row_index']}: {result['prediction_label']} "
          f"(Confidence: {result['confidence_pct']}, "
          f"Churn Prob: {result['probability_churn']})")

Output:

🔍 Running .verify() — local pre-flight check...

✅ .verify() passed! Here are the results:
   Keys in response : ['predictions', 'model_version', 'total_rows']
   Total rows       : 3

📊 Predictions:
   Row 0: Will Churn ⚠️ (Confidence: 72.0%, Churn Prob: 0.72)
   Row 1: Will Stay ✅ (Confidence: 86.0%, Churn Prob: 0.14)
   Row 2: Will Stay ✅ (Confidence: 68.0%, Churn Prob: 0.32)
✅ Golden Rule:
NEVER skip .verify()! Every single time you change your score.py — even one line — run .verify() again. Bugs caught here take seconds to fix. Bugs discovered after deployment take hours to fix and cause production outages! ⏰
❌ Common Mistake:
Passing the wrong data format to .verify(). Use the exact same JSON structure that your real API callers will use. If they'll send {"data": [...]}, then verify with that exact format too!

🐛 Common .verify() Errors and Fixes

Error Message What it Means Fix
AttributeError: 'NoneType' Model didn't load — wrong file path in load_model() Use os.path.realpath(__file__)
KeyError: 'tenure' Column name mismatch between input data and model Check column names match exactly (case-sensitive!)
ModuleNotFoundError A Python library in score.py isn't in the conda env Add it to your conda env or use a different env
JSONSerializationError Your predict() returns numpy types (not JSON-safe) Wrap predictions with .tolist() or int()

🏛️ Step 4: .save() — Upload to OCI Model Catalog

🧠 What is .save()?

After verifying your model works locally, it's time to upload it to the cloud! .save() takes your entire artifact folder and uploads it to the OCI Model Catalog.

Think of it like submitting your final project to a shared Google Drive 📁 that your entire team (and OCI) can access anytime, from anywhere.

📝 What the code below does:
This uploads your artifact folder to OCI Model Catalog, gives your model a clear name and description, and returns a unique Model OCID — the permanent ID for this model version. Save that OCID! You'll need it for deployment and future reference. 🔑
# Step 4: Save the model to OCI Model Catalog

print("📤 Uploading model to OCI Model Catalog...\n")

model_id = model_artifact.save(

    # How this model version will appear in OCI Console
    display_name = "Customer_Churn_RF_v1",

    # A clear description helps your team understand what this model does
    description  = (
        "Random Forest classifier for customer churn prediction. "
        "Trained on Jan 2026 data. Test Accuracy: 76.25%. "
        "Features: tenure, monthly_charges, support_calls, contract_type, internet."
    ),

    # Which OCI project and compartment to store it in
    project_id     = "ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID",
    compartment_id = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID",

    # Override settings (optional)
    ignore_pending_changes = True,  # Save even if there are minor uncommitted changes
    overwrite_existing_artifact = False  # Don't overwrite — create a new version
)

print(f"✅ Model saved to OCI Model Catalog!")
print(f"   Model OCID  : {model_id}")
print(f"   Name        : Customer_Churn_RF_v1")
print(f"   Status      : Active")
print(f"\n   ⚠️  IMPORTANT: Save this OCID — you'll need it for deployment!")
print(f"   {model_id}")

📋 Adding Rich Metadata Before .save()

Before calling .save(), you can attach useful information to your model that makes it easier to find, track, and audit later. This is especially important in enterprise teams with many model versions!

📝 What the code below does:
This adds custom "tags" to the model before saving — like attaching sticky notes 📌 with important details: who trained it, when, what data it used, accuracy achieved. In the OCI Console, your team can search and filter models by these metadata fields!
# Add custom metadata BEFORE calling .save()
# These appear as searchable labels in OCI Model Catalog

# Custom key-value pairs — any info you want to track
model_artifact.metadata_custom.add(
    key         = "accuracy",
    value       = "0.7625",
    description = "Test set accuracy on Jan 2026 holdout dataset"
)

model_artifact.metadata_custom.add(
    key         = "training_date",
    value       = "2026-01-20",
    description = "Date this model was trained"
)

model_artifact.metadata_custom.add(
    key         = "data_source",
    value       = "oci://ml-bucket@namespace/churn/train_jan2026.csv",
    description = "OCI Object Storage path of training data"
)

model_artifact.metadata_custom.add(
    key         = "team",
    value       = "ML-Platform-Team",
    description = "Which team owns this model"
)

model_artifact.metadata_custom.add(
    key         = "psi_threshold",
    value       = "0.20",
    description = "PSI score that triggers automatic retraining of this model"
)

# Taxonomy metadata — standard OCI categories
model_artifact.metadata_taxonomy.set("UseCaseType",   "binary_classification")
model_artifact.metadata_taxonomy.set("Framework",     "scikit-learn")
model_artifact.metadata_taxonomy.set("Algorithm",     "RandomForestClassifier")
model_artifact.metadata_taxonomy.set("ArtifactTestResults", "PASSED")

print("✅ Metadata added! Now calling .save()...")

# NOW call .save()
model_id = model_artifact.save(
    display_name   = "Customer_Churn_RF_v1",
    description    = "Churn classifier with full metadata. Jan 2026.",
    project_id     = "ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID",
    compartment_id = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID"
)

print(f"\n✅ Model saved with full metadata!")
print(f"   OCID : {model_id}")
⭐ Pro Tip — Versioning Strategy:
Use a consistent naming pattern like ModelName_v{version}_{YYYY-MM}. For example: Churn_RF_v3_2026-01. This way, when you retrain monthly, you can instantly see the timeline of all model versions in OCI Console! 📅

🚀 Step 5: .deploy() — Go Live!

🧠 What is .deploy()?

This is the moment of truth! 🎉 .deploy() takes your saved model from the catalog and turns it into a live REST API endpoint running on OCI servers.

After this command completes, any application in the world can send HTTP requests to your endpoint URL and get predictions back in milliseconds! Your model is now a real product. 🌐

🏗️ What Happens Inside OCI When You Call .deploy()

1. Fetch model from catalog → OCI pulls your artifact from Model Catalog
2. Spin up VMs → Creates cloud servers with your chosen shape (CPU/GPU)
3. Install conda env → Sets up Python environment with all required libraries
4. Run load_model() → Calls your score.py to load model into memory
5. ✅ Endpoint is LIVE! → REST API URL is active and accepting prediction requests
📝 What the code below does:
This takes the model OCID from .save() and launches it as a live REST API on OCI's infrastructure. You configure the compute shape (CPU power), number of instances, and bandwidth. OCI handles all the infrastructure — no DevOps skills needed! 💪
from ads.model.deployment import ModelDeployment, ModelDeploymentProperties

# Configure the deployment
deployment_config = ModelDeploymentProperties(

    # Link to the saved model in Model Catalog
    model_id       = model_id,  # OCID from .save() step

    # OCI project and compartment
    project_id     = "ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID",
    compartment_id = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID",

    # Friendly name for this deployment (you can have multiple deployments per model!)
    display_name   = "churn-prediction-prod-v1",
    description    = "Production deployment of Customer Churn RF Model v1",

    # Compute settings — how powerful the serving VM is
    instance_shape = "VM.Standard.E4.Flex",
    instance_count = 2,       # Run 2 instances = high availability!
    ocpus          = 1,
    memory_in_gbs  = 16,

    # Network bandwidth for the endpoint
    bandwidth_mbps = 10,

    # Logging — send prediction logs to OCI Log Group for monitoring
    access_log = {
        "logGroupId": "ocid1.loggroup.oc1..YOUR_LOG_GROUP_OCID",
        "logId"     : "ocid1.log.oc1..YOUR_ACCESS_LOG_OCID"
    },
    predict_log = {
        "logGroupId": "ocid1.loggroup.oc1..YOUR_LOG_GROUP_OCID",
        "logId"     : "ocid1.log.oc1..YOUR_PREDICT_LOG_OCID"
    }
)

# 🚀 Deploy! (Takes approximately 5-15 minutes)
print("🚀 Deploying model to OCI... (this takes ~10 minutes)")
print("   Grab a coffee! ☕\n")

deployment = ModelDeployment()
deployment.deploy(
    properties          = deployment_config,
    wait_for_completion = True   # Script waits until deployment is 100% live
)

print("\n🎉 DEPLOYMENT SUCCESSFUL!")
print(f"   Endpoint URL : {deployment.url}")
print(f"   Status       : {deployment.state.name}")
print(f"   Instances    : 2 x VM.Standard.E4.Flex (1 OCPU, 16GB RAM each)")
print(f"\n   Your model is now live! Send prediction requests to:")
print(f"   POST {deployment.url}/predict")

Output:

🚀 Deploying model to OCI... (this takes ~10 minutes)
   Grab a coffee! ☕

🎉 DEPLOYMENT SUCCESSFUL!
   Endpoint URL : https://modeldeployment.us-ashburn-1.oci.customer-oci.com/ocid1.datasciencemodeldeployment.oc1..xxxx
   Status       : ACTIVE
   Instances    : 2 x VM.Standard.E4.Flex (1 OCPU, 16GB RAM each)

   Your model is now live! Send prediction requests to:
   POST https://modeldeployment.../predict

🔥 Calling Your Live Deployed Model

📝 What the code below does:
This sends a real HTTP request to your live deployed model endpoint — exactly like a mobile app or website would do it in production! It uses OCI's request signer to authenticate the call securely. 🔐
import requests
import oci
from oci.signer import Signer

# Set up OCI request signing (authentication for the API call)
config = oci.config.from_file("~/.oci/config")
signer = Signer(
    tenancy              = config["tenancy"],
    user                 = config["user"],
    fingerprint          = config["fingerprint"],
    private_key_file_location = config["key_file"]
)

# The customer data to predict
payload = {
    "data": [
        {
            "tenure"            : 3,
            "monthly_charges"   : 99.99,
            "num_support_calls" : 7,
            "contract_type"     : 0,
            "has_internet"      : 1
        }
    ]
}

# Send prediction request!
endpoint = f"{deployment.url}/predict"
response = requests.post(
    url  = endpoint,
    json = payload,
    auth = signer
)

result = response.json()
pred   = result["predictions"][0]

print(f"🤖 Live Prediction from Deployed Model:")
print(f"   Endpoint response time : {response.elapsed.total_seconds()*1000:.0f}ms")
print(f"   Prediction             : {pred['prediction_label']}")
print(f"   Churn Probability      : {pred['probability_churn']*100:.1f}%")
print(f"   Confidence             : {pred['confidence_pct']}")
print(f"\n   ✅ Your AI model is serving real predictions in production! 🎉")
⭐ Cost Warning! 💰
Deployed models keep running (and charging you money) 24/7 until you delete them! If you're just testing, remember to deactivate or delete the deployment when you're done. You can reactivate it anytime from the OCI Console. ⏸️

🔄 Updating a Deployment (Zero-Downtime!)

📝 What the code below does:
When you retrain your model and want to update the live deployment, this code swaps the old model for the new one — without any downtime! OCI does a rolling update: new instances come up, old ones go down gracefully. Your users never see an interruption. 🔄
# When you have a new model version (after retraining),
# update the deployment without taking it offline!

new_model_id = "ocid1.datasciencemodel.oc1..NEW_MODEL_OCID"

# Update the deployment to point to the new model
deployment.update(
    ModelDeploymentProperties(
        model_id      = new_model_id,
        display_name  = "churn-prediction-prod-v2",
        description   = "Updated to RF v2 — retrained with Feb 2026 data"
        # All other settings (shape, instances, etc.) remain the same
    )
)

print("✅ Deployment updated to new model version!")
print("   Zero downtime — users kept receiving predictions during the update!")

📊 Step 6: .summary_status() — The Health Dashboard

🧠 What is .summary_status()?

You wouldn't drive a car without checking the dashboard, right? 🚗 You'd want to know: Is the engine okay? How much fuel? Any warning lights?

.summary_status() is your model's dashboard! It gives you a complete, real-time status report on where your model is in its lifecycle — and whether each step succeeded or failed.

📝 What the code below does:
This displays a table showing the status of every step in your model's lifecycle. You can call it after any step — after .prepare(), after .save(), after .deploy() — to see exactly what has completed and what hasn't. Green ✅ = done. Red ❌ = failed. Yellow ⚪ = not started yet.
# Check the full status of your model's deployment journey
# Call this after any step to see where you are!

print("📊 Model Lifecycle Status Report\n")
model_artifact.summary_status()

# This generates an interactive table in your Jupyter notebook showing:
#
# ┌────────────────────┬──────────┬──────────────────────────────────────────────┐
# │ Step               │ Status   │ Details                                      │
# ├────────────────────┼──────────┼──────────────────────────────────────────────┤
# │ initiate           │ ✅ Done  │ Model artifact directory created              │
# │ prepare()          │ ✅ Done  │ score.py, model.pkl, runtime.yaml created     │
# │ verify()           │ ✅ Done  │ Local prediction test passed (3 samples)      │
# │ save()             │ ✅ Done  │ Model uploaded to OCI Model Catalog           │
# │ deploy()           │ ✅ Done  │ REST API endpoint is ACTIVE                   │
# └────────────────────┴──────────┴──────────────────────────────────────────────┘

📈 Checking Deployment Health Programmatically

📝 What the code below does:
Instead of just viewing a table, this code reads the status information programmatically — so you can use it in an automated monitoring script! If the deployment state is not ACTIVE, it automatically triggers an alert. 🚨
import datetime

def check_deployment_health(deployment):
    """
    A simple health check function you can run in a monitoring script.
    Checks if the deployment is active and prints a status report.
    """
    current_time = datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")

    print(f"🔍 Deployment Health Check — {current_time}")
    print("-" * 55)

    # Check the deployment state
    state = deployment.state.name  # 'ACTIVE', 'FAILED', 'CREATING', 'DELETING'

    print(f"   Deployment Name  : {deployment.display_name}")
    print(f"   Endpoint URL     : {deployment.url}")
    print(f"   Current State    : {state}")
    print(f"   Instances        : {deployment.instance_count}")
    print(f"   Shape            : {deployment.instance_shape}")

    # Interpret the state
    if state == "ACTIVE":
        print(f"\n   ✅ HEALTHY — Deployment is live and accepting requests!")

    elif state == "FAILED":
        print(f"\n   🔥 CRITICAL — Deployment has FAILED!")
        print(f"   ⚡ Action: Check OCI Log Explorer for error details.")
        print(f"   ⚡ Consider: Delete and redeploy with corrected score.py")
        # In a real script, you'd trigger a Slack/email alert here!

    elif state == "CREATING":
        print(f"\n   ⏳ PENDING — Deployment is still being created.")
        print(f"   ⚡ Wait 5-10 more minutes, then check again.")

    elif state == "INACTIVE":
        print(f"\n   💤 INACTIVE — Deployment exists but is paused.")
        print(f"   ⚡ Activate it from OCI Console or call deployment.activate()")

    else:
        print(f"\n   ⚠️  UNKNOWN STATE: {state}")
        print(f"   ⚡ Check OCI Console for more details.")

    print("-" * 55)
    return state == "ACTIVE"


# Run the health check
is_healthy = check_deployment_health(deployment)
print(f"\n   Ready for traffic? {'YES ✅' if is_healthy else 'NO ❌'}")

Output:

🔍 Deployment Health Check — 2026-01-20 14:32:07
-------------------------------------------------------
   Deployment Name  : churn-prediction-prod-v1
   Endpoint URL     : https://modeldeployment.us-ashburn-1.oci...
   Current State    : ACTIVE
   Instances        : 2
   Shape            : VM.Standard.E4.Flex
-------------------------------------------------------

   ✅ HEALTHY — Deployment is live and accepting requests!

   Ready for traffic? YES ✅

🔍 Additional Useful Status Methods

ADS provides several other helpful status and inspection methods that work alongside summary_status():

📝 What the code below does:
This shows a collection of extra methods that let you inspect different aspects of your model artifact — like reading a car's manual 📖 to understand every switch and button on the dashboard!
# ============================================================
# BONUS: Extra Inspection Methods Available After .prepare()
# ============================================================

# 1. List everything in the artifact folder
print("📂 Artifact folder contents:")
model_artifact.populate_metadata()   # Refresh metadata
print(model_artifact.artifact_dir)

# 2. Check what conda environment will be used at inference
print("\n🐍 Inference Conda Environment:")
print(model_artifact.metadata_taxonomy.to_dataframe())

# 3. View the auto-generated input schema
print("\n📥 Input Schema (what data the model expects):")
print(model_artifact.schema_input.to_dict())

# 4. View the auto-generated output schema
print("\n📤 Output Schema (what data the model returns):")
print(model_artifact.schema_output.to_dict())

# 5. View all custom metadata you attached
print("\n🏷️ Custom Metadata:")
for item in model_artifact.metadata_custom.to_dataframe().itertuples():
    print(f"   {item.key:<30 6.="" :="" after="" check="" code="" created="" deployment.state.name="" deployment.time_created="" deployment.url="" deployment="" directly="" f="" item.value="" n="" print="" state="" time="" url="">

🏁 Putting It All Together — The Complete 6-Step Flow

Now let's see all 6 steps in one clean, complete script that you can use as a template for ANY model deployment! 📋

📝 What the code below does:
This is your master deployment template! It chains all 6 steps together with proper error checking between each step. Copy this, replace the model and OCIDs, and you have a production-ready deployment script for any sklearn model! 🏆
"""
============================================================
MASTER OCI MODEL DEPLOYMENT SCRIPT
Covers all 6 steps: prepare → score.py → verify → save → deploy → summary_status
============================================================
"""

import ads
import os
from ads.model.framework.sklearn_model import SklearnModel
from ads.model.deployment import ModelDeployment, ModelDeploymentProperties

# ─── CONFIG ──────────────────────────────────────────────────
ARTIFACT_DIR   = "./churn_model_artifact_final"
COMPARTMENT_ID = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID"
PROJECT_ID     = "ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID"
MODEL_NAME     = "Customer_Churn_RF_v1"
CONDA_ENV      = "generalml_p38_cpu_v1"
# ─────────────────────────────────────────────────────────────

ads.set_auth("resource_principal")

print("=" * 55)
print("🚀 OCI Model Deployment — 6-Step Journey")
print("=" * 55)

# ─── STEP 1: PREPARE ─────────────────────────────────────────
print("\n📦 STEP 1: .prepare() — Packaging the model...")

model_artifact = SklearnModel(estimator=clf, artifact_dir=ARTIFACT_DIR)
model_artifact.prepare(
    inference_conda_env = CONDA_ENV,
    training_dataset    = X_train,
    force_overwrite     = True
)
print("   ✅ Artifact folder created with score.py, model.pkl, schemas\n")


# ─── STEP 2: INSPECT score.py ────────────────────────────────
print("📝 STEP 2: score.py — Verifying inference script exists...")

score_py_path = os.path.join(ARTIFACT_DIR, "score.py")
assert os.path.exists(score_py_path), "❌ score.py not found!"

with open(score_py_path, 'r') as f:
    content = f.read()
    assert "def load_model" in content, "❌ load_model() missing from score.py!"
    assert "def predict"    in content, "❌ predict() missing from score.py!"

print("   ✅ score.py exists with load_model() and predict() functions\n")


# ─── STEP 3: VERIFY ──────────────────────────────────────────
print("✅ STEP 3: .verify() — Running local pre-flight test...")

test_payload = {
    "data": [{"tenure": 6, "monthly_charges": 89.5,
               "num_support_calls": 4, "contract_type": 0, "has_internet": 1}]
}

verify_result = model_artifact.verify(test_payload)
assert "predictions" in verify_result, "❌ predict() returned unexpected output!"
print(f"   ✅ Local test passed! Sample prediction: "
      f"{verify_result['predictions'][0]['prediction_label']}\n")


# ─── STEP 4: SAVE ────────────────────────────────────────────
print("🏛️ STEP 4: .save() — Uploading to OCI Model Catalog...")

model_artifact.metadata_custom.add("accuracy",  "0.7625", "Test accuracy Jan 2026")
model_artifact.metadata_custom.add("team",      "ML-Platform", "Owning team")
model_artifact.metadata_taxonomy.set("UseCaseType", "binary_classification")
model_artifact.metadata_taxonomy.set("Framework",   "scikit-learn")

model_id = model_artifact.save(
    display_name   = MODEL_NAME,
    description    = "Churn RF v1. Accuracy 76.25%. Jan 2026 data.",
    project_id     = PROJECT_ID,
    compartment_id = COMPARTMENT_ID
)
print(f"   ✅ Model saved! OCID: {model_id}\n")


# ─── STEP 5: DEPLOY ──────────────────────────────────────────
print("🚀 STEP 5: .deploy() — Going live! (wait ~10 mins...)")

deployment = ModelDeployment()
deployment.deploy(
    properties = ModelDeploymentProperties(
        model_id       = model_id,
        project_id     = PROJECT_ID,
        compartment_id = COMPARTMENT_ID,
        display_name   = f"{MODEL_NAME}-prod",
        instance_shape = "VM.Standard.E4.Flex",
        instance_count = 2,
        ocpus          = 1,
        memory_in_gbs  = 16,
        bandwidth_mbps = 10
    ),
    wait_for_completion = True
)
print(f"   ✅ Deployment ACTIVE!")
print(f"   🌐 Endpoint: {deployment.url}\n")


# ─── STEP 6: SUMMARY STATUS ──────────────────────────────────
print("📊 STEP 6: .summary_status() — Final health check\n")
model_artifact.summary_status()

print("\n" + "=" * 55)
print("🎉 ALL 6 STEPS COMPLETE — MODEL IS LIVE IN PRODUCTION!")
print("=" * 55)
print(f"   Endpoint URL : {deployment.url}/predict")
print(f"   Method       : POST")
print(f"   Auth         : OCI Request Signing")
print(f"   Instances    : 2 (High Availability)")

📋 Complete Method Reference Table

Method Analogy What It Creates / Does Returns
.prepare() Pack the cake in a box 📦 Creates artifact folder with score.py, model.pkl, schemas, runtime.yaml Local folder path
score.py The waiter 🧑‍🍽️ Define load_model() and predict() — the inference logic you control Python file
.verify() Quality inspector ✅ Runs score.py locally, validates input/output, catches bugs before cloud upload Prediction dict
.save() Ship to the warehouse 🏛️ Uploads artifact to OCI Model Catalog with metadata and versioning Model OCID string
.deploy() Open the shop 🏪 Creates live REST API endpoint on OCI infrastructure with autoscaling ModelDeployment object
.summary_status() Dashboard check 📊 Shows completion status of every lifecycle step in a formatted table Interactive table

🏆 Best Practices — The Golden Rules

✅ DOs — What Experts Do Every Time:

  • ✅ Always run .verify() before .save() — no exceptions!
  • ✅ Add logging to every function in score.py (logger.info())
  • ✅ Use try/except in predict() to return clean error messages
  • ✅ Include rich metadata before .save() — accuracy, date, data source
  • ✅ Use consistent model naming: Name_Algorithm_vN_YYYY-MM
  • ✅ Enable prediction logging to OCI Log Groups for monitoring
  • ✅ Deploy with instance_count=2 for production (high availability)
  • ✅ Test your .verify() with edge cases (empty input, null values)
  • ✅ Call deployment.update() (not redeploy) when rolling out new models
  • ✅ Deactivate deployments when not in use to save costs! 💰
❌ DON'Ts — Costly Mistakes to Avoid:

  • ❌ Don't skip .verify() and go straight to .save() + .deploy()
  • ❌ Don't use hardcoded file paths in score.py
  • ❌ Don't return numpy dtypes from predict() — JSON can't serialize them!
  • ❌ Don't deploy without logging enabled — you'll be blind in production
  • ❌ Don't save models without meaningful display names or descriptions
  • ❌ Don't forget to delete or deactivate test deployments — they cost money!
  • ❌ Don't import large libraries at the TOP of score.py — it slows startup
  • ❌ Don't modify score.py after .verify() passes without re-verifying!

📝 Quick Summary — Everything in One Place

  • .prepare() = Creates the artifact folder with score.py, model file, and schemas
  • score.py = You write this! Defines load_model() and predict() functions
  • .verify() = Runs score.py locally — catches bugs before the cloud. NEVER SKIP!
  • .save() = Uploads the artifact to OCI Model Catalog. Returns the Model OCID.
  • .deploy() = Turns the saved model into a live REST API on OCI infrastructure
  • .summary_status() = Shows a health dashboard of all 6 lifecycle steps

Comments