Skip to main content

Deploy Your Fine-Tuned AI Model on OCI

Calculating read time…

Imagine you adopted a puppy 🐶 and spent months training it to do amazing tricks — sit, shake hands, fetch, roll over! Your puppy (the fine-tuned model) is brilliant.

But right now it only performs tricks in your living room (your laptop).  To show the world, you need to take it to a dog show stage (OCI deployment) where thousands of visitors can watch it perform anytime, on demand! 🏟️

That's exactly what this guide covers — taking your fine-tuned AI model from a private notebook all the way to a live, production-ready REST API on Oracle Cloud Infrastructure.





📚 Your 7-Step Journey

  • 🔐 Step 1: Authenticate — Show Your OCI Badge
  • 🔧 Step 2: Initialize Pipeline — Set Up the Assembly Line
  • 📦 Step 3: Prepare Model Artifact — Pack Everything Properly
  • 🔬 Step 4: Run Introspection — The Smart Health Checker
  • ✅ Step 5: Verify the Generated Model Artifacts — Final Pre-Flight Test
  • 🚀 Step 6: Deploy — Go Live on OCI!
  • 📡 Step 7: Invoke — Send Real Predictions
🗺️ The Complete Fine-Tuned Model Deployment Pipeline

🔐 Step 1
Authenticate
Prove who you are
→
🔧 Step 2
Init Pipeline
Configure the tool
→
📦 Step 3
Prepare Artifact
Build the package
→
🔬 Step 4
Introspection
Auto-validate files
→
✅ Step 5
Verify
Local test run
→
🚀 Step 6
Deploy
Go live on OCI
→
📡 Step 7
Invoke
Call the live API

⬆️ Every fine-tuned model deployment follows this exact sequence. Master it once — use it forever!


🧠 Quick Recap: What is a Fine-Tuned Model?

A foundation model (like Llama 3, Mistral, or Falcon) is a general-purpose AI — like a very smart university graduate. 🎓 It knows a little about everything.

Fine-tuning is extra specialized training — like that graduate completing a medical residency! 🏥 After fine-tuning on your company's data, the model becomes an expert specifically for your use case (customer support, legal docs, code review, etc.).

Deploying this fine-tuned model on OCI means your specialized AI expert becomes available as a live service that your apps can call 24/7. 🌐

⭐ Context:
The most common fine-tuned models being deployed on OCI are: Llama 3 variants (7B, 13B, 70B), Mistral 7B, Falcon, and custom healthcare/legal LLMs. OCI supports GPU shapes (A10, A100, H100) for serving these large models efficiently.

⚙️ Before You Start — Prerequisites

Make sure you have these ready before beginning:

Requirement What It Is Where to Get
OCI Account Your Oracle Cloud account cloud.oracle.com (Free Tier available!)
OCI Data Science Project A workspace container in OCI OCI Console → Data Science
Fine-tuned model weights Your trained model files (.bin, .safetensors) From HuggingFace fine-tuning or OCI AI Fine-Tuning
oracle-ads Python library The ADS toolkit (pre-installed in OCI notebooks) pip install oracle-ads[complete]
GPU-enabled notebook Notebook with GPU for LLM handling OCI Data Science → GPU shapes (A10, V100)

🔐 Step 1: Authenticate — Show Your OCI Badge

🧠 What is Authentication?

Imagine you walk into a secure government building. 🏛️ Before you can access ANY room, you must show your ID badge at reception. Without it — no access, no exceptions!

OCI authentication is exactly this — you must prove to OCI who you are before it lets you create projects, upload models, or deploy anything. There are 3 ways to show your badge, depending on where you're working.

🔑 3 Authentication Methods — Pick One Based on Where You Work

✅ Resource Principal
Inside OCI Notebook.
Most secure — no keys needed.
Use this in production!
⚪ API Key
On your local laptop.
Uses ~/.oci/config file.
Good for development.
⚪ Instance Principal
Inside OCI Cloud Shell
or OCI Compute VM.
Good for automation.
📝 What the code below does:
This is the very first thing you run in your notebook. It tells the ADS library how to connect to your OCI account — like typing your PIN code at a bank ATM! 🏦 We set up authentication AND verify we can reach the right OCI region and compartment.
import ads
import oci
import os
import json

# ─── STEP 1A: Set Authentication ────────────────────────────────
# Inside OCI Data Science Notebook — use resource_principal (BEST!)
ads.set_auth("resource_principal")

# On your LOCAL machine — use API key from ~/.oci/config
# ads.set_auth("api_key")

print("✅ Authentication configured!")

# ─── STEP 1B: Set Your OCI Configuration Constants ─────────────
# Replace ALL of these with YOUR actual OCI values!
# Find them in OCI Console → Identity → Compartments

COMPARTMENT_ID = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID"
PROJECT_ID     = "ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID"
REGION         = "us-ashburn-1"   # e.g., us-ashburn-1, eu-frankfurt-1, ap-mumbai-1

# The base model that was fine-tuned (used for tokenizer loading)
BASE_MODEL_NAME = "meta-llama/Llama-3-8b-hf"

# Where your fine-tuned model weights are stored in OCI Object Storage
FINETUNED_MODEL_PATH = "oci://your-bucket@your-namespace/finetuned-models/llama3-customer-support/"

# Local folder for building the deployment artifact
ARTIFACT_DIR = "./llama3_finetuned_artifact"

# Which GPU conda environment to use for inference
INFERENCE_CONDA_ENV = "pytorch21_p39_gpu_v1"   # OCI pre-built PyTorch + GPU env

print(f"✅ Configuration loaded!")
print(f"   Region        : {REGION}")
print(f"   Compartment   : {COMPARTMENT_ID[:30]}...")
print(f"   Base Model    : {BASE_MODEL_NAME}")
print(f"   Model Path    : {FINETUNED_MODEL_PATH}")

🔍 Verify Your Authentication Works

📝 What the code below does:
This makes a simple test call to OCI to confirm authentication is working. Think of it as dialling a test call after setting up a new phone — just to make sure it actually connects! 📞 If it returns your compartment details, you're good to go!
import oci

# Test authentication by listing your OCI Data Science projects
config  = oci.config.from_file()   # Reads ~/.oci/config
signer  = oci.auth.signers.get_resource_principals_signer()

# Try to connect to OCI Data Science service
ds_client = oci.data_science.DataScienceClient(config={}, signer=signer)

try:
    projects = ds_client.list_projects(compartment_id=COMPARTMENT_ID)
    print(f"✅ Authentication CONFIRMED! OCI connection is working.")
    print(f"   Found {len(projects.data)} project(s) in your compartment:")
    for proj in projects.data:
        print(f"   → {proj.display_name} ({proj.lifecycle_state})")
except Exception as e:
    print(f"❌ Authentication FAILED: {e}")
    print(f"   Check: Is resource_principal enabled for this notebook?")
    print(f"   Check: Does your policy allow 'manage data-science-models'?")
❌ Most Common Auth Error — "Authorization failed":
This means your OCI IAM policies don't grant you access to Data Science. Ask your OCI admin to add this policy:
Allow dynamic-group my-notebook-dg to manage data-science-models in compartment my-compartment
Policies are set in OCI Console → Identity → Policies. 🔑

🔧 Step 2: Initialize Pipeline — Set Up the Assembly Line

🧠 What Does "Initialize Pipeline" Mean?

Imagine you're setting up a car assembly line in a factory. 🏭 Before building any car, you need to: choose which car model to build, configure all the machines, and decide what materials to use.

In OCI, "Initialize Pipeline" means setting up the HuggingFace Pipeline that will serve your fine-tuned model. This pipeline knows how to load your model, tokenize input text, run inference, and return results — all in one clean chain!

⚙️ What a HuggingFace Inference Pipeline Contains

🔤 Tokenizer
Converts text → numbers
→
🤖 Model Weights
Your fine-tuned brain
→
⚡ GPU Device
Runs computations fast
→
📤 Decoded Output
Text response
📝 What the code below does:
This loads your fine-tuned model weights and tokenizer from OCI Object Storage into a HuggingFace inference pipeline. Think of it as hiring the chef (loading the model) and giving them the recipe book (tokenizer) before the restaurant opens! 👨‍🍳📖

We also apply memory optimizations so large models like Llama 3 can fit on GPU efficiently.
import torch
from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
    pipeline,
    BitsAndBytesConfig
)
import os

print("🔧 STEP 2: Initializing Inference Pipeline...\n")

# ─── 2A: Download model from OCI Object Storage ────────────────
# The fine-tuned weights live in OCI Object Storage.
# We download them to a local temp folder first.

LOCAL_MODEL_DIR = "./local_finetuned_model"
os.makedirs(LOCAL_MODEL_DIR, exist_ok=True)

print(f"📥 Downloading fine-tuned model from OCI Object Storage...")
print(f"   Source : {FINETUNED_MODEL_PATH}")

# Use OCI SDK to download the model files
import oci

signer      = oci.auth.signers.get_resource_principals_signer()
os_client   = oci.object_storage.ObjectStorageClient(config={}, signer=signer)

# Parse bucket info from the OCI path
# Format: oci://bucket@namespace/prefix/
bucket_name = "your-bucket"
namespace   = "your-namespace"
prefix      = "finetuned-models/llama3-customer-support/"

# List and download all model files
objects = os_client.list_objects(namespace, bucket_name, prefix=prefix)
for obj in objects.data.objects:
    local_file = os.path.join(LOCAL_MODEL_DIR, os.path.basename(obj.name))
    if not obj.name.endswith("/"):   # Skip directory entries
        print(f"   → Downloading: {os.path.basename(obj.name)}")
        response = os_client.get_object(namespace, bucket_name, obj.name)
        with open(local_file, 'wb') as f:
            f.write(response.data.content)

print(f"✅ Model downloaded to: {LOCAL_MODEL_DIR}\n")

# ─── 2B: Configure 4-bit Quantization for Memory Efficiency ────
# Large models like Llama 3 need LOTS of GPU memory.
# 4-bit quantization is like compressing a 4K video to 720p —
# you lose a tiny bit of quality but use 4x less GPU memory!
# This is ESSENTIAL for deploying 7B+ parameter models on OCI.

quantization_config = BitsAndBytesConfig(
    load_in_4bit              = True,
    bnb_4bit_compute_dtype    = torch.float16,
    bnb_4bit_use_double_quant = True,   # Extra compression layer
    bnb_4bit_quant_type       = "nf4"   # NormalFloat4 — best quality/size tradeoff
)

print("⚙️  Loading tokenizer and model weights...")

# ─── 2C: Load Tokenizer ────────────────────────────────────────
# The tokenizer converts human text ↔ numbers the model understands.
# Even though we fine-tuned the model, we usually reuse the BASE model's tokenizer.

tokenizer = AutoTokenizer.from_pretrained(
    LOCAL_MODEL_DIR,
    trust_remote_code = True   # Required for some model architectures
)

# Set padding token (needed for batch inference)
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token

print(f"   ✅ Tokenizer loaded. Vocabulary size: {tokenizer.vocab_size:,} tokens")

# ─── 2D: Load Fine-Tuned Model ────────────────────────────────
# This loads the actual brain — billions of parameters!
# Using quantization_config to keep GPU memory usage manageable.

model = AutoModelForCausalLM.from_pretrained(
    LOCAL_MODEL_DIR,
    quantization_config = quantization_config,
    device_map          = "auto",       # Automatically distributes across available GPUs
    trust_remote_code   = True,
    torch_dtype         = torch.float16
)

model.eval()  # Set to evaluation mode (no gradient tracking — saves more memory!)
print(f"   ✅ Model loaded. Device: {next(model.parameters()).device}")

# ─── 2E: Create the Inference Pipeline ────────────────────────
# This chains tokenizer + model into one easy-to-use object!
# You send text → it returns text. Simple!

inference_pipeline = pipeline(
    task      = "text-generation",
    model     = model,
    tokenizer = tokenizer,
    device_map = "auto",

    # Generation parameters — control HOW the model responds
    max_new_tokens  = 512,       # Max words in the response
    temperature     = 0.7,       # 0 = deterministic, 1 = creative
    do_sample       = True,      # Enable sampling (makes responses varied)
    top_p           = 0.9,       # Nucleus sampling — cuts off unlikely words
    repetition_penalty = 1.15,   # Penalizes repeated phrases
    return_full_text   = False   # Only return new text, not the input prompt
)

print("\n✅ STEP 2 COMPLETE — Inference pipeline is ready!")
print(f"   Model  : Fine-tuned Llama 3 8B")
print(f"   Device : GPU (4-bit quantized)")
print(f"   Task   : Text Generation")

# Quick sanity test — make sure the pipeline works!
test_output = inference_pipeline("Hello! Who are you?", max_new_tokens=30)
print(f"\n🧪 Quick sanity test:")
print(f"   Input  : Hello! Who are you?")
print(f"   Output : {test_output[0]['generated_text'][:100]}...")
✅ GPU Memory Tips :
  • Llama 3 8B (4-bit) → needs ~6 GB GPU VRAM → OCI A10 shape works perfectly
  • Llama 3 70B (4-bit) → needs ~40 GB GPU VRAM → Use OCI BM.GPU.A100 (80GB)
  • Mistral 7B (4-bit) → needs ~5 GB GPU VRAM → OCI A10 shape works
  • Always test quantization locally before deploying to catch compatibility issues!

📦 Step 3: Prepare Model Artifact — Build the Deployment Package

🧠 What is a Model Artifact?

Your HuggingFace pipeline runs beautifully in your notebook. But OCI can't just take a running notebook and deploy it! 🚫

OCI needs a standardized package — called an artifact — that contains everything needed to run your model from scratch. Think of it like writing a complete IKEA instruction manual 📋 so anyone (or any OCI server) can rebuild your model from scratch!

🗂️ What the Artifact Folder Must Contain

📁 Required Artifact Folder Structure for LLM Deployment

📁 llama3_finetuned_artifact/
    ├── 🐍 score.py               ← YOUR inference logic (load + predict)
    ├── 📋 runtime.yaml           ← Python env + dependencies
    ├── 📄 input_schema.json      ← What inputs the model accepts
    ├── 📄 output_schema.json     ← What the model returns
    ├── ⚙️ config.json            ← Model architecture config
    └── 🔗 model_weights_location.txt ← Points to OCI Object Storage (for large models!)

⚠️ Note: For LLMs (multi-GB files), we DON'T put model weights in the artifact! We put a pointer file instead. Weights stay in OCI Object Storage. This avoids hitting the 6GB artifact size limit!

📝 What the code below does:
This uses ADS's HuggingFacePipelineModel to automatically generate most of the artifact files for you — and then we write a custom score.py that tells OCI exactly how to load our fine-tuned model and handle inference. Think of it as ADS doing 80% of the paperwork and you filling in the details! 📝
from ads.model.framework.hugging_face_model import HuggingFacePipelineModel
import os, json

print("📦 STEP 3: Preparing Model Artifact...\n")

# ─── 3A: Create ADS model wrapper ──────────────────────────────
# ADS wraps your HuggingFace pipeline into an OCI-compatible model object.
# It auto-generates schemas and runtime config!

model_artifact = HuggingFacePipelineModel(
    estimator    = inference_pipeline,
    artifact_dir = ARTIFACT_DIR
)

# ─── 3B: Run .prepare() ────────────────────────────────────────
# This creates the artifact folder structure automatically.
model_artifact.prepare(
    inference_conda_env = INFERENCE_CONDA_ENV,
    force_overwrite     = True,

    # Use case for the auto-generated schema hints
    # "text_generation" or "text_classification" etc.
    use_case_type = "text_generation"
)

print(f"✅ Base artifact structure created at: {ARTIFACT_DIR}\n")

# ─── 3C: Write Custom score.py for Fine-Tuned LLM ─────────────
# ADS generates a basic score.py, but for fine-tuned LLMs
# we need a CUSTOM one that:
# (1) Loads model weights from OCI Object Storage (not artifact folder)
# (2) Applies 4-bit quantization
# (3) Handles the specific prompt format our model was fine-tuned on
# (4) Returns clean, well-structured responses

score_py_content = '''
# =============================================================
# score.py — Fine-Tuned Llama 3 Customer Support Model
# OCI calls load_model() on startup, predict() for each request.
# =============================================================

import json
import os
import logging
import torch
from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
    pipeline,
    BitsAndBytesConfig
)
import oci

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

# OCI Object Storage config — where model weights live
NAMESPACE   = "your-namespace"
BUCKET      = "your-bucket"
MODEL_KEY   = "finetuned-models/llama3-customer-support/"
LOCAL_DIR   = "/tmp/model_weights"

# The prompt template this model was fine-tuned on (MUST match training!)
PROMPT_TEMPLATE = """<|system|>
You are a helpful customer support agent for TechCorp.
Answer questions about our products clearly and professionally.
<|user|> {user_message} <|assistant|> """ def _download_model_from_oci(): """Downloads model weights from OCI Object Storage to /tmp.""" logger.info("Downloading model weights from OCI Object Storage...") os.makedirs(LOCAL_DIR, exist_ok=True) signer = oci.auth.signers.get_resource_principals_signer() os_client = oci.object_storage.ObjectStorageClient(config={}, signer=signer) objects = os_client.list_objects(NAMESPACE, BUCKET, prefix=MODEL_KEY) for obj in objects.data.objects: if obj.name.endswith("/"): continue file_name = os.path.basename(obj.name) local_path = os.path.join(LOCAL_DIR, file_name) if not os.path.exists(local_path): # Skip if already downloaded logger.info(f" Downloading: {file_name}") response = os_client.get_object(NAMESPACE, BUCKET, obj.name) with open(local_path, "wb") as f: f.write(response.data.content) logger.info("✅ All model files downloaded!") def load_model(): """ Called ONCE when OCI starts the deployment server. Downloads weights and loads the inference pipeline into memory. """ logger.info("=== load_model() called — initializing model ===") # Download model from OCI Object Storage _download_model_from_oci() # Load tokenizer tokenizer = AutoTokenizer.from_pretrained(LOCAL_DIR, trust_remote_code=True) if tokenizer.pad_token is None: tokenizer.pad_token = tokenizer.eos_token # Load model with 4-bit quantization quant_config = BitsAndBytesConfig( load_in_4bit = True, bnb_4bit_compute_dtype = torch.float16, bnb_4bit_quant_type = "nf4" ) model = AutoModelForCausalLM.from_pretrained( LOCAL_DIR, quantization_config = quant_config, device_map = "auto", trust_remote_code = True, torch_dtype = torch.float16 ) model.eval() # Build the inference pipeline pipe = pipeline( task = "text-generation", model = model, tokenizer = tokenizer, device_map = "auto", max_new_tokens = 512, temperature = 0.7, do_sample = True, top_p = 0.9, repetition_penalty = 1.15, return_full_text = False ) logger.info("✅ Model loaded and ready for inference!") return pipe def predict(data, model=load_model()): """ Called for EVERY prediction request. data = JSON payload from the API caller model = the loaded pipeline from load_model() """ logger.info(f"predict() called with data type: {type(data)}") try: # Parse the incoming request if isinstance(data, dict): user_message = data.get("input", data.get("prompt", "")) system_override = data.get("system_prompt", None) max_tokens = int(data.get("max_tokens", 512)) elif isinstance(data, str): user_message = data system_override = None max_tokens = 512 else: return {"error": "Unsupported input format. Send JSON with 'input' key."} if not user_message.strip(): return {"error": "Empty input received. Please provide a question or prompt."} # Build the prompt using the fine-tuning template if system_override: prompt = f"<|system|>\\n{system_override}\\n\\n<|user|>\\n{user_message}\\n\\n<|assistant|>\\n" else: prompt = PROMPT_TEMPLATE.format(user_message=user_message) logger.info(f"Running inference. Input length: {len(user_message)} chars") # Run inference! output = model( prompt, max_new_tokens = min(max_tokens, 1024) # Cap at 1024 for safety ) generated_text = output[0]["generated_text"].strip() # Return a clean, structured response return { "response" : generated_text, "input_length" : len(user_message), "output_length" : len(generated_text), "model_version" : "llama3-8b-customer-support-v1", "status" : "success" } except torch.cuda.OutOfMemoryError: logger.error("GPU Out of Memory!") torch.cuda.empty_cache() return {"error": "GPU memory exhausted. Try a shorter input or lower max_tokens."} except Exception as e: logger.error(f"Prediction error: {str(e)}") return {"error": f"Prediction failed: {str(e)}", "status": "error"} ''' # Write the custom score.py to the artifact folder score_py_path = os.path.join(ARTIFACT_DIR, "score.py") with open(score_py_path, "w") as f: f.write(score_py_content) print(f"✅ Custom score.py written to: {score_py_path}") # ─── 3D: Write model_weights_location.txt ───────────────────── # This is a pointer file telling anyone where the actual weights live. # Helps with traceability — which version of the model is this artifact using? weights_info = { "storage_type" : "OCI Object Storage", "bucket" : "your-bucket", "namespace" : "your-namespace", "prefix" : "finetuned-models/llama3-customer-support/", "base_model" : "meta-llama/Llama-3-8b-hf", "fine_tune_date": "2026-01-20", "quantization" : "4-bit NF4 (bitsandbytes)" } with open(os.path.join(ARTIFACT_DIR, "model_weights_location.json"), "w") as f: json.dump(weights_info, f, indent=2) # ─── 3E: List what was created ──────────────────────────────── print(f"\n📂 Artifact folder contents:") for fname in sorted(os.listdir(ARTIFACT_DIR)): fsize = os.path.getsize(os.path.join(ARTIFACT_DIR, fname)) print(f" 📄 {fname:<35 3="" artifact="" bytes="" code="" complete="" f="" fsize:="" is="" n="" package="" print="" ready="" step="">
⭐ Critical LLM Deployment Rule :
OCI Model Artifact has a 6GB size limit. LLM model weights are 5-140GB! 😱 NEVER put weights in the artifact folder. Always store weights in OCI Object Storage and download them in load_model(). Your artifact should only contain: score.py, runtime.yaml, schemas, and a pointer file.

🔬 Step 4: Run Introspection — The Smart Health Checker

🧠 What is Introspection?

Imagine a pilot doing a pre-flight safety checklist before takeoff. ✈️ They check: fuel levels ✅, engine status ✅, landing gear ✅, navigation system ✅... Every single thing. No exceptions.

Introspection is ADS's automatic pre-flight checklist for your model! It examines every file in your artifact folder and verifies that: (1) All required files are present, (2) score.py has the right functions, (3) runtime.yaml specifies a valid environment, (4) schemas are well-formed.

✈️ What Introspection Checks — The Full Checklist

🔍 File Presence → score.py, runtime.yaml, input/output schemas exist?
🐍 score.py Structure → load_model() and predict() defined?
📋 runtime.yaml → Valid conda environment specified?
📥 Input Schema → Valid JSON schema format?
📤 Output Schema → Valid JSON schema format?
🔗 Import Check → All imports in score.py are valid Python?
🏷️ Metadata → Required model metadata present?
📝 What the code below does:
This runs the built-in ADS introspection test on your artifact folder. It's like a spell-checker for your deployment package — it reads every file and reports any problems it finds. Each check either passes ✅ or fails ❌ with a clear reason!
print("🔬 STEP 4: Running Introspection...\n")

# Run the introspection test
# This scans every file in the artifact and validates structure
introspection_result = model_artifact.introspect()

# Display the results
print("📊 Introspection Report:")
print("-" * 60)

all_passed = True
for check_name, result in introspection_result.items():
    status = result.get("status", "unknown")
    message = result.get("message", "")

    if status == "passed":
        icon = "✅"
    elif status == "warning":
        icon = "⚠️ "
        all_passed = False
    else:
        icon = "❌"
        all_passed = False

    print(f"   {icon} {check_name:<35 60="" :="" above="" all="" all_passed:="" and="" are="" artifact="" attention.="" before="" check="" checks="" code="" common="" create="" defined="" else:="" ensure="" f="" file:="" files="" fix="" fixes:="" folder="" for="" format="" if="" in="" introspection="" is="" issue:="" issues="" it="" json="" load_model="" message="" missing="" n="" need="" passed="" predict="" print="" proceeding.="" ready="" schema="" score.py="" some="" status.upper="" status="" that="" the="" valid="" verification.="">

🛠️ Fixing Common Introspection Failures

📝 What the code below does:
This is a helper script that manually creates any missing required files in case introspection found something missing. Think of it as going back to pack the things you forgot! 🎒
import json, yaml, os

def fix_missing_artifacts(artifact_dir):
    """
    Checks for and creates any missing required artifact files.
    Run this if introspection fails with 'file not found' errors.
    """

    # ─── Fix 1: Create input_schema.json if missing ────────────
    input_schema_path = os.path.join(artifact_dir, "input_schema.json")
    if not os.path.exists(input_schema_path):
        print("   Creating input_schema.json...")
        input_schema = {
            "$schema"    : "http://json-schema.org/draft-07/schema#",
            "type"       : "object",
            "properties" : {
                "input" : {
                    "type"        : "string",
                    "description" : "The user's question or prompt text",
                    "example"     : "How do I reset my account password?"
                },
                "max_tokens" : {
                    "type"        : "integer",
                    "description" : "Maximum tokens in response (default: 512)",
                    "default"     : 512
                },
                "temperature" : {
                    "type"        : "number",
                    "description" : "Response creativity (0.0-1.0)",
                    "default"     : 0.7
                }
            },
            "required": ["input"]
        }
        with open(input_schema_path, "w") as f:
            json.dump(input_schema, f, indent=2)
        print("   ✅ input_schema.json created")

    # ─── Fix 2: Create output_schema.json if missing ───────────
    output_schema_path = os.path.join(artifact_dir, "output_schema.json")
    if not os.path.exists(output_schema_path):
        print("   Creating output_schema.json...")
        output_schema = {
            "$schema"    : "http://json-schema.org/draft-07/schema#",
            "type"       : "object",
            "properties" : {
                "response"      : {"type": "string", "description": "The model's generated response"},
                "input_length"  : {"type": "integer", "description": "Character count of input"},
                "output_length" : {"type": "integer", "description": "Character count of response"},
                "model_version" : {"type": "string", "description": "Model version identifier"},
                "status"        : {"type": "string", "enum": ["success", "error"]}
            }
        }
        with open(output_schema_path, "w") as f:
            json.dump(output_schema, f, indent=2)
        print("   ✅ output_schema.json created")

    # ─── Fix 3: Create/Update runtime.yaml ─────────────────────
    runtime_path = os.path.join(artifact_dir, "runtime.yaml")
    runtime_config = {
        "kind"    : "conda_environment",
        "version" : "1.0",
        "inference": {
            "python_version"   : "3.9",
            "conda_env"        : INFERENCE_CONDA_ENV,
            "type"             : "published",   # published = OCI-managed env
            "input_schema_uri" : "predict_input_schema.json",
            "output_schema_uri": "predict_output_schema.json"
        },
        "MODEL_DEPLOYMENT" : {
            "INFERENCE_CONDA_ENV" : INFERENCE_CONDA_ENV
        }
    }
    with open(runtime_path, "w") as f:
        yaml.dump(runtime_config, f, default_flow_style=False)
    print("   ✅ runtime.yaml updated")

    print("\n✅ All missing artifacts fixed! Run introspection again to confirm.")


# Run the fixer if needed
print("🔧 Running artifact fixer...")
fix_missing_artifacts(ARTIFACT_DIR)

# Re-run introspection to confirm everything passes
print("\n🔬 Re-running introspection after fixes...")
introspection_result = model_artifact.introspect()
all_clean = all(r.get("status") == "passed" for r in introspection_result.values())
print(f"\n{'✅ All checks passed!' if all_clean else '⚠️  Still some issues — check errors above'}")
✅ DO: Always run introspection AND fix all warnings — not just critical failures. Warnings often become errors when deployed to OCI's production environment. A warning on your laptop = a crash in production! ⚠️

✅ Step 5: Verify the Generated Model Artifacts

🧠 Why Verify After Introspection?

Introspection checks the structure of your files. Verification actually runs your code! 🏃

Think of introspection as checking your car's lights and mirrors. Verification is actually starting the engine and doing a test drive around the parking lot before hitting the highway! 🚗

📝 What the code below does:
This calls .verify() with sample inputs that simulate what a real API caller would send. ADS loads your score.py, calls load_model(), then predict(), and checks that everything works end-to-end — all locally, before touching OCI. If this passes, your deployment will almost certainly succeed! 🎯
print("✅ STEP 5: Verifying Model Artifacts...\n")

# Test Case 1: Standard customer support question
test_case_1 = {
    "input"      : "How do I reset my password? I've been locked out of my account.",
    "max_tokens" : 200,
    "temperature": 0.7
}

# Test Case 2: Short input
test_case_2 = {
    "input": "What are your business hours?"
}

# Test Case 3: Edge case — system prompt override
test_case_3 = {
    "input"         : "What is the refund policy?",
    "system_prompt" : "You are a concise assistant. Answer in under 50 words.",
    "max_tokens"    : 100
}

# Test Case 4: Edge case — empty input (should return a clean error!)
test_case_4 = {
    "input": ""
}

test_cases = [
    ("Normal customer question", test_case_1),
    ("Short question"          , test_case_2),
    ("Custom system prompt"    , test_case_3),
    ("Empty input (error case)", test_case_4)
]

print("🧪 Running verification tests:\n")

all_tests_passed = True
for test_name, test_data in test_cases:
    print(f"   Test: {test_name}")
    print(f"   Input: {str(test_data)[:60]}...")

    try:
        result = model_artifact.verify(test_data)

        # Check basic response structure
        if "error" in result and test_name != "Empty input (error case)":
            print(f"   ❌ FAILED — Got error: {result['error']}\n")
            all_tests_passed = False
        elif "response" in result:
            print(f"   ✅ PASSED — Response preview: {result['response'][:80]}...")
            print(f"   Status: {result.get('status', 'unknown')}, "
                  f"Length: {result.get('output_length', '?')} chars\n")
        elif "error" in result and test_name == "Empty input (error case)":
            print(f"   ✅ PASSED — Error handled cleanly: {result['error']}\n")
        else:
            print(f"   ⚠️  Unexpected output structure: {list(result.keys())}\n")
            all_tests_passed = False

    except Exception as e:
        print(f"   ❌ EXCEPTION — {str(e)}\n")
        all_tests_passed = False

print("-" * 55)
if all_tests_passed:
    print("🎉 ALL VERIFICATION TESTS PASSED!")
    print("   The model artifact is ready for OCI deployment!")
else:
    print("❌ Some tests failed. Fix score.py before proceeding to .save() and .deploy()")

Sample Output:

🧪 Running verification tests:

   Test: Normal customer question
   Input: {'input': 'How do I reset my password?...'}...
   ✅ PASSED — Response preview: To reset your password, please visit our login page
                                 and click "Forgot Password"...
   Status: success, Length: 187 chars

   Test: Short question
   Input: {'input': 'What are your business hours?'}...
   ✅ PASSED — Response preview: Our support team is available Monday through Friday...
   Status: success, Length: 124 chars

   Test: Empty input (error case)
   Input: {'input': ''}...
   ✅ PASSED — Error handled cleanly: Empty input received.

-------------------------------------------------------
🎉 ALL VERIFICATION TESTS PASSED!
   The model artifact is ready for OCI deployment!
❌ DON'T skip verification tests!
Many beginners go straight from .prepare() to .save() to .deploy() and only discover their score.py is broken after spending 15 minutes waiting for deployment — only to get a FAILED status! Verification catches 95% of issues in under 2 minutes locally. ⏱️

🚀 Step 6: Deploy — Go Live on OCI!

🧠 What Happens During Deployment?

Deployment is when OCI takes your verified artifact, spins up real GPU servers, installs your Python environment, loads your model, and makes it available as a live HTTPS endpoint that anyone can call!

⚙️ What OCI Does Behind the Scenes During .deploy()

1. 📥 Fetches your artifact from OCI Model Catalog
2. 🖥️ Provisions GPU VM(s) with your chosen shape (A10, A100, etc.)
3. 🐍 Downloads and activates your conda environment
4. 📦 Extracts artifact files to the server's filesystem
5. 🔄 Calls load_model() from your score.py
6. 🌐 Starts an HTTPS server that routes requests to predict()
7. ✅ Marks the deployment as ACTIVE — live and serving!

First, we need to .save() to OCI Model Catalog, then call .deploy(). Let's do both:

📝 What the code below does:
This uploads your verified artifact to OCI Model Catalog (.save()) and then launches it as a live GPU-powered REST API (.deploy()). After this runs, your fine-tuned AI is live on the internet! 🌐

Important: We choose a GPU shape because LLMs absolutely need GPU power to respond in reasonable time. On CPU, a 7B model takes 2-5 minutes per response. On GPU, it's under 3 seconds! ⚡
from ads.model.deployment import ModelDeployment, ModelDeploymentProperties

print("🏛️ Saving model to OCI Model Catalog...")

# ─── 6A: Add Metadata Before Saving ───────────────────────────
model_artifact.metadata_custom.add(
    key         = "base_model",
    value       = BASE_MODEL_NAME,
    description = "Foundation model this was fine-tuned from"
)
model_artifact.metadata_custom.add(
    key         = "fine_tune_task",
    value       = "customer-support-qa",
    description = "The specific task this model was fine-tuned for"
)
model_artifact.metadata_custom.add(
    key         = "quantization",
    value       = "4-bit NF4 bitsandbytes",
    description = "Quantization applied to reduce GPU memory footprint"
)
model_artifact.metadata_custom.add(
    key         = "training_date",
    value       = "2026-01-15",
    description = "Date fine-tuning was completed"
)
model_artifact.metadata_custom.add(
    key         = "prompt_template",
    value       = "Zephyr-style: <|system|>...<|user|>...<|assistant|>",
    description = "Prompt format this model was trained with"
)

model_artifact.metadata_taxonomy.set("UseCaseType", "text_generation")
model_artifact.metadata_taxonomy.set("Framework",   "PyTorch")
model_artifact.metadata_taxonomy.set("Algorithm",   "LLaMA-3-8B-fine-tuned")

# ─── 6B: Save to Model Catalog ────────────────────────────────
model_id = model_artifact.save(
    display_name   = "Llama3-8B-CustomerSupport-v1",
    description    = (
        "Fine-tuned Llama 3 8B for TechCorp customer support. "
        "Trained on 50K Q&A pairs. 4-bit quantized. Jan 2026."
    ),
    project_id     = PROJECT_ID,
    compartment_id = COMPARTMENT_ID
)

print(f"✅ Model saved! OCID: {model_id}\n")

# ─── 6C: Deploy the Model ─────────────────────────────────────
print("🚀 Deploying to OCI (this takes 10-20 minutes for LLMs)...")
print("   Provisioning GPU VM, installing conda env, loading weights...")
print("   ☕ Perfect time for a coffee break!\n")

deployment = ModelDeployment()
deployment.deploy(
    properties = ModelDeploymentProperties(
        model_id       = model_id,
        project_id     = PROJECT_ID,
        compartment_id = COMPARTMENT_ID,

        display_name   = "llama3-customer-support-prod-v1",
        description    = "Production LLM API for customer support chatbot",

        # ── GPU Shape Selection ──────────────────────────────
        # For Llama 3 8B (4-bit): A10 (24GB VRAM) works great!
        # For Llama 3 70B (4-bit): Need A100 (80GB VRAM)
        instance_shape = "VM.GPU.A10.1",    # 1x NVIDIA A10 GPU, 24GB VRAM
        instance_count = 1,                  # 1 for dev/test, 2+ for production HA

        bandwidth_mbps = 10,

        # ── Logging Configuration ────────────────────────────
        # Essential for debugging production issues!
        access_log = {
            "logGroupId": "ocid1.loggroup.oc1..YOUR_LOG_GROUP_OCID",
            "logId"     : "ocid1.log.oc1..YOUR_ACCESS_LOG_OCID"
        },
        predict_log = {
            "logGroupId": "ocid1.loggroup.oc1..YOUR_LOG_GROUP_OCID",
            "logId"     : "ocid1.log.oc1..YOUR_PREDICT_LOG_OCID"
        },

        # ── Environment Variables ────────────────────────────
        # Passed into score.py's runtime environment
        environment_variables = {
            "MODEL_VERSION"      : "llama3-8b-v1",
            "MAX_BATCH_SIZE"     : "1",        # LLMs typically serve 1 request at a time
            "TRANSFORMERS_CACHE" : "/tmp/hf_cache",
            "TOKENIZERS_PARALLELISM": "false"  # Prevents HuggingFace tokenizer warnings
        }
    ),

    wait_for_completion = True   # Block until deployment is ACTIVE
)

print("\n🎉 DEPLOYMENT SUCCESSFUL!")
print(f"   Status       : {deployment.state.name}")
print(f"   Endpoint URL : {deployment.url}")
print(f"   GPU Shape    : VM.GPU.A10.1 (NVIDIA A10, 24GB VRAM)")
print(f"\n   📡 Your fine-tuned AI is now LIVE on OCI!")
print(f"   Send POST requests to: {deployment.url}/predict")
⭐ GPU Shape Guide for LLMs:
  • VM.GPU.A10.1 → 24GB VRAM → Good for 7B/8B models (4-bit)
  • BM.GPU.A10.4 → 96GB VRAM → Good for 13B-34B models (4-bit)
  • BM.GPU.A100-v2.8 → 640GB VRAM → For 70B models and multi-GPU inference
  • BM.GPU.H100.8 → 640GB VRAM → For fastest inference, production at scale

📡 Step 7: Invoke — Send Real Predictions!

🧠 What Does "Invoke" Mean?

Your fine-tuned model is now live! 🎉 "Invoke" simply means sending an HTTP request to your deployed endpoint URL and getting back a response from your AI.

Just like calling a friend on the phone 📞 — you dial their number (endpoint URL), say something (your input), and they respond (the AI's answer)!

📝 What the code below does:
This sends a real HTTP request to your live deployed LLM endpoint using proper OCI authentication. It's exactly what your mobile app, website, or backend service would do in a real production system! The AI responds in seconds. 🤖⚡
import requests
import json
import oci
from oci.signer import Signer
import time

# ─── 7A: Set Up OCI Request Signing ────────────────────────────
# OCI requires all API calls to be signed with your credentials.
# The Signer handles this automatically — like a digital signature on a letter.

signer = oci.auth.signers.get_resource_principals_signer()

endpoint_url = f"{deployment.url}/predict"

print(f"📡 Invoking Fine-Tuned LLM at:")
print(f"   {endpoint_url}\n")

# ─── 7B: Function to Call the Deployed Model ──────────────────

def call_llm(question, max_tokens=300, temperature=0.7):
    """
    Sends a question to your deployed fine-tuned model and returns the response.
    
    question    = The user's question
    max_tokens  = Maximum words in response
    temperature = Creativity level (0=focused, 1=creative)
    """
    payload = {
        "input"      : question,
        "max_tokens" : max_tokens,
        "temperature": temperature
    }

    start_time = time.time()

    response = requests.post(
        url     = endpoint_url,
        json    = payload,
        auth    = signer,
        headers = {"Content-Type": "application/json"},
        timeout = 120   # LLMs can take 10-60 seconds — set a generous timeout!
    )

    elapsed = (time.time() - start_time) * 1000   # milliseconds

    if response.status_code == 200:
        result = response.json()
        return {
            "answer"       : result.get("response", "No response"),
            "status"       : result.get("status", "unknown"),
            "response_time": f"{elapsed:.0f}ms",
            "output_length": result.get("output_length", 0)
        }
    else:
        return {
            "error"        : f"HTTP {response.status_code}: {response.text}",
            "response_time": f"{elapsed:.0f}ms"
        }


# ─── 7C: Test with Multiple Customer Questions ─────────────────

test_questions = [
    "How do I reset my password?",
    "I was charged twice this month. Can you help me get a refund?",
    "What subscription plans do you offer?",
    "My product stopped working after the latest update. What should I do?"
]

print("🤖 Testing Fine-Tuned Customer Support LLM:\n")
print("=" * 60)

for i, question in enumerate(test_questions, 1):
    print(f"\n💬 Question {i}: {question}")
    print(f"   {'─' * 50}")

    result = call_llm(question)

    if "error" in result:
        print(f"   ❌ Error: {result['error']}")
    else:
        print(f"   🤖 AI Response:")
        # Wrap long responses for readability
        answer = result["answer"]
        words  = answer.split()
        line   = "   "
        for word in words:
            if len(line) + len(word) > 65:
                print(line)
                line = "   "
            line += word + " "
        if line.strip():
            print(line)

        print(f"\n   ⏱️  Response time : {result['response_time']}")
        print(f"   📊 Output length : {result['output_length']} chars")
        print(f"   ✅ Status        : {result['status']}")

print("\n" + "=" * 60)
print("✅ STEP 7 COMPLETE — Invocation tests successful!")

Sample Output:

💬 Question 1: How do I reset my password?
   ──────────────────────────────────────────────────
   🤖 AI Response:
   To reset your password, please visit our login page and click
   on "Forgot Password". You'll receive a secure reset link to your
   registered email address within 2 minutes. If you don't receive
   it, check your spam folder. Need further help? Call us at 1-800-TECHCORP.

   ⏱️  Response time : 4,312ms
   📊 Output length : 287 chars
   ✅ Status        : success

🔧 Advanced Invocation Patterns

📝 What the code below does:
This shows three production-grade invocation patterns you'll need in a real application: (1) Streaming responses for real-time chat, (2) retry logic for reliability, (3) batch processing for efficiency. These patterns are what separate hobbyist deployments from professional production systems! 🏆
import time, random

# ─── Pattern 1: Retry Logic (for production reliability) ───────
def call_llm_with_retry(question, max_retries=3, backoff_base=2):
    """
    Calls the LLM with automatic retry on failure.
    Uses exponential backoff: waits 2s, 4s, 8s between retries.
    Essential for production systems where occasional timeouts happen!
    """
    for attempt in range(1, max_retries + 1):
        try:
            result = call_llm(question)
            if "error" not in result:
                return result   # Success!

            # If first or second attempt failed, wait and retry
            if attempt < max_retries:
                wait_time = backoff_base ** attempt + random.uniform(0, 1)
                print(f"   ⚠️  Attempt {attempt} failed. Retrying in {wait_time:.1f}s...")
                time.sleep(wait_time)

        except requests.exceptions.Timeout:
            if attempt < max_retries:
                print(f"   ⏳ Timeout on attempt {attempt}. Retrying...")
                time.sleep(backoff_base ** attempt)

    return {"error": f"All {max_retries} attempts failed.", "status": "error"}


# ─── Pattern 2: Multi-Turn Conversation ────────────────────────
def call_llm_with_history(conversation_history, new_message, max_tokens=400):
    """
    Maintains conversation context across multiple turns.
    Sends the full conversation history with each request
    so the model remembers what was said before!
    """
    # Build a context string from history
    context = ""
    for turn in conversation_history[-4:]:   # Only last 4 turns (context window limit!)
        role     = turn.get("role", "user")
        content  = turn.get("content", "")
        context += f"{role.capitalize()}: {content}\n"

    # Add the new message
    full_prompt = f"{context}User: {new_message}"

    return call_llm(
        question   = full_prompt,
        max_tokens = max_tokens
    )


# Example multi-turn conversation
print("💬 Multi-Turn Conversation Demo:\n")

history = []

questions = [
    "Hi! I need help with my account.",
    "I forgot the email I used to sign up.",
    "My phone number is +1-555-0123"
]

for q in questions:
    print(f"   👤 User: {q}")
    result = call_llm_with_history(history, q)
    answer = result.get("answer", "Error")
    print(f"   🤖 AI  : {answer[:100]}...\n")

    # Add to history for next turn
    history.append({"role": "user",      "content": q})
    history.append({"role": "assistant", "content": answer})


# ─── Pattern 3: Health Check Before Sending Traffic ────────────
def health_check(endpoint_url, signer):
    """
    Sends a lightweight test request to verify the deployment is responding.
    Use this in your monitoring cron job every 5 minutes!
    """
    try:
        test_response = requests.post(
            url     = endpoint_url,
            json    = {"input": "ping", "max_tokens": 10},
            auth    = signer,
            timeout = 30
        )
        if test_response.status_code == 200:
            return {"healthy": True, "status_code": 200}
        else:
            return {"healthy": False, "status_code": test_response.status_code}
    except Exception as e:
        return {"healthy": False, "error": str(e)}

health = health_check(endpoint_url, signer)
print(f"💓 Health Check: {'✅ HEALTHY' if health['healthy'] else '❌ UNHEALTHY'}")

🗺️ Complete Architecture — Everything in One View

🏛️ Fine-Tuned LLM Deployment on OCI — Full Architecture

Step Action OCI Service Used Time Required
🔐 1. Authenticate ads.set_auth("resource_principal") OCI IAM / Identity ~10 seconds
🔧 2. Init Pipeline Load tokenizer + model, build HF pipeline OCI Object Storage + GPU notebook 5-15 min (download)
📦 3. Prepare Artifact HuggingFacePipelineModel.prepare() + custom score.py Local filesystem ~2 minutes
🔬 4. Introspection model_artifact.introspect() Local ADS validation ~30 seconds
✅ 5. Verify model_artifact.verify(test_payload) Local execution 30s - 5 min
🚀 6. Deploy .save() + .deploy() with GPU shape OCI Model Catalog + OCI Model Deployment 10-20 minutes
📡 7. Invoke requests.post(endpoint_url) with OCI signer OCI Model Deployment REST API 2-10 sec per call

🛠️ Troubleshooting Guide — When Things Go Wrong

Problem Likely Cause Fix
Deployment stays in CREATING forever load_model() taking too long (model download) Pre-download weights to OCI Object Storage, increase timeout
Deployment FAILED state score.py syntax error or import failure Check OCI Log Explorer → predict logs for the exact error
HTTP 500 on invoke Exception in predict() function Add try/except in score.py, check predict logs
CUDA Out of Memory Model too large for GPU shape Use larger GPU shape or add 4-bit quantization
HTTP 401 Unauthorized Wrong authentication method Use OCI Signer — not raw API keys — for invoke calls
Response takes 5+ minutes Running on CPU instead of GPU Verify GPU shape is selected AND device_map="auto" in score.py
Wrong model outputs (hallucinations) Wrong prompt template used in score.py Verify prompt template EXACTLY matches what was used in fine-tuning!

🏆 Best Practices — The Expert Checklist

✅ DOs — What Production LLM Teams Do:

  • ✅ Store model weights in OCI Object Storage — NEVER in the artifact folder
  • ✅ Use 4-bit quantization for 7B+ models — saves 4x GPU memory
  • ✅ Use the EXACT same prompt template in score.py as used in fine-tuning
  • ✅ Always run introspection + verify before saving to catalog
  • ✅ Enable prediction logs → they're your only debug tool in production
  • ✅ Add retry logic in your invoke client — LLMs occasionally timeout
  • ✅ Set TOKENIZERS_PARALLELISM=false environment variable to avoid warnings
  • ✅ Test with edge cases: empty input, very long input, special characters
  • ✅ Use deployment.update() to swap models without downtime
  • ✅ Set up OCI Monitoring alarms on deployment state and response latency
❌ DON'Ts — Mistakes That Cost Hours of Debugging:

  • ❌ Don't use a different prompt template in score.py than you used in training
  • ❌ Don't include large model weight files in the artifact folder
  • ❌ Don't deploy without running .verify() with edge case inputs
  • ❌ Don't use CPU shapes for LLM deployment — response times will be unacceptable
  • ❌ Don't hardcode absolute paths in score.py — they won't work on OCI servers
  • ❌ Don't skip logging in score.py — you'll be blind when debugging production issues
  • ❌ Don't leave test deployments running — GPU shapes cost $3-15/hour!
  • ❌ Don't return numpy or tensor types from predict() — use .tolist() or str()

📝 Quick Summary — The Complete 7-Step Journey

  • Authenticate = Set resource_principal auth and configure OCI region, project, compartment
  • Initialize Pipeline = Download fine-tuned weights, load tokenizer + model, build HuggingFace pipeline with 4-bit quantization
  • Prepare Artifact = Use HuggingFacePipelineModel.prepare(), write custom score.py with correct prompt template
  • Introspection = model_artifact.introspect() validates file structure, function presence, and schema validity
  • Verify = model_artifact.verify(payload) runs score.py end-to-end locally — catches bugs before OCI upload
  • Deploy = .save() uploads to Model Catalog, .deploy() launches a GPU-powered REST API
  • Invoke = requests.post(endpoint_url, json=payload, auth=signer) calls your live AI from anywhere!

Comments