Skip to main content

Oracle ADS Library: Explore AI and Machine Learning Capabilities on OCI

Calculating read time…

Imagine you are building a LEGO castle. 🏰 You could carve each brick from raw stone yourself — OR you could open a box where all the bricks are already shaped, colored, and ready to click together!

That's exactly what Oracle's ADS (Accelerated Data Science) Library is. It's a massive pre-built toolbox that makes building AI and ML models on OCI dramatically faster, easier, and more professional — even for beginners!




📚 What We Will Learn

  • 🧠 What is Oracle ADS? (The big picture)
  • 🔧 ADS vs Doing It Manually — Why ADS Wins Every Time
  • 📦 Installing and Setting Up ADS
  • 📂 Loading and Exploring Data with ADS
  • 🤖 Building ML Models with AutoML (Zero-Code AI!)
  • 📊 Model Evaluation and Explanation with ADS
  • 🏛️ Saving Models to OCI Model Catalog
  • 🚀 Deploying Models as REST APIs
  • 🔄 ADS Pipelines — Automate Everything
  • 🏆 Best Practices Checklist 

🧠 Section 1: What is Oracle ADS?

ADS = Accelerated Data Science. It is an open-source Python library built by Oracle specifically for working inside OCI Data Science.

Think of it like this — if regular Python is a bicycle 🚲, then ADS is a rocket ship 🚀. They both get you there, but ADS gets you there in a fraction of the time!

ADS wraps together the best of pandas, scikit-learn, AutoML, MLflow, OCI services, and model deployment — all into one clean, simple Python API. You write fewer lines of code and get more done.

⭐ One-Line Summary:
ADS = The easiest way to go from raw data → trained model → deployed API entirely within OCI, using Python you already know!
🗺️ The ADS Universe — Everything It Covers

📂 Data Loading
CSV, Parquet, DB, Object Storage
🔍 EDA
Auto data profiling & visualizations
🤖 AutoML
Auto model selection & tuning
📊 Evaluation
Metrics, plots, explainability
🏛️ Model Catalog
Version, tag, store models
🚀 Deployment
REST API with one command
🔄 Pipelines
End-to-end ML automation
🧪 Feature Store
Reusable feature engineering

⚔️ Section 2: ADS vs Doing It Manually

Let's look at how much work ADS saves you. We'll use a real example: loading data from OCI Object Storage.

Task Without ADS (Manual) With ADS ✅
Load CSV from Object Storage 20+ lines (OCI SDK + boto3 + pandas) 2 lines
Data profiling / EDA Write custom plots, stats manually 1 method call
Train best ML model Try 10+ algorithms manually AutoML does it automatically
Save model to OCI Manual SDK calls, file packaging model.save() — one line
Deploy as REST API Docker + OCI setup (hours) model.deploy() — one line

The difference is massive. A senior engineer might spend 3 days setting up a proper ML pipeline manually. With ADS, a beginner can do it in a few hours on Day 1! 🎉


📦 Section 3: Installing and Setting Up ADS

ADS is pre-installed in every OCI Data Science notebook session! 🎉 You don't need to do anything special when working inside OCI.

But if you want to use ADS on your local computer or in a custom environment, here's how to install it:

📝 What the code below does:
This installs the Oracle ADS library on your computer — like downloading a new app. The [complete] part means "install ADS with ALL its extra features included." Think of it like getting the full version instead of the lite version! 📱
# Install ADS with all optional features
# Run this in your terminal or Jupyter notebook cell

pip install oracle-ads[complete]

# If you only want the core features (lighter install):
pip install oracle-ads

# Verify the installation worked:
python -c "import ads; print(f'ADS version: {ads.__version__}')"

🔑 Step 2: Authenticate with OCI

Before ADS can talk to OCI services (Object Storage, Model Catalog, etc.), it needs to know who you are. This is called authentication — like showing your ID card at the entrance! 🪪

📝 What the code below does:
This tells ADS how to connect to your OCI account. Inside OCI Data Science notebooks, resource_principal is the best choice — it's like the notebook automatically shows its OCI employee badge! 🏢
On your local machine, it uses your ~/.oci/config file (API key method).
import ads

# Method 1: Inside OCI Data Science notebook (RECOMMENDED)
# The notebook authenticates itself automatically — no passwords needed!
ads.set_auth("resource_principal")

# Method 2: On your local machine using your OCI config file
# (You create this file when you set up OCI CLI on your laptop)
ads.set_auth("api_key")

# Method 3: Inside OCI Cloud Shell
ads.set_auth("instance_principal")

print("✅ ADS is authenticated and ready to use!")
✅ DO: Always use resource_principal when inside OCI Data Science notebooks. It is the most secure method because no passwords or keys are stored in your code!
❌ DON'T: Never hardcode your OCI API keys or secrets directly inside Python code. If you share the notebook, anyone can steal your credentials! 😱

📂 Section 4: Loading and Exploring Data with ADS

The core data object in ADS is called ADSDataset. Think of it like a supercharged pandas DataFrame that also knows about OCI, can automatically profile your data, and understands ML concepts!

📥 Loading Data from OCI Object Storage

📝 What the code below does:
This loads a CSV file stored in OCI Object Storage (Oracle's cloud file storage — like Google Drive but for code) directly into Python. Without ADS, this would take 20+ lines of complex SDK code. With ADS, it's just 2 lines! Like magic! 🪄
import ads
from ads.dataset.factory import DatasetFactory

ads.set_auth("resource_principal")

# Load a CSV file directly from OCI Object Storage
# Replace with your actual bucket name and file path
dataset = DatasetFactory.open(
    "oci://my-bucket@my-namespace/data/customer_churn.csv",
    target="churn"   # Tell ADS which column is the label we want to predict
)

print(f"✅ Dataset loaded!")
print(f"   Shape      : {dataset.shape}")
print(f"   Target col : churn")
print(f"   Features   : {list(dataset.feature_names)[:5]} ...")

Output:

✅ Dataset loaded!
   Shape      : (10000, 21)
   Target col : churn
   Features   : ['age', 'tenure', 'monthly_charges', 'contract_type', 'payment_method'] ...

🔍 Auto Exploratory Data Analysis (EDA)

EDA means looking at your data carefully before building any model — like reading the recipe before cooking! 🍳 ADS can do a full data health report with ONE line of code:

📝 What the code below does:
This runs a full automatic analysis of your dataset. It checks every column, finds missing values, detects outliers, shows distributions, and even highlights which features might be important for prediction. It's like having a data scientist review your data and hand you a full report! 📋
# ADS auto-profile gives you a full dashboard of your data
# This generates interactive charts inside your Jupyter notebook

dataset.show_in_notebook()

# Want a quick summary? Use this:
print(dataset.summary())

# Check data types and missing values at a glance:
print(dataset.head())
⭐ Pro Tip:
show_in_notebook() generates interactive visualizations for EVERY column automatically! It shows histograms for numbers, bar charts for categories, correlation heatmaps, and even flags potential data quality issues — all without you writing a single plot!

📊 Loading Data from Different Sources

📝 What the code below does:
ADS can load data from many different places — not just CSV files! This shows examples of loading from a database, a pandas DataFrame, a local file, and even directly from a URL on the internet. ADS figures out the format automatically — you just give it the path! 🗂️
from ads.dataset.factory import DatasetFactory
import pandas as pd

# --- From a local CSV file ---
ds_local = DatasetFactory.open("./my_data.csv", target="label")

# --- From an existing pandas DataFrame ---
raw_df = pd.read_csv("./data.csv")
ds_from_df = DatasetFactory.open(raw_df, target="price")

# --- From OCI Object Storage (Parquet format) ---
ds_parquet = DatasetFactory.open(
    "oci://analytics-bucket@my-namespace/sales/q4_2025.parquet",
    target="revenue"
)

# --- From an Oracle Autonomous Database ---
ds_db = DatasetFactory.open(
    "oracle+cx_oracle://user:password@hostname:1521/ORCL",
    table="CUSTOMER_FEATURES",
    target="is_premium"
)

print("✅ ADS can load from virtually any source!")

🤖 Section 5: AutoML — Let AI Build Your AI!

Here's where ADS gets truly magical. ✨ AutoML (Automated Machine Learning) means you don't have to decide which algorithm to use, how to tune it, or how to preprocess the data. ADS tries many options automatically and picks the best one!

Imagine you want to bake the best cake. 🎂 Instead of trying one recipe, AutoML bakes 50 different cakes simultaneously and tells you which one tastes best. Then you serve that one!

🔄 How AutoML Works Inside ADS

📂 Your Dataset
→
🧹 Auto Data Cleaning
→
⚙️ Feature Engineering
→
🏁 Try Many Algorithms
→
🏆 Best Model Returned!

ADS AutoML tries: Random Forest, XGBoost, LightGBM, SVM, Logistic Regression, Neural Networks & more — all automatically!

📝 What the code below does:
This is the most powerful thing you can do with ADS in just a few lines! It automatically cleans your data, engineers features, tries many ML algorithms, tunes their settings, and returns the best model — ready to make predictions. It's like hiring an entire data science team and getting results in minutes! ⚡
import ads
from ads.dataset.factory import DatasetFactory
from ads.automl.provider import OracleAutoMLProvider
from ads.automl.operator import AutoMLOperator

ads.set_auth("resource_principal")

# Step 1: Load dataset with target column specified
dataset = DatasetFactory.open(
    "oci://ml-bucket@namespace/customer_churn.csv",
    target="churn"
)

# Step 2: Split into training and test sets
# 80% for training, 20% for final evaluation
train, test = dataset.train_test_split(test_size=0.2, random_state=42)

print(f"Training samples : {train.shape[0]}")
print(f"Test samples     : {test.shape[0]}")

# Step 3: Set up Oracle's AutoML engine
automl_provider = OracleAutoMLProvider()

# Step 4: Launch AutoML!
# ADS will now automatically try many algorithms and pick the best one.
# This may take a few minutes depending on dataset size.
model, baseline = AutoMLOperator()\
    .with_estimator(automl_provider)\
    .with_dataset(train)\
    .train()

print("\n✅ AutoML complete!")
print(f"   Best Algorithm  : {model.estimator.__class__.__name__}")
print(f"   Best Score      : {model.score(test.X, test.y.values):.4f}")

Sample Output:

Training samples : 8000
Test samples     : 2000

✅ AutoML complete!
   Best Algorithm  : LGBMClassifier
   Best Score      : 0.9312

With just these few lines, ADS tried RandomForest, XGBoost, LightGBM, Logistic Regression, and more — then automatically selected LightGBM as the winner with 93% accuracy. 🏆

✅ DO: Always let AutoML run first before manually experimenting. It gives you an excellent baseline and often beats hand-tuned models! Use the best AutoML model as your starting point, then fine-tune if needed.

📊 Section 6: Model Evaluation and Explainability

Training a model is only half the job. You also need to understand how good it is and why it makes each decision. ADS makes both of these super easy!

📈 Evaluating Model Performance

📝 What the code below does:
This creates a full evaluation report for your trained model. It automatically calculates accuracy, precision, recall, F1 score, and draws the ROC curve and Confusion Matrix — all in one call! Think of it as a school report card for your AI model. 📄
from ads.evaluations.evaluator import ADSEvaluator
from ads.common.model import ADSModel

# Wrap both models (AutoML winner + baseline) for comparison
automl_model   = ADSModel.from_estimator(model.estimator)
baseline_model = ADSModel.from_estimator(baseline.estimator)

# Create an evaluator comparing both models side by side
evaluator = ADSEvaluator(
    test,                             # your test dataset
    models=[automl_model, baseline_model],
    training_data=train
)

# Generate the full evaluation dashboard in your notebook
evaluator.show_in_notebook()

# Access specific metrics programmatically
metrics = evaluator.metrics

for model_name, scores in metrics.items():
    print(f"\n📊 {model_name}")
    print(f"   Accuracy  : {scores.get('accuracy', 'N/A'):.4f}")
    print(f"   F1 Score  : {scores.get('f1', 'N/A'):.4f}")
    print(f"   AUC-ROC   : {scores.get('roc_auc', 'N/A'):.4f}")

🧠 Explainability — Why Did the Model Decide That?

Imagine your AI rejects a loan application. 🏦 The customer asks: "Why was I rejected?" If you can't explain your model's decision, you're in trouble — legally and ethically!

ADS has built-in SHAP (SHapley Additive exPlanations) support, which is the industry standard for explaining AI decisions .

📝 What the code below does:
This calculates SHAP values for your model — basically a "blame score" for each feature. It shows you which features pushed the prediction higher and which pulled it lower. For a loan rejection: "Your credit score (-0.34) and high debt ratio (-0.28) are the main reasons." 📉
from ads.explanations.explainer import ADSExplainer

# Create the explainer for your model
explainer = ADSExplainer(
    model    = automl_model,
    dataset  = test,
    training = train
)

# Get global feature importance (which features matter most overall?)
global_explanation = explainer.global_explanation

# Show a bar chart of feature importance in your notebook
global_explanation.show_in_notebook(
    mode="bar",
    num_features=10   # Show top 10 most important features
)

# Get local explanation for a single prediction (why THIS one?)
# Let's explain what the model decided for customer #42
single_customer = test.iloc[42:43]
local_explanation = explainer.local_explanation

local_explanation.show_in_notebook(
    mode="waterfall",   # Waterfall chart shows each feature's contribution
    num_features=8
)

# Print feature importances as text
print("\n🏆 Top 5 Most Important Features (Global):")
for feature, score in global_explanation.feature_importance[:5]:
    print(f"   {feature:<25 :="" code="" score:.4f="">

Sample Output:

🏆 Top 5 Most Important Features (Global):
   tenure                    : 0.3841
   monthly_charges           : 0.2917
   contract_type             : 0.1823
   tech_support              : 0.0934
   payment_method            : 0.0712
⭐ Regulation Alert!
AI explainability is no longer optional in many industries. EU's AI Act and various financial regulators now require that AI decisions can be explained to customers. ADS's built-in SHAP support makes compliance much easier! ✅

🏛️ Section 7: Saving Models to OCI Model Catalog

After training a great model, you need to store it safely so your team can use it, track its version history, and deploy it anytime.

The OCI Model Catalog is like a professional library 📚 for your AI models. Every model gets its own shelf with a label showing: who created it, when, what data it used, and how well it performs.

🏛️ What OCI Model Catalog Stores for Each Model

  • 📦 Model artifacts — the actual trained model files
  • 🐍 Conda environment — all Python packages needed to run it
  • 📊 Custom metadata — accuracy, PSI scores, retrain dates
  • 📄 Schema — what inputs it expects and what outputs it gives
  • 📋 Taxonomy tags — framework, algorithm, use case labels
  • 🔒 Access control — who can use or modify the model
📝 What the code below does:
This packages your trained model into a format OCI understands, adds useful metadata labels (like sticking a detailed tag on a jar!), and uploads the whole thing to OCI Model Catalog with one command. Later, your entire team can download and use this exact model! 🤝
import ads
from ads.model.generic_model import GenericModel

ads.set_auth("resource_principal")

# Step 1: Wrap your trained AutoML model for OCI packaging
oci_model = GenericModel(
    estimator    = model.estimator,   # your best AutoML model
    artifact_dir = "./churn_model_v1" # temp folder to build the package
)

# Step 2: Prepare — ADS creates all required files automatically
# (model pickle, score.py inference script, conda env details)
oci_model.prepare(
    inference_conda_env = "generalml_p38_cpu_v1",  # OCI pre-built environment
    force_overwrite     = True,
    use_case_type       = "binary_classification"
)

# Step 3: Add custom metadata — attach useful information to the model
oci_model.metadata_custom.add(
    key         = "accuracy",
    value       = "0.9312",
    description = "Test set accuracy achieved by AutoML winner (LGBMClassifier)"
)
oci_model.metadata_custom.add(
    key         = "training_data_date",
    value       = "2026-01-15",
    description = "Date the training dataset was extracted from production"
)
oci_model.metadata_custom.add(
    key         = "use_case",
    value       = "Customer Churn Prediction — Telecom Dataset",
    description = "What business problem this model solves"
)

# Step 4: Add taxonomy — helps team search and categorize models
oci_model.metadata_taxonomy.set(
    category  = "UseCaseType",
    value     = "binary_classification"
)
oci_model.metadata_taxonomy.set(
    category  = "Framework",
    value     = "LightGBM"
)

# Step 5: Save to OCI Model Catalog!
model_id = oci_model.save(
    display_name = "Customer_Churn_AutoML_v1",
    description  = "Best model from AutoML run. LightGBM. 93.1% accuracy on Jan 2026 data.",
    project_id   = "ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID",
    compartment_id = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID"
)

print(f"✅ Model saved to OCI Model Catalog!")
print(f"   Model ID    : {model_id}")
print(f"   Name        : Customer_Churn_AutoML_v1")
print(f"   Accuracy    : 93.12%")
print(f"   Ready for deployment! 🚀")

🚀 Section 8: Deploying Models as REST APIs

A trained model sitting in a catalog doesn't help anyone by itself. You need to deploy it so other applications can send it data and get predictions back. This is called a REST API endpoint.

Think of it like opening a restaurant. 🍽️ Your model is the chef. The REST API is the counter where customers place orders (send data) and pick up their food (receive predictions)!

📝 What the code below does:
This takes your saved model from OCI Model Catalog and turns it into a live web service (REST API)! After this runs, any application anywhere in the world can send customer data to a URL and instantly get a churn prediction back. It's like flipping the "OPEN" sign on your restaurant! 🪧
from ads.model.deployment import ModelDeployment, ModelDeploymentProperties

# Step 1: Configure the deployment settings
deployment_props = ModelDeploymentProperties(
    model_id       = model_id,    # OCID from the save step above
    project_id     = "ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID",
    compartment_id = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID",

    display_name   = "churn-prediction-api-v1",
    description    = "Customer Churn Prediction REST API — Production",

    # Compute shape for serving predictions
    instance_shape = "VM.Standard.E4.Flex",
    instance_count = 2,           # Run 2 instances for high availability
    ocpus          = 1,
    memory_in_gbs  = 16,

    # Auto-scaling settings (best practice!)
    bandwidth_mbps = 10
)

# Step 2: Deploy! (This takes ~5-10 minutes)
deployment = ModelDeployment()
deployment.deploy(
    properties    = deployment_props,
    wait_for_completion = True   # Wait until deployment is live
)

print("✅ Model deployed successfully!")
print(f"   Endpoint URL : {deployment.url}")
print(f"   Status       : {deployment.state.name}")

📡 Making Predictions from the Deployed API

📝 What the code below does:
This sends a real customer's data to your live AI model and gets a churn prediction back — just like a real production application! Any app (website, mobile app, another service) can call this URL to get instant AI-powered predictions. 📱
import requests
import json
import oci
from oci.signer import Signer

# Set up OCI authentication for the API call
config = oci.config.from_file("~/.oci/config")
signer = Signer(
    tenancy  = config["tenancy"],
    user     = config["user"],
    fingerprint = config["fingerprint"],
    private_key_file_location = config["key_file"]
)

# The customer data we want to predict churn for
customer_data = {
    "data": [
        {
            "tenure"           : 24,
            "monthly_charges"  : 89.50,
            "contract_type"    : "month-to-month",
            "tech_support"     : "No",
            "payment_method"   : "electronic_check",
            "total_charges"    : 2148.0
        }
    ]
}

# Send prediction request to your deployed API
endpoint_url = deployment.url + "/predict"

response = requests.post(
    url     = endpoint_url,
    json    = customer_data,
    auth    = signer
)

result = response.json()
print("🤖 Churn Prediction Result:")
print(f"   Prediction  : {'Will Churn ⚠️' if result['prediction'][0] == 1 else 'Will Stay ✅'}")
print(f"   Confidence  : {result['probability'][0][1]*100:.1f}%")
print(f"   Response ms : {response.elapsed.total_seconds()*1000:.0f}ms")

Sample Output:

🤖 Churn Prediction Result:
   Prediction  : Will Churn ⚠️
   Confidence  : 87.3%
   Response ms : 42ms
✅ Amazing!
Your model is now live on the internet, returning predictions in 42 milliseconds! You can now build a customer retention dashboard, a mobile app, or an automated email campaign that uses these predictions in real time! 🎉

🔄 Section 9: ADS ML Pipelines — Automate Everything!

So far we've done each step manually in a notebook. But in production, you want everything to run automatically — data loads, model retrains, evaluations, deployments — on a schedule.

ADS ML Pipelines let you chain all these steps together like an assembly line in a factory. 🏭 Once set up, the pipeline runs by itself every week, month, or whenever new data arrives!

🏭 ADS ML Pipeline — Assembly Line for AI

📥 Step 1
Load Data
Object Storage
→
🔬 Step 2
Validate
Quality checks
→
🏋️ Step 3
Train
AutoML
→
📊 Step 4
Evaluate
Metrics check
→
🚀 Step 5
Deploy
If accuracy ✅

Each step runs as an OCI Data Science Job. The pipeline orchestrates them automatically!

📝 What the code below does:
This creates a full ML Pipeline with 3 steps: (1) load & validate data, (2) train the model, (3) evaluate and save. Each step runs as a separate OCI Job that can use different compute resources. Once created, you can trigger this pipeline manually or on a schedule! ⏰
from ads.pipeline import Pipeline, PipelineStep
from ads.jobs import DataScienceJob, PythonRuntime

# Step 1: Define the data validation step
validate_step = PipelineStep("validate-data")\
    .with_job_id("ocid1.datasciencejob.oc1..VALIDATE_JOB_OCID")\
    .with_description("Load data from OCI Object Storage, run quality checks")

# Step 2: Define the model training step
train_step = PipelineStep("train-model")\
    .with_job_id("ocid1.datasciencejob.oc1..TRAIN_JOB_OCID")\
    .with_description("Run AutoML on validated data, pick best model")\
    .with_depends_on([validate_step])  # Only runs AFTER validation passes!

# Step 3: Define evaluation + save step
evaluate_step = PipelineStep("evaluate-and-save")\
    .with_job_id("ocid1.datasciencejob.oc1..EVALUATE_JOB_OCID")\
    .with_description("Evaluate accuracy. If > 85%, save to Model Catalog.")\
    .with_depends_on([train_step])

# Create the full pipeline connecting all steps
pipeline = Pipeline("Weekly_Churn_Model_Retrain")\
    .with_compartment_id("ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID")\
    .with_project_id("ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID")\
    .with_description("Weekly automated retraining pipeline for churn model")\
    .with_step_details([validate_step, train_step, evaluate_step])

# Create and run the pipeline!
pipeline.create()
pipeline_run = pipeline.run()

print(f"✅ Pipeline created and running!")
print(f"   Pipeline ID  : {pipeline.id}")
print(f"   Run ID       : {pipeline_run.id}")
print(f"   Steps        : validate-data → train-model → evaluate-and-save")
print(f"\n   🕐 Schedule this to run every Sunday night for automatic retraining!")

🧪 Section 10: ADS Feature Store — Reuse Your Best Features!

In big companies, 10 different ML teams might all need the same features — like "customer age," "days since last purchase," or "credit score bucket."

Without a Feature Store, each team recalculates these features from scratch. 😱 That's wasteful, inconsistent, and slow! The Feature Store is like a shared kitchen pantry 🥫 — one team prepares the ingredients, everyone else just grabs what they need!

📝 What the code below does:
This creates a Feature Store in OCI, defines a feature group (a set of related features about customers), and saves computed features so any other team or model can instantly use the same data without recalculating everything from scratch! 🏪
from ads.feature_store.feature_store import FeatureStore
from ads.feature_store.entity import Entity
from ads.feature_store.feature_group import FeatureGroup
import pandas as pd

# Step 1: Create the Feature Store (the pantry!)
feature_store = FeatureStore()\
    .with_compartment_id("ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID")\
    .with_name("TelecomMLFeatureStore")\
    .with_description("Shared features for all telecom ML models")\
    .create()

print(f"✅ Feature Store created: {feature_store.id}")

# Step 2: Define an Entity (what the features describe — a Customer)
customer_entity = Entity()\
    .with_feature_store_id(feature_store.id)\
    .with_name("Customer")\
    .with_description("A telecom subscriber — the main entity for churn prediction")\
    .create()

# Step 3: Create a Feature Group (a set of related features)
# These are the features the data engineering team prepared:
churn_features_df = pd.DataFrame({
    'customer_id'          : ['C001', 'C002', 'C003'],
    'tenure_months'        : [24, 6, 48],
    'avg_monthly_spend'    : [89.5, 45.0, 120.0],
    'support_tickets_90d'  : [0, 3, 1],
    'has_streaming_tv'     : [True, False, True],
    'contract_score'       : [0.8, 0.2, 0.9],   # Engineered feature
    'churn_risk_segment'   : ['Low', 'High', 'Low']
})

feature_group = FeatureGroup()\
    .with_feature_store_id(feature_store.id)\
    .with_entity_id(customer_entity.id)\
    .with_name("CustomerChurnFeatures")\
    .with_primary_keys(["customer_id"])\
    .with_input_feature_details(churn_features_df)\
    .create()

# Step 4: Ingest the actual feature data
feature_group.ingest(churn_features_df)

print(f"✅ Features saved to Feature Store!")
print(f"   Any team can now fetch these features with:")
print(f"   feature_group.select().read()")
✅ Best Practice:
OCI Feature Store now supports streaming feature ingestion via OCI Streaming! This means your features can update in real time as customers take actions — perfect for real-time fraud detection and live recommendation engines! ⚡

🤖 Section 11: ADS with OCI GenAI — The Frontier!

ADS doesn't just support classical ML models. It now integrates directly with OCI Generative AI Service — meaning you can build pipelines that include LLMs (Large Language Models) like Command R+, Llama 3, and Cohere directly from Python!

📝 What the code below does:
This connects to OCI's Generative AI service and sends a text prompt to a powerful LLM (like Cohere's Command R+) directly from Python using ADS. No API keys to manage, no external services — everything stays inside OCI! 🔒
from ads.llm import GenerativeAI

# Connect to OCI GenAI with your compartment
genai_model = GenerativeAI(
    compartment_id  = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID",
    service_endpoint = "https://inference.generativeai.us-chicago-1.oci.oraclecloud.com",
    model_id        = "cohere.command-r-plus"
)

# Ask the LLM a question about one of your churned customers
prompt = """
You are a customer retention specialist at a telecom company.
A customer with the following profile has been predicted to churn:
- Tenure: 6 months
- Monthly charges: $45
- Has submitted 3 support tickets in the last 90 days
- Currently on a month-to-month contract
- No streaming services subscribed

Write a short, personalized retention email offering a discount and upgrade.
Keep it under 100 words and friendly in tone.
"""

response = genai_model.generate(
    prompt      = prompt,
    max_tokens  = 200,
    temperature = 0.7   # 0 = predictable, 1 = creative
)

print("📧 AI-Generated Retention Email:")
print("-" * 50)
print(response.generations[0].text)

Now imagine combining this with your churn prediction pipeline: your ML model identifies at-risk customers, and your GenAI model instantly writes personalized retention messages for each one! That's the power of combining classical ML + GenAI . 🔥

⭐ Trend — Compound AI Systems:
The hottest trend in AI right now is combining classical ML models (prediction, classification) with LLMs (reasoning, text generation) in a single automated pipeline. ADS is one of the few libraries that natively supports BOTH in the same workflow!

💼 Section 12: ADS Data Science Jobs — Run Code Without a Notebook

Notebooks are great for experimenting. But for production, you need code that runs automatically — even when your laptop is off!

OCI Data Science Jobs (accessible through ADS) let you run Python scripts on powerful cloud machines on a schedule. Think of it as setting up a robot employee who works 24/7! 🤖

📝 What the code below does:
This creates a scheduled job that runs your Python training script every Monday morning automatically on an OCI GPU machine! You set it up once, and it runs forever without any manual work. ⏰
from ads.jobs import DataScienceJob, PythonRuntime, ScriptRuntime

# Step 1: Define the compute environment for the job
job = DataScienceJob()\
    .with_compartment_id("ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID")\
    .with_project_id("ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID")\
    .with_name("Weekly_Model_Retrain_Job")\
    .with_description("Runs every Monday: loads fresh data, retrains churn model, saves to catalog")\
    .with_shape_name("VM.Standard.E4.Flex")\
    .with_shape_config_details(ocpus=4, memory_in_gbs=32)\
    .with_block_storage_size(100)

# Step 2: Define what to run — point to your Python script
runtime = ScriptRuntime()\
    .with_source("./retrain_churn_model.py")\
    .with_service_conda("generalml_p38_cpu_v1")\
    .with_environment_variable(
        PSI_THRESHOLD    = "0.2",
        RETRAIN_WINDOW   = "90",
        MODEL_NAME       = "Customer_Churn_AutoML",
        SLACK_WEBHOOK    = "https://hooks.slack.com/services/YOUR/WEBHOOK"
    )

# Step 3: Create the job
job.with_runtime(runtime).create()

print(f"✅ Retraining Job created!")
print(f"   Job ID : {job.id}")
print(f"   Shape  : VM.Standard.E4.Flex (4 OCPUs, 32GB RAM)")
print(f"\n   ⏰ Schedule this job via OCI Resource Scheduler:")
print(f"   Every Monday at 02:00 UTC for fresh weekly retraining!")

📋 Section 13: ADS Quick Reference Cheat Sheet

Here's your one-stop reference card for the most commonly used ADS commands. Bookmark this! 🔖

What You Want To Do ADS Command
Set authentication ads.set_auth("resource_principal")
Load dataset from anywhere DatasetFactory.open("path", target="col")
Auto EDA profile dataset.show_in_notebook()
Train-test split train, test = dataset.train_test_split(test_size=0.2)
Run AutoML AutoMLOperator().with_estimator(OracleAutoMLProvider()).with_dataset(train).train()
Evaluate model ADSEvaluator(test, models=[model]).show_in_notebook()
Explain predictions (SHAP) ADSExplainer(model, test).global_explanation.show_in_notebook()
Save to Model Catalog GenericModel(estimator).prepare(...).save(display_name="...")
Deploy as REST API ModelDeployment().deploy(properties=...)
Create ML Pipeline Pipeline("name").with_step_details([...]).create()
Create a scheduled Job DataScienceJob().with_runtime(ScriptRuntime()).create()
Call OCI GenAI LLM GenerativeAI(compartment_id=...).generate(prompt=...)

🏆 Section 14: Best Practices — The Golden Rules of ADS

✅ DOs — What Every Great ADS Developer Does:

  • ✅ Always use resource_principal auth inside OCI notebooks
  • ✅ Call show_in_notebook() before ANY model training — understand your data first!
  • ✅ Let AutoML run first. Use its best model as your baseline.
  • ✅ Add meaningful metadata to every model saved in Model Catalog
  • ✅ Always evaluate with ADSEvaluator before deploying
  • ✅ Use SHAP explainability for any model used in customer-facing decisions
  • ✅ Use ML Pipelines for production — not manual notebook runs
  • ✅ Store reusable features in Feature Store for team-wide sharing
  • ✅ Set model version names clearly: ModelName_v2_YYYY-MM
  • ✅ Monitor deployed endpoints with OCI Monitoring + set accuracy alerts
❌ DON'Ts — Mistakes That Cost Teams Time and Money:

  • ❌ Don't hardcode credentials anywhere in your code
  • ❌ Don't skip data profiling and go straight to model training
  • ❌ Don't deploy a model without evaluating it on a proper test set
  • ❌ Don't ignore model explainability for regulated use cases
  • ❌ Don't run production retraining manually from a notebook — use Pipelines/Jobs
  • ❌ Don't save models without version numbers — you'll lose track fast
  • ❌ Don't forget to delete unused deployments — they cost money even when idle!
  • ❌ Don't use large compute shapes during development — use Flex shapes wisely

🗺️ Section 15: Full ADS Architecture — The Complete Picture

🏛️ End-to-End ML with ADS on OCI

Every row maps ADS code to the underlying OCI service it uses.

# Stage ADS Tool OCI Service Behind It
1 📂 Data Storage DatasetFactory.open("oci://...") OCI Object Storage
2 🔍 EDA dataset.show_in_notebook() OCI Data Science Notebooks
3 🤖 Training AutoMLOperator.train() OCI Data Science (GPU/CPU compute)
4 📊 Evaluation ADSEvaluator + ADSExplainer OCI Data Science Notebooks
5 🏛️ Storage GenericModel.save() OCI Model Catalog
6 🚀 Deployment ModelDeployment.deploy() OCI Model Deployment (REST API)
7 🔄 Automation Pipeline + DataScienceJob OCI Data Science Jobs + Pipelines
8 🧪 Features FeatureStore + FeatureGroup OCI Feature Store
9 🤖 GenAI GenerativeAI.generate() OCI Generative AI Service

📝 Quick Summary — Everything You Learned Today

  • ADS = Oracle's open-source Python library that makes ML on OCI dramatically faster and easier
  • Authentication = Use resource_principal inside OCI notebooks — most secure!
  • DatasetFactory = Load data from Object Storage, databases, or local files in 2 lines
  • AutoML = Automatically tries many algorithms and picks the best one for you
  • ADSEvaluator = Full model evaluation with one method call — metrics + plots
  • ADSExplainer = SHAP-powered explanations for why the model made each decision
  • GenericModel = Package and save any trained model to OCI Model Catalog
  • ModelDeployment = Deploy any model as a live REST API with one command
  • Pipeline = Chain steps into a production-grade automated ML workflow
  • Feature Store = Shared repository for reusable, consistent ML features
  • GenerativeAI = Call OCI LLMs directly from Python — combine ML + GenAI!

🦸 Keep building! 📈

Comments