Imagine you are building a LEGO castle. 🏰 You could carve each brick from raw stone yourself — OR you could open a box where all the bricks are already shaped, colored, and ready to click together!
That's exactly what Oracle's ADS (Accelerated Data Science) Library is. It's a massive pre-built toolbox that makes building AI and ML models on OCI dramatically faster, easier, and more professional — even for beginners!
📚 What We Will Learn
- 🧠 What is Oracle ADS? (The big picture)
- 🔧 ADS vs Doing It Manually — Why ADS Wins Every Time
- 📦 Installing and Setting Up ADS
- 📂 Loading and Exploring Data with ADS
- 🤖 Building ML Models with AutoML (Zero-Code AI!)
- 📊 Model Evaluation and Explanation with ADS
- 🏛️ Saving Models to OCI Model Catalog
- 🚀 Deploying Models as REST APIs
- 🔄 ADS Pipelines — Automate Everything
- 🏆 Best Practices Checklist
🧠 Section 1: What is Oracle ADS?
ADS = Accelerated Data Science. It is an open-source Python library built by Oracle specifically for working inside OCI Data Science.
Think of it like this — if regular Python is a bicycle 🚲, then ADS is a rocket ship 🚀. They both get you there, but ADS gets you there in a fraction of the time!
ADS wraps together the best of pandas, scikit-learn, AutoML, MLflow, OCI services, and model deployment — all into one clean, simple Python API. You write fewer lines of code and get more done.
ADS = The easiest way to go from raw data → trained model → deployed API entirely within OCI, using Python you already know!
CSV, Parquet, DB, Object Storage
Auto data profiling & visualizations
Auto model selection & tuning
Metrics, plots, explainability
Version, tag, store models
REST API with one command
End-to-end ML automation
Reusable feature engineering
⚔️ Section 2: ADS vs Doing It Manually
Let's look at how much work ADS saves you. We'll use a real example: loading data from OCI Object Storage.
| Task | Without ADS (Manual) | With ADS ✅ |
|---|---|---|
| Load CSV from Object Storage | 20+ lines (OCI SDK + boto3 + pandas) | 2 lines |
| Data profiling / EDA | Write custom plots, stats manually | 1 method call |
| Train best ML model | Try 10+ algorithms manually | AutoML does it automatically |
| Save model to OCI | Manual SDK calls, file packaging | model.save() — one line |
| Deploy as REST API | Docker + OCI setup (hours) | model.deploy() — one line |
The difference is massive. A senior engineer might spend 3 days setting up a proper ML pipeline manually. With ADS, a beginner can do it in a few hours on Day 1! 🎉
📦 Section 3: Installing and Setting Up ADS
ADS is pre-installed in every OCI Data Science notebook session! 🎉 You don't need to do anything special when working inside OCI.
But if you want to use ADS on your local computer or in a custom environment, here's how to install it:
This installs the Oracle ADS library on your computer — like downloading a new app. The
[complete] part means "install ADS with ALL its extra features included."
Think of it like getting the full version instead of the lite version! 📱
# Install ADS with all optional features
# Run this in your terminal or Jupyter notebook cell
pip install oracle-ads[complete]
# If you only want the core features (lighter install):
pip install oracle-ads
# Verify the installation worked:
python -c "import ads; print(f'ADS version: {ads.__version__}')"
🔑 Step 2: Authenticate with OCI
Before ADS can talk to OCI services (Object Storage, Model Catalog, etc.), it needs to know who you are. This is called authentication — like showing your ID card at the entrance! 🪪
This tells ADS how to connect to your OCI account. Inside OCI Data Science notebooks,
resource_principal is the best choice —
it's like the notebook automatically shows its OCI employee badge! 🏢On your local machine, it uses your
~/.oci/config file (API key method).
import ads
# Method 1: Inside OCI Data Science notebook (RECOMMENDED)
# The notebook authenticates itself automatically — no passwords needed!
ads.set_auth("resource_principal")
# Method 2: On your local machine using your OCI config file
# (You create this file when you set up OCI CLI on your laptop)
ads.set_auth("api_key")
# Method 3: Inside OCI Cloud Shell
ads.set_auth("instance_principal")
print("✅ ADS is authenticated and ready to use!")
resource_principal when inside OCI Data Science notebooks.
It is the most secure method because no passwords or keys are stored in your code!
📂 Section 4: Loading and Exploring Data with ADS
The core data object in ADS is called ADSDataset. Think of it like a supercharged pandas DataFrame that also knows about OCI, can automatically profile your data, and understands ML concepts!
📥 Loading Data from OCI Object Storage
This loads a CSV file stored in OCI Object Storage (Oracle's cloud file storage — like Google Drive but for code) directly into Python. Without ADS, this would take 20+ lines of complex SDK code. With ADS, it's just 2 lines! Like magic! 🪄
import ads
from ads.dataset.factory import DatasetFactory
ads.set_auth("resource_principal")
# Load a CSV file directly from OCI Object Storage
# Replace with your actual bucket name and file path
dataset = DatasetFactory.open(
"oci://my-bucket@my-namespace/data/customer_churn.csv",
target="churn" # Tell ADS which column is the label we want to predict
)
print(f"✅ Dataset loaded!")
print(f" Shape : {dataset.shape}")
print(f" Target col : churn")
print(f" Features : {list(dataset.feature_names)[:5]} ...")
Output:
✅ Dataset loaded! Shape : (10000, 21) Target col : churn Features : ['age', 'tenure', 'monthly_charges', 'contract_type', 'payment_method'] ...
🔍 Auto Exploratory Data Analysis (EDA)
EDA means looking at your data carefully before building any model — like reading the recipe before cooking! 🍳 ADS can do a full data health report with ONE line of code:
This runs a full automatic analysis of your dataset. It checks every column, finds missing values, detects outliers, shows distributions, and even highlights which features might be important for prediction. It's like having a data scientist review your data and hand you a full report! 📋
# ADS auto-profile gives you a full dashboard of your data # This generates interactive charts inside your Jupyter notebook dataset.show_in_notebook() # Want a quick summary? Use this: print(dataset.summary()) # Check data types and missing values at a glance: print(dataset.head())
show_in_notebook() generates interactive visualizations for EVERY column automatically!
It shows histograms for numbers, bar charts for categories, correlation heatmaps,
and even flags potential data quality issues — all without you writing a single plot!
📊 Loading Data from Different Sources
ADS can load data from many different places — not just CSV files! This shows examples of loading from a database, a pandas DataFrame, a local file, and even directly from a URL on the internet. ADS figures out the format automatically — you just give it the path! 🗂️
from ads.dataset.factory import DatasetFactory
import pandas as pd
# --- From a local CSV file ---
ds_local = DatasetFactory.open("./my_data.csv", target="label")
# --- From an existing pandas DataFrame ---
raw_df = pd.read_csv("./data.csv")
ds_from_df = DatasetFactory.open(raw_df, target="price")
# --- From OCI Object Storage (Parquet format) ---
ds_parquet = DatasetFactory.open(
"oci://analytics-bucket@my-namespace/sales/q4_2025.parquet",
target="revenue"
)
# --- From an Oracle Autonomous Database ---
ds_db = DatasetFactory.open(
"oracle+cx_oracle://user:password@hostname:1521/ORCL",
table="CUSTOMER_FEATURES",
target="is_premium"
)
print("✅ ADS can load from virtually any source!")
🤖 Section 5: AutoML — Let AI Build Your AI!
Here's where ADS gets truly magical. ✨ AutoML (Automated Machine Learning) means you don't have to decide which algorithm to use, how to tune it, or how to preprocess the data. ADS tries many options automatically and picks the best one!
Imagine you want to bake the best cake. 🎂 Instead of trying one recipe, AutoML bakes 50 different cakes simultaneously and tells you which one tastes best. Then you serve that one!
ADS AutoML tries: Random Forest, XGBoost, LightGBM, SVM, Logistic Regression, Neural Networks & more — all automatically!
This is the most powerful thing you can do with ADS in just a few lines! It automatically cleans your data, engineers features, tries many ML algorithms, tunes their settings, and returns the best model — ready to make predictions. It's like hiring an entire data science team and getting results in minutes! ⚡
import ads
from ads.dataset.factory import DatasetFactory
from ads.automl.provider import OracleAutoMLProvider
from ads.automl.operator import AutoMLOperator
ads.set_auth("resource_principal")
# Step 1: Load dataset with target column specified
dataset = DatasetFactory.open(
"oci://ml-bucket@namespace/customer_churn.csv",
target="churn"
)
# Step 2: Split into training and test sets
# 80% for training, 20% for final evaluation
train, test = dataset.train_test_split(test_size=0.2, random_state=42)
print(f"Training samples : {train.shape[0]}")
print(f"Test samples : {test.shape[0]}")
# Step 3: Set up Oracle's AutoML engine
automl_provider = OracleAutoMLProvider()
# Step 4: Launch AutoML!
# ADS will now automatically try many algorithms and pick the best one.
# This may take a few minutes depending on dataset size.
model, baseline = AutoMLOperator()\
.with_estimator(automl_provider)\
.with_dataset(train)\
.train()
print("\n✅ AutoML complete!")
print(f" Best Algorithm : {model.estimator.__class__.__name__}")
print(f" Best Score : {model.score(test.X, test.y.values):.4f}")
Sample Output:
Training samples : 8000 Test samples : 2000 ✅ AutoML complete! Best Algorithm : LGBMClassifier Best Score : 0.9312
With just these few lines, ADS tried RandomForest, XGBoost, LightGBM, Logistic Regression, and more — then automatically selected LightGBM as the winner with 93% accuracy. 🏆
📊 Section 6: Model Evaluation and Explainability
Training a model is only half the job. You also need to understand how good it is and why it makes each decision. ADS makes both of these super easy!
📈 Evaluating Model Performance
This creates a full evaluation report for your trained model. It automatically calculates accuracy, precision, recall, F1 score, and draws the ROC curve and Confusion Matrix — all in one call! Think of it as a school report card for your AI model. 📄
from ads.evaluations.evaluator import ADSEvaluator
from ads.common.model import ADSModel
# Wrap both models (AutoML winner + baseline) for comparison
automl_model = ADSModel.from_estimator(model.estimator)
baseline_model = ADSModel.from_estimator(baseline.estimator)
# Create an evaluator comparing both models side by side
evaluator = ADSEvaluator(
test, # your test dataset
models=[automl_model, baseline_model],
training_data=train
)
# Generate the full evaluation dashboard in your notebook
evaluator.show_in_notebook()
# Access specific metrics programmatically
metrics = evaluator.metrics
for model_name, scores in metrics.items():
print(f"\n📊 {model_name}")
print(f" Accuracy : {scores.get('accuracy', 'N/A'):.4f}")
print(f" F1 Score : {scores.get('f1', 'N/A'):.4f}")
print(f" AUC-ROC : {scores.get('roc_auc', 'N/A'):.4f}")
🧠 Explainability — Why Did the Model Decide That?
Imagine your AI rejects a loan application. 🏦 The customer asks: "Why was I rejected?" If you can't explain your model's decision, you're in trouble — legally and ethically!
ADS has built-in SHAP (SHapley Additive exPlanations) support, which is the industry standard for explaining AI decisions .
This calculates SHAP values for your model — basically a "blame score" for each feature. It shows you which features pushed the prediction higher and which pulled it lower. For a loan rejection: "Your credit score (-0.34) and high debt ratio (-0.28) are the main reasons." 📉
from ads.explanations.explainer import ADSExplainer
# Create the explainer for your model
explainer = ADSExplainer(
model = automl_model,
dataset = test,
training = train
)
# Get global feature importance (which features matter most overall?)
global_explanation = explainer.global_explanation
# Show a bar chart of feature importance in your notebook
global_explanation.show_in_notebook(
mode="bar",
num_features=10 # Show top 10 most important features
)
# Get local explanation for a single prediction (why THIS one?)
# Let's explain what the model decided for customer #42
single_customer = test.iloc[42:43]
local_explanation = explainer.local_explanation
local_explanation.show_in_notebook(
mode="waterfall", # Waterfall chart shows each feature's contribution
num_features=8
)
# Print feature importances as text
print("\n🏆 Top 5 Most Important Features (Global):")
for feature, score in global_explanation.feature_importance[:5]:
print(f" {feature:<25 :="" code="" score:.4f="">25>
Sample Output:
🏆 Top 5 Most Important Features (Global): tenure : 0.3841 monthly_charges : 0.2917 contract_type : 0.1823 tech_support : 0.0934 payment_method : 0.0712
AI explainability is no longer optional in many industries. EU's AI Act and various financial regulators now require that AI decisions can be explained to customers. ADS's built-in SHAP support makes compliance much easier! ✅
🏛️ Section 7: Saving Models to OCI Model Catalog
After training a great model, you need to store it safely so your team can use it, track its version history, and deploy it anytime.
The OCI Model Catalog is like a professional library 📚 for your AI models. Every model gets its own shelf with a label showing: who created it, when, what data it used, and how well it performs.
- 📦 Model artifacts — the actual trained model files
- 🐍 Conda environment — all Python packages needed to run it
- 📊 Custom metadata — accuracy, PSI scores, retrain dates
- 📄 Schema — what inputs it expects and what outputs it gives
- 📋 Taxonomy tags — framework, algorithm, use case labels
- 🔒 Access control — who can use or modify the model
This packages your trained model into a format OCI understands, adds useful metadata labels (like sticking a detailed tag on a jar!), and uploads the whole thing to OCI Model Catalog with one command. Later, your entire team can download and use this exact model! 🤝
import ads
from ads.model.generic_model import GenericModel
ads.set_auth("resource_principal")
# Step 1: Wrap your trained AutoML model for OCI packaging
oci_model = GenericModel(
estimator = model.estimator, # your best AutoML model
artifact_dir = "./churn_model_v1" # temp folder to build the package
)
# Step 2: Prepare — ADS creates all required files automatically
# (model pickle, score.py inference script, conda env details)
oci_model.prepare(
inference_conda_env = "generalml_p38_cpu_v1", # OCI pre-built environment
force_overwrite = True,
use_case_type = "binary_classification"
)
# Step 3: Add custom metadata — attach useful information to the model
oci_model.metadata_custom.add(
key = "accuracy",
value = "0.9312",
description = "Test set accuracy achieved by AutoML winner (LGBMClassifier)"
)
oci_model.metadata_custom.add(
key = "training_data_date",
value = "2026-01-15",
description = "Date the training dataset was extracted from production"
)
oci_model.metadata_custom.add(
key = "use_case",
value = "Customer Churn Prediction — Telecom Dataset",
description = "What business problem this model solves"
)
# Step 4: Add taxonomy — helps team search and categorize models
oci_model.metadata_taxonomy.set(
category = "UseCaseType",
value = "binary_classification"
)
oci_model.metadata_taxonomy.set(
category = "Framework",
value = "LightGBM"
)
# Step 5: Save to OCI Model Catalog!
model_id = oci_model.save(
display_name = "Customer_Churn_AutoML_v1",
description = "Best model from AutoML run. LightGBM. 93.1% accuracy on Jan 2026 data.",
project_id = "ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID",
compartment_id = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID"
)
print(f"✅ Model saved to OCI Model Catalog!")
print(f" Model ID : {model_id}")
print(f" Name : Customer_Churn_AutoML_v1")
print(f" Accuracy : 93.12%")
print(f" Ready for deployment! 🚀")
🚀 Section 8: Deploying Models as REST APIs
A trained model sitting in a catalog doesn't help anyone by itself. You need to deploy it so other applications can send it data and get predictions back. This is called a REST API endpoint.
Think of it like opening a restaurant. 🍽️ Your model is the chef. The REST API is the counter where customers place orders (send data) and pick up their food (receive predictions)!
This takes your saved model from OCI Model Catalog and turns it into a live web service (REST API)! After this runs, any application anywhere in the world can send customer data to a URL and instantly get a churn prediction back. It's like flipping the "OPEN" sign on your restaurant! 🪧
from ads.model.deployment import ModelDeployment, ModelDeploymentProperties
# Step 1: Configure the deployment settings
deployment_props = ModelDeploymentProperties(
model_id = model_id, # OCID from the save step above
project_id = "ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID",
compartment_id = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID",
display_name = "churn-prediction-api-v1",
description = "Customer Churn Prediction REST API — Production",
# Compute shape for serving predictions
instance_shape = "VM.Standard.E4.Flex",
instance_count = 2, # Run 2 instances for high availability
ocpus = 1,
memory_in_gbs = 16,
# Auto-scaling settings (best practice!)
bandwidth_mbps = 10
)
# Step 2: Deploy! (This takes ~5-10 minutes)
deployment = ModelDeployment()
deployment.deploy(
properties = deployment_props,
wait_for_completion = True # Wait until deployment is live
)
print("✅ Model deployed successfully!")
print(f" Endpoint URL : {deployment.url}")
print(f" Status : {deployment.state.name}")
📡 Making Predictions from the Deployed API
This sends a real customer's data to your live AI model and gets a churn prediction back — just like a real production application! Any app (website, mobile app, another service) can call this URL to get instant AI-powered predictions. 📱
import requests
import json
import oci
from oci.signer import Signer
# Set up OCI authentication for the API call
config = oci.config.from_file("~/.oci/config")
signer = Signer(
tenancy = config["tenancy"],
user = config["user"],
fingerprint = config["fingerprint"],
private_key_file_location = config["key_file"]
)
# The customer data we want to predict churn for
customer_data = {
"data": [
{
"tenure" : 24,
"monthly_charges" : 89.50,
"contract_type" : "month-to-month",
"tech_support" : "No",
"payment_method" : "electronic_check",
"total_charges" : 2148.0
}
]
}
# Send prediction request to your deployed API
endpoint_url = deployment.url + "/predict"
response = requests.post(
url = endpoint_url,
json = customer_data,
auth = signer
)
result = response.json()
print("🤖 Churn Prediction Result:")
print(f" Prediction : {'Will Churn ⚠️' if result['prediction'][0] == 1 else 'Will Stay ✅'}")
print(f" Confidence : {result['probability'][0][1]*100:.1f}%")
print(f" Response ms : {response.elapsed.total_seconds()*1000:.0f}ms")
Sample Output:
🤖 Churn Prediction Result: Prediction : Will Churn ⚠️ Confidence : 87.3% Response ms : 42ms
Your model is now live on the internet, returning predictions in 42 milliseconds! You can now build a customer retention dashboard, a mobile app, or an automated email campaign that uses these predictions in real time! 🎉
🔄 Section 9: ADS ML Pipelines — Automate Everything!
So far we've done each step manually in a notebook. But in production, you want everything to run automatically — data loads, model retrains, evaluations, deployments — on a schedule.
ADS ML Pipelines let you chain all these steps together like an assembly line in a factory. 🏭 Once set up, the pipeline runs by itself every week, month, or whenever new data arrives!
Load Data
Object Storage
Validate
Quality checks
Train
AutoML
Evaluate
Metrics check
Deploy
If accuracy ✅
Each step runs as an OCI Data Science Job. The pipeline orchestrates them automatically!
This creates a full ML Pipeline with 3 steps: (1) load & validate data, (2) train the model, (3) evaluate and save. Each step runs as a separate OCI Job that can use different compute resources. Once created, you can trigger this pipeline manually or on a schedule! ⏰
from ads.pipeline import Pipeline, PipelineStep
from ads.jobs import DataScienceJob, PythonRuntime
# Step 1: Define the data validation step
validate_step = PipelineStep("validate-data")\
.with_job_id("ocid1.datasciencejob.oc1..VALIDATE_JOB_OCID")\
.with_description("Load data from OCI Object Storage, run quality checks")
# Step 2: Define the model training step
train_step = PipelineStep("train-model")\
.with_job_id("ocid1.datasciencejob.oc1..TRAIN_JOB_OCID")\
.with_description("Run AutoML on validated data, pick best model")\
.with_depends_on([validate_step]) # Only runs AFTER validation passes!
# Step 3: Define evaluation + save step
evaluate_step = PipelineStep("evaluate-and-save")\
.with_job_id("ocid1.datasciencejob.oc1..EVALUATE_JOB_OCID")\
.with_description("Evaluate accuracy. If > 85%, save to Model Catalog.")\
.with_depends_on([train_step])
# Create the full pipeline connecting all steps
pipeline = Pipeline("Weekly_Churn_Model_Retrain")\
.with_compartment_id("ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID")\
.with_project_id("ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID")\
.with_description("Weekly automated retraining pipeline for churn model")\
.with_step_details([validate_step, train_step, evaluate_step])
# Create and run the pipeline!
pipeline.create()
pipeline_run = pipeline.run()
print(f"✅ Pipeline created and running!")
print(f" Pipeline ID : {pipeline.id}")
print(f" Run ID : {pipeline_run.id}")
print(f" Steps : validate-data → train-model → evaluate-and-save")
print(f"\n 🕐 Schedule this to run every Sunday night for automatic retraining!")
🧪 Section 10: ADS Feature Store — Reuse Your Best Features!
In big companies, 10 different ML teams might all need the same features — like "customer age," "days since last purchase," or "credit score bucket."
Without a Feature Store, each team recalculates these features from scratch. 😱 That's wasteful, inconsistent, and slow! The Feature Store is like a shared kitchen pantry 🥫 — one team prepares the ingredients, everyone else just grabs what they need!
This creates a Feature Store in OCI, defines a feature group (a set of related features about customers), and saves computed features so any other team or model can instantly use the same data without recalculating everything from scratch! 🏪
from ads.feature_store.feature_store import FeatureStore
from ads.feature_store.entity import Entity
from ads.feature_store.feature_group import FeatureGroup
import pandas as pd
# Step 1: Create the Feature Store (the pantry!)
feature_store = FeatureStore()\
.with_compartment_id("ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID")\
.with_name("TelecomMLFeatureStore")\
.with_description("Shared features for all telecom ML models")\
.create()
print(f"✅ Feature Store created: {feature_store.id}")
# Step 2: Define an Entity (what the features describe — a Customer)
customer_entity = Entity()\
.with_feature_store_id(feature_store.id)\
.with_name("Customer")\
.with_description("A telecom subscriber — the main entity for churn prediction")\
.create()
# Step 3: Create a Feature Group (a set of related features)
# These are the features the data engineering team prepared:
churn_features_df = pd.DataFrame({
'customer_id' : ['C001', 'C002', 'C003'],
'tenure_months' : [24, 6, 48],
'avg_monthly_spend' : [89.5, 45.0, 120.0],
'support_tickets_90d' : [0, 3, 1],
'has_streaming_tv' : [True, False, True],
'contract_score' : [0.8, 0.2, 0.9], # Engineered feature
'churn_risk_segment' : ['Low', 'High', 'Low']
})
feature_group = FeatureGroup()\
.with_feature_store_id(feature_store.id)\
.with_entity_id(customer_entity.id)\
.with_name("CustomerChurnFeatures")\
.with_primary_keys(["customer_id"])\
.with_input_feature_details(churn_features_df)\
.create()
# Step 4: Ingest the actual feature data
feature_group.ingest(churn_features_df)
print(f"✅ Features saved to Feature Store!")
print(f" Any team can now fetch these features with:")
print(f" feature_group.select().read()")
OCI Feature Store now supports streaming feature ingestion via OCI Streaming! This means your features can update in real time as customers take actions — perfect for real-time fraud detection and live recommendation engines! ⚡
🤖 Section 11: ADS with OCI GenAI — The Frontier!
ADS doesn't just support classical ML models. It now integrates directly with OCI Generative AI Service — meaning you can build pipelines that include LLMs (Large Language Models) like Command R+, Llama 3, and Cohere directly from Python!
This connects to OCI's Generative AI service and sends a text prompt to a powerful LLM (like Cohere's Command R+) directly from Python using ADS. No API keys to manage, no external services — everything stays inside OCI! 🔒
from ads.llm import GenerativeAI
# Connect to OCI GenAI with your compartment
genai_model = GenerativeAI(
compartment_id = "ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID",
service_endpoint = "https://inference.generativeai.us-chicago-1.oci.oraclecloud.com",
model_id = "cohere.command-r-plus"
)
# Ask the LLM a question about one of your churned customers
prompt = """
You are a customer retention specialist at a telecom company.
A customer with the following profile has been predicted to churn:
- Tenure: 6 months
- Monthly charges: $45
- Has submitted 3 support tickets in the last 90 days
- Currently on a month-to-month contract
- No streaming services subscribed
Write a short, personalized retention email offering a discount and upgrade.
Keep it under 100 words and friendly in tone.
"""
response = genai_model.generate(
prompt = prompt,
max_tokens = 200,
temperature = 0.7 # 0 = predictable, 1 = creative
)
print("📧 AI-Generated Retention Email:")
print("-" * 50)
print(response.generations[0].text)
Now imagine combining this with your churn prediction pipeline: your ML model identifies at-risk customers, and your GenAI model instantly writes personalized retention messages for each one! That's the power of combining classical ML + GenAI . 🔥
The hottest trend in AI right now is combining classical ML models (prediction, classification) with LLMs (reasoning, text generation) in a single automated pipeline. ADS is one of the few libraries that natively supports BOTH in the same workflow!
💼 Section 12: ADS Data Science Jobs — Run Code Without a Notebook
Notebooks are great for experimenting. But for production, you need code that runs automatically — even when your laptop is off!
OCI Data Science Jobs (accessible through ADS) let you run Python scripts on powerful cloud machines on a schedule. Think of it as setting up a robot employee who works 24/7! 🤖
This creates a scheduled job that runs your Python training script every Monday morning automatically on an OCI GPU machine! You set it up once, and it runs forever without any manual work. ⏰
from ads.jobs import DataScienceJob, PythonRuntime, ScriptRuntime
# Step 1: Define the compute environment for the job
job = DataScienceJob()\
.with_compartment_id("ocid1.compartment.oc1..YOUR_COMPARTMENT_OCID")\
.with_project_id("ocid1.datascienceproject.oc1..YOUR_PROJECT_OCID")\
.with_name("Weekly_Model_Retrain_Job")\
.with_description("Runs every Monday: loads fresh data, retrains churn model, saves to catalog")\
.with_shape_name("VM.Standard.E4.Flex")\
.with_shape_config_details(ocpus=4, memory_in_gbs=32)\
.with_block_storage_size(100)
# Step 2: Define what to run — point to your Python script
runtime = ScriptRuntime()\
.with_source("./retrain_churn_model.py")\
.with_service_conda("generalml_p38_cpu_v1")\
.with_environment_variable(
PSI_THRESHOLD = "0.2",
RETRAIN_WINDOW = "90",
MODEL_NAME = "Customer_Churn_AutoML",
SLACK_WEBHOOK = "https://hooks.slack.com/services/YOUR/WEBHOOK"
)
# Step 3: Create the job
job.with_runtime(runtime).create()
print(f"✅ Retraining Job created!")
print(f" Job ID : {job.id}")
print(f" Shape : VM.Standard.E4.Flex (4 OCPUs, 32GB RAM)")
print(f"\n ⏰ Schedule this job via OCI Resource Scheduler:")
print(f" Every Monday at 02:00 UTC for fresh weekly retraining!")
📋 Section 13: ADS Quick Reference Cheat Sheet
Here's your one-stop reference card for the most commonly used ADS commands. Bookmark this! 🔖
| What You Want To Do | ADS Command |
|---|---|
| Set authentication | ads.set_auth("resource_principal") |
| Load dataset from anywhere | DatasetFactory.open("path", target="col") |
| Auto EDA profile | dataset.show_in_notebook() |
| Train-test split | train, test = dataset.train_test_split(test_size=0.2) |
| Run AutoML | AutoMLOperator().with_estimator(OracleAutoMLProvider()).with_dataset(train).train() |
| Evaluate model | ADSEvaluator(test, models=[model]).show_in_notebook() |
| Explain predictions (SHAP) | ADSExplainer(model, test).global_explanation.show_in_notebook() |
| Save to Model Catalog | GenericModel(estimator).prepare(...).save(display_name="...") |
| Deploy as REST API | ModelDeployment().deploy(properties=...) |
| Create ML Pipeline | Pipeline("name").with_step_details([...]).create() |
| Create a scheduled Job | DataScienceJob().with_runtime(ScriptRuntime()).create() |
| Call OCI GenAI LLM | GenerativeAI(compartment_id=...).generate(prompt=...) |
🏆 Section 14: Best Practices — The Golden Rules of ADS
- ✅ Always use
resource_principalauth inside OCI notebooks - ✅ Call
show_in_notebook()before ANY model training — understand your data first! - ✅ Let AutoML run first. Use its best model as your baseline.
- ✅ Add meaningful metadata to every model saved in Model Catalog
- ✅ Always evaluate with ADSEvaluator before deploying
- ✅ Use SHAP explainability for any model used in customer-facing decisions
- ✅ Use ML Pipelines for production — not manual notebook runs
- ✅ Store reusable features in Feature Store for team-wide sharing
- ✅ Set model version names clearly:
ModelName_v2_YYYY-MM - ✅ Monitor deployed endpoints with OCI Monitoring + set accuracy alerts
- ❌ Don't hardcode credentials anywhere in your code
- ❌ Don't skip data profiling and go straight to model training
- ❌ Don't deploy a model without evaluating it on a proper test set
- ❌ Don't ignore model explainability for regulated use cases
- ❌ Don't run production retraining manually from a notebook — use Pipelines/Jobs
- ❌ Don't save models without version numbers — you'll lose track fast
- ❌ Don't forget to delete unused deployments — they cost money even when idle!
- ❌ Don't use large compute shapes during development — use Flex shapes wisely
🗺️ Section 15: Full ADS Architecture — The Complete Picture
Every row maps ADS code to the underlying OCI service it uses.
| # | Stage | ADS Tool | OCI Service Behind It |
|---|---|---|---|
| 1 | 📂 Data Storage | DatasetFactory.open("oci://...") |
OCI Object Storage |
| 2 | 🔍 EDA | dataset.show_in_notebook() |
OCI Data Science Notebooks |
| 3 | 🤖 Training | AutoMLOperator.train() |
OCI Data Science (GPU/CPU compute) |
| 4 | 📊 Evaluation | ADSEvaluator + ADSExplainer |
OCI Data Science Notebooks |
| 5 | 🏛️ Storage | GenericModel.save() |
OCI Model Catalog |
| 6 | 🚀 Deployment | ModelDeployment.deploy() |
OCI Model Deployment (REST API) |
| 7 | 🔄 Automation | Pipeline + DataScienceJob |
OCI Data Science Jobs + Pipelines |
| 8 | 🧪 Features | FeatureStore + FeatureGroup |
OCI Feature Store |
| 9 | 🤖 GenAI | GenerativeAI.generate() |
OCI Generative AI Service |
📝 Quick Summary — Everything You Learned Today
- ADS = Oracle's open-source Python library that makes ML on OCI dramatically faster and easier
- Authentication = Use
resource_principalinside OCI notebooks — most secure! - DatasetFactory = Load data from Object Storage, databases, or local files in 2 lines
- AutoML = Automatically tries many algorithms and picks the best one for you
- ADSEvaluator = Full model evaluation with one method call — metrics + plots
- ADSExplainer = SHAP-powered explanations for why the model made each decision
- GenericModel = Package and save any trained model to OCI Model Catalog
- ModelDeployment = Deploy any model as a live REST API with one command
- Pipeline = Chain steps into a production-grade automated ML workflow
- Feature Store = Shared repository for reusable, consistent ML features
- GenerativeAI = Call OCI LLMs directly from Python — combine ML + GenAI!
Comments
Post a Comment