Skip to main content

OCI Data Science Notebook Sessions: Build, Train, and Run Machine Learning Models

Calculating read time…

Imagine you are a scientist who needs a fully equipped laboratory to run experiments. You need powerful computers, all the right tools pre-installed, a clean workspace, and the ability to share your work with your team. But setting all of that up from scratch takes weeks!

OCI Data Science Notebook Sessions are exactly that — a fully equipped AI and Data Science laboratory, ready in the cloud in under 5 minutes. No setup. No server management. Just open your browser and start building AI models!



What Is OCI Data Science?

OCI Data Science is Oracle's fully managed platform for building, training, and deploying machine learning models. "Fully managed" means Oracle handles all the infrastructure — the servers, the networking, the storage, the software updates. You focus entirely on your data and your models. Oracle handles everything else!

OCI Data Science organises your work around two key concepts — Projects and Notebook Sessions:

  • 📁 Project — Think of a Project like a folder for a specific piece of work. For example: "Customer Churn Prediction" or "NLP Sentiment Analysis". A project groups together all the notebook sessions and models related to that piece of work. Your whole team can share access to a project.
  • 📓 Notebook Session — This is the actual working environment inside a project. It's a full JupyterLab interface running on real compute power (CPU or GPU) in the cloud. Think of it like opening a new lab bench to run your experiments on.

What Is a Notebook Session? 🧪

A Notebook Session is an interactive coding environment running in OCI's cloud. It's powered by JupyterLab — the most popular data science IDE (Integrated Development Environment) in the world.

💡 What is JupyterLab? Think of JupyterLab like a super-smart digital notebook that can also run code. You write text to explain your thinking, then write Python code, run it, and see the results (charts, tables, numbers) all in the same page. Scientists love it because you can tell a story with your data and show your work at the same time — like a science lab report that runs code!

WHAT YOU GET IN EVERY NOTEBOOK SESSION:
═════════════════════════════════════════════════════════════

📓 JupyterLab Interface
   └── Write Python, run code, see charts — all in the browser

💻 Compute Power
   ├── CPU shapes: Standard VMs (great for data prep, light training)
   └── GPU shapes: NVIDIA A10, A100, H100, H200 (for deep learning!)

📦 Pre-installed Libraries (via Conda Environments)
   ├── Data Science: Pandas, NumPy, Scikit-learn, Matplotlib
   ├── Deep Learning: TensorFlow, PyTorch, Keras
   ├── NLP: Hugging Face Transformers, NLTK, SpaCy
   └── OCI Tools: ADS SDK, oci-python-sdk, ocifs

💾 Persistent Storage
   └── Your notebooks and files are saved to block storage
       (don't disappear when you stop the session!)

🔗 OCI Integration
   ├── Read/write OCI Object Storage directly
   ├── Connect to Autonomous Database, Data Flow, Vault
   └── Call OCI AI Services (Vision, Language, Generative AI)

🤖 AI Quick Actions (2025 feature!)
   └── Deploy and fine-tune Llama, Mistral, GPT-OSS — no code needed!

The Notebook Session Lifecycle — Start, Pause, Resume 🔄

Understanding the lifecycle is important — especially for cost management! A notebook session can be in one of these states:

  • 🟢 Active — The compute is running, JupyterLab is open, you can code. This is when you're being billed for the compute (CPU or GPU).
  • 🔴 Inactive (Deactivated) — The compute is stopped. Your notebooks and files are still saved in block storage, but no compute is running. You're only paying for the storage (very cheap!), NOT the compute. You can reactivate anytime — even change the compute shape when you do!
  • 🗑️ Deleted — Everything is gone. The compute AND the storage are released. Only do this when you're completely done with a project.
✅ Money-saving tip: When you're done working for the day, always Deactivate your notebook session — NOT Delete. Deactivating stops the compute billing while keeping all your saved notebooks and files. You can resume the next morning exactly where you left off! Only pay for GPU when your GPU is actually running code.
❌ DON'T: Never leave a GPU notebook session running overnight when you're not using it. NVIDIA H100 GPUs cost a significant amount per hour. A single forgotten active GPU session can result in hundreds of dollars of unexpected charges! Always deactivate when you're stepping away.
NOTEBOOK SESSION LIFECYCLE:

          CREATE
            │
            ▼
    ┌─── ACTIVE ───┐     ← You are working, compute is running, you are billed
    │  JupyterLab  │
    │  is open     │
    └──────┬───────┘
           │ Deactivate
           ▼
    ┌─ INACTIVE ───┐     ← Compute stopped, files saved, only storage billed
    │  Files safe  │
    │  Compute off │
    └──────┬───────┘
           │ Activate (can change compute shape here!)
           ▼
    ┌─── ACTIVE ───┐     ← Back to work!
           │
           │ Delete (WARNING: files gone forever!)
           ▼
        DELETED

Compute Shapes — CPU vs GPU, Which Do You Need? 💻

When creating a notebook session, you choose a compute shape — the type and amount of computing power. This is like choosing between a regular car (CPU) and a Formula 1 race car (GPU) for your journey.

CPU Shapes — The Reliable Everyday Car 🚗

  • Good for: Data exploration, data cleaning, feature engineering, training simple models (linear regression, decision trees, random forests)
  • Example shapes: VM.Standard.E4.Flex (AMD), VM.Standard3.Flex (Intel), VM.Standard.A1.Flex (ARM)
  • Flexible CPU and RAM: You choose how many CPUs (1 to 64) and how much RAM (1GB to 1TB)
  • Cost: Much cheaper than GPU. Great for most daily data science work

GPU Shapes — The Formula 1 Race Car 🏎️

  • Good for: Training neural networks, fine-tuning large language models, computer vision, any deep learning task
  • Available GPU shapes in OCI Data Science:
    • A10 (24GB VRAM) — Great for small-medium LLM fine-tuning, inference
    • A100 (40GB or 80GB VRAM) — Industry standard for LLM training
    • H100 (80GB VRAM) — State-of-the-art, fastest available GPU
    • H200 (141GB VRAM) — Latest generation, massive memory for huge models
  • Cost: Significantly higher than CPU. Use only when your workload genuinely needs a GPU
💡 GPU limits are ZERO by default! When you first create an OCI account, GPU shape limits are set to zero for all GPU types. Before you can use a GPU, you must submit a Service Limit Increase request through OCI Console → Governance & Administration → Limits, Quotas and Usage. This is to prevent accidental high GPU charges. CPU shapes are available immediately with no limit requests needed!

Step-by-Step: Create Your First Notebook Session 🛠️

Prerequisites: Set Up IAM Policies First

📌 What IAM Policies do:
OCI requires an administrator to explicitly grant permission for users to use each service. Before anyone can create a notebook session, an admin must add Policy statements. Think of it like a school requiring a permission slip from the principal before students can use the science lab!

An administrator must add these policies in OCI Console → Identity → Policies:

# Allow your user group to use Data Science service
allow group <your-group> to manage data-science-family in tenancy

# Allow Data Science to access Object Storage (to read/write data and conda envs)
allow service datascience to use object-family in tenancy

# Allow Data Science to use VCN networking
allow service datascience to use virtual-network-family in tenancy

# Allow your notebooks to call other OCI services (Resource Principal)
allow dynamic-group <notebook-dynamic-group> to use generative-ai-family in tenancy
allow dynamic-group <notebook-dynamic-group> to read secret-family in tenancy
allow dynamic-group <notebook-dynamic-group> to manage object-family in tenancy

Create the Notebook Session — Click by Click

  • Step 1: OCI Console → hamburger menu (☰ top-left) → Analytics & AI → Data Science
  • Step 2: Click Create Project → give it a name like my-first-ds-project → click Create
  • Step 3: Inside the project, click Create Notebook Session
  • Step 4: Configure the notebook session:
    • Name: e.g. exploration-notebook
    • Compute Shape: For beginners, choose VM.Standard.E4.Flex with 4 OCPUs and 32 GB RAM
    • Block Storage: 100 GB is a good starting point (where your notebooks are saved)
    • Networking: Choose "Managed" to let OCI configure networking automatically — easiest option!
  • Step 5: Click Create — the session will be in "Creating" state for 2-5 minutes
  • Step 6: When status changes to Active, click Open — JupyterLab opens in your browser! 🎉
✅ Managed Networking is your best friend as a beginner! When you select "Managed" networking, OCI automatically creates all the required VCN, subnet, and internet gateway settings for you. You don't need to know anything about networking to get started. Save the custom VCN setup for later when you need private access or custom security configurations.

Inside JupyterLab — Your Data Science Cockpit 🖥️

When you open a notebook session, you see JupyterLab — a browser-based development environment. Let's understand the key parts of the interface:

JUPYTERLAB INTERFACE OVERVIEW:

┌──────────────────────────────────────────────────────────────────────────┐
│  MENU BAR: File | Edit | View | Run | Kernel | Tabs | Settings | Help   │
├──────────────┬───────────────────────────────────────────────────────────┤
│              │                                                           │
│  LEFT PANEL  │              MAIN WORK AREA                               │
│  (File       │                                                           │
│   Browser)   │  ┌────────────────────────────────────────────────┐       │
│              │  │  notebook.ipynb                                │       │
│  📁 Folders  │  │  ┌────────────────────────────────────────┐   │       │
│  📓 Files    │  │  │ [Markdown Cell] — write explanations   │   │       │
│              │  │  └────────────────────────────────────────┘   │       │
│  [Launcher]  │  │  ┌────────────────────────────────────────┐   │       │
│              │  │  │ [Code Cell] import pandas as pd        │   │       │
│  Environment │  │  │              df = pd.read_csv(...)     │   │       │
│  Explorer 🔍 │  │  │ ▶ RUN  [Output appears here]           │   │       │
│              │  │  └────────────────────────────────────────┘   │       │
│  AI Quick    │  └────────────────────────────────────────────────┘       │
│  Actions 🤖  │                                                           │
│              │  TERMINAL (bash shell — for conda, pip, git commands)     │
└──────────────┴───────────────────────────────────────────────────────────┘

The most important parts for you as a beginner:

  • 📁 File Browser (left panel) — Navigate your notebook files and folders
  • 📓 Notebook (.ipynb files) — Your main working area with code + markdown cells
  • 💻 Terminal — A command line where you run conda and pip commands to install packages
  • 🔍 Environment Explorer — Browse and install pre-built Conda environments with one click
  • 🤖 AI Quick Actions — Deploy and fine-tune LLMs without writing any code!

Conda Environments — Your Tool Boxes 📦

What Is a Conda Environment?

Imagine you are a chef with two kitchens. One kitchen is for baking (with ovens, mixers, flour). The other is for sushi (with knives, bamboo mats, fresh fish). The tools are completely separate — the baking kitchen doesn't have sushi tools and vice versa.

That's exactly what a Conda environment is — a completely separate, self-contained collection of Python packages. You can have one environment for TensorFlow projects and another for PyTorch projects — and they never conflict with each other!

OCI Data Science provides pre-built, managed Conda environments maintained by Oracle's Data Science team. These are carefully tested environments with all the right packages already installed. You just install one in your notebook session and start coding!

Key Pre-Built Conda Environments in OCI

  • 🤖 General Machine Learning (generalml)
    The go-to environment for most data science tasks. Includes: Scikit-learn, XGBoost, LightGBM, Pandas, NumPy, Matplotlib, Seaborn, ADS SDK. Perfect for: classification, regression, clustering, time series analysis.
  • 🧠 TensorFlow
    GPU-optimised environment for TensorFlow deep learning. Includes: TensorFlow 2.x, Keras, TensorBoard, ADS SDK. Perfect for: neural networks, image classification, sequence modelling.
  • 🔥 PyTorch
    GPU-optimised environment for PyTorch deep learning. Includes: PyTorch, TorchVision, TorchAudio, Hugging Face Transformers, ADS SDK. Perfect for: LLM fine-tuning, computer vision, custom neural architectures.
  • 💬 NLP and Text
    Environment focused on natural language processing. Includes: Hugging Face Transformers, NLTK, SpaCy, Gensim, ADS SDK. Perfect for: text classification, NER, sentiment analysis, text generation.
  • 💰 Financial Services
    Specialised environment for quantitative finance. Includes: QuantLib, TA-Lib (technical analysis), portfolio optimisation tools. Perfect for: stock analysis, risk modelling, algorithmic trading signals.
  • ⚡ PySpark
    For big data processing at scale. Includes: PySpark, Delta Lake, integration with OCI Data Flow. Perfect for: processing datasets larger than your RAM, distributed computing.

How to Install a Conda Environment in Your Notebook Session

📌 What the commands below do:
These terminal commands install a pre-built OCI Conda environment into your notebook session. Open a Terminal in JupyterLab (Launcher → Terminal) and run these commands. The odsc tool is the OCI Data Science CLI — it manages conda environments and more. Installing a conda environment can take 5-15 minutes. You only need to do this once per notebook session!
# List all available pre-built OCI conda environments
odsc conda list

# Install the General Machine Learning environment (CPU/Python 3.11 version)
odsc conda install -s generalml_p311_cpu_x86_64_v1

# Install the PyTorch environment for GPU workloads
odsc conda install -s pytorch21_p39_gpu_v1

# Check which environments are installed in this session
odsc conda list --installed

After installation, you'll see the new environment in the JupyterLab Kernel dropdown — just select it when creating or switching notebook kernels!

Your First Notebook — Loading Data from Object Storage 📊

What Is Oracle ADS (Accelerated Data Science) SDK?

The ADS (Accelerated Data Science) SDK is Oracle's own Python library that makes working in OCI Data Science much easier. Think of it as a collection of shortcuts — instead of writing 30 lines of OCI API code to read a file from Object Storage, you write 2 lines using ADS! ADS comes pre-installed in all OCI-managed Conda environments.

Authentication — Resource Principal vs API Key

Before your notebook code can access OCI services (Object Storage, Generative AI, Vault, etc.), it needs to prove it's authorised. There are two ways to do this inside a notebook session:

  • 🔑 Resource Principal (Recommended) — Your notebook session itself becomes an authorised "identity" in OCI. No keys or passwords to manage — the session automatically gets temporary credentials that rotate every 15 minutes. This is the modern, secure approach used in production!
  • 🗝️ API Key — You manually create an API key in OCI Console, download the private key file, and put it in a config file. Simpler to understand at first, but less secure for production. Good for local development outside a notebook session.
📌 What the code below does — in plain English:
This code authenticates your notebook session with OCI using Resource Principal (the recommended method). After setting authentication, it reads a CSV file directly from OCI Object Storage into a Pandas DataFrame — just like reading a local file, but the file is in the cloud! The oci:// path format is just like s3:// for Amazon but for Oracle Object Storage.
import ads
import pandas as pd

# ── STEP 1: Authenticate using Resource Principal ──────────────────────────────
# This tells ADS to use the notebook session's own OCI identity (no passwords needed!)
# The notebook session must be in a Dynamic Group with the right IAM policies.
ads.set_auth(auth="resource_principal")

print("✅ Authenticated with OCI using Resource Principal!")

# ── STEP 2: Read a CSV directly from OCI Object Storage ────────────────────────
# Format: oci://{bucket-name}@{namespace}/{file-path}
# You can find your namespace in OCI Console → Object Storage → Namespace
NAMESPACE   = "your-tenancy-namespace"     # e.g. "abcde12345"
BUCKET_NAME = "my-data-bucket"
FILE_PATH   = "sales/q3_2026_sales.csv"

print(f"📂 Reading data from: oci://{BUCKET_NAME}@{NAMESPACE}/{FILE_PATH}")

df = pd.read_csv(
    f"oci://{BUCKET_NAME}@{NAMESPACE}/{FILE_PATH}",
    storage_options=ads.common.auth.default_signer()  # Passes OCI auth to the reader
)

print(f"\n✅ Data loaded successfully!")
print(f"   Shape: {df.shape[0]:,} rows × {df.shape[1]} columns")
print(f"\nFirst 5 rows:")
print(df.head())
print(f"\nColumn names: {list(df.columns)}")
print(f"\nData types:\n{df.dtypes}")

Example output:

✅ Authenticated with OCI using Resource Principal!
📂 Reading data from: oci://my-data-bucket@abcde12345/sales/q3_2026_sales.csv

✅ Data loaded successfully!
   Shape: 45,231 rows × 8 columns

First 5 rows:
   order_id  customer_id     product     region  amount  discount  date
0     10001        C4521  Laptop Pro     EMEA    1299.0       0.0   2026-07-01
1     10002        C7832  Phone X12     APAC     599.0      10.0   2026-07-01
2     10003        C2201  Headphones    AMER     149.0       0.0   2026-07-02
3     10004        C9910  Smart Watch   EMEA     349.0       5.0   2026-07-02
4     10005        C3341  Laptop Pro    APAC    1299.0      15.0   2026-07-03

45,000+ rows of data loaded from the cloud into Pandas in seconds! And you didn't need to download anything to your local computer. 

Training a Machine Learning Model in Your Notebook 

📌 What the code below does — in plain English:
This is a complete, beginner-friendly machine learning pipeline:
  1. Uses the sales data we loaded from Object Storage
  2. Prepares the data (handles missing values, encodes text columns)
  3. Splits the data into training and testing portions
  4. Trains a Random Forest model to predict whether a sale has a discount or not
  5. Evaluates how accurate the model is on data it has never seen before
Think of it like teaching a dog tricks: you show it examples (training), then test if it learned by trying new examples (testing)!
import pandas as pd
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder
from sklearn.metrics import accuracy_score, classification_report
import matplotlib.pyplot as plt
import warnings
warnings.filterwarnings('ignore')

# ── STEP 1: Prepare features for training ────────────────────────────────────
# We'll predict: did this customer get a discount? (Yes = 1, No = 0)

# Create our target variable: 1 if discount > 0, else 0
df["got_discount"] = (df["discount"] > 0).astype(int)

# Encode text column 'region' to numbers (ML needs numbers, not text)
# LabelEncoder converts: EMEA→0, APAC→1, AMER→2
le_region = LabelEncoder()
df["region_encoded"] = le_region.fit_transform(df["region"])

# Encode text column 'product' to numbers
le_product = LabelEncoder()
df["product_encoded"] = le_product.fit_transform(df["product"])

# Extract useful features from the date
df["date"] = pd.to_datetime(df["date"])
df["day_of_week"] = df["date"].dt.dayofweek   # 0=Monday, 6=Sunday
df["month"]       = df["date"].dt.month

# ── STEP 2: Select feature columns (inputs) and target (what we want to predict)
features = ["region_encoded", "product_encoded", "amount", "day_of_week", "month"]
target   = "got_discount"

X = df[features]
y = df[target]

print(f"📊 Dataset prepared:")
print(f"   Features: {features}")
print(f"   Total samples: {len(X):,}")
print(f"   Discount rate: {y.mean()*100:.1f}% of orders had a discount\n")

# ── STEP 3: Split data — 80% for training, 20% for testing ───────────────────
# We train on 80% and test on the remaining 20% (data the model has never seen)
X_train, X_test, y_train, y_test = train_test_split(
    X, y,
    test_size=0.2,          # 20% goes to test set
    random_state=42,        # Fixed seed = reproducible results
    stratify=y              # Keep same discount ratio in both splits
)

print(f"   Training samples: {len(X_train):,}")
print(f"   Testing samples:  {len(X_test):,}\n")

# ── STEP 4: Train the Random Forest model ────────────────────────────────────
# Random Forest = an ensemble of 100 decision trees working together
# Like asking 100 experts to vote — better than asking just one!
print("🌲 Training Random Forest model...")
model = RandomForestClassifier(
    n_estimators=100,   # Build 100 decision trees
    max_depth=8,        # Each tree can be at most 8 levels deep
    random_state=42
)
model.fit(X_train, y_train)
print("✅ Training complete!\n")

# ── STEP 5: Evaluate the model on the test set ───────────────────────────────
y_pred  = model.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)

print(f"📊 MODEL PERFORMANCE:")
print(f"   Accuracy: {accuracy*100:.1f}%")
print(f"\n   Detailed report:")
print(classification_report(y_test, y_pred,
                             target_names=["No Discount", "Got Discount"]))

# ── STEP 6: Show which features matter most ──────────────────────────────────
importance_df = pd.DataFrame({
    "Feature":    features,
    "Importance": model.feature_importances_
}).sort_values("Importance", ascending=False)

print("🔍 Feature Importance (what the model relies on most):")
for _, row in importance_df.iterrows():
    bar = "▓" * int(row["Importance"] * 50)
    print(f"   {row['Feature']:<20} {bar} {row['Importance']:.3f}")

Example output:

📊 Dataset prepared:
   Features: ['region_encoded', 'product_encoded', 'amount', 'day_of_week', 'month']
   Total samples: 45,231
   Discount rate: 23.7% of orders had a discount

   Training samples: 36,184
   Testing samples:  9,047

🌲 Training Random Forest model...
✅ Training complete!

📊 MODEL PERFORMANCE:
   Accuracy: 87.4%

   Detailed report:
                 precision    recall  f1-score   support
   No Discount       0.91      0.93      0.92      6,903
   Got Discount       0.76      0.71      0.73      2,144

🔍 Feature Importance (what the model relies on most):
   product_encoded     ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░░░░░░░░░  0.341
   amount              ▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░░░░░░░░░░░░  0.289
   region_encoded      ▓▓▓▓▓▓▓▓░░░░░░░░░░░░░░░░░░░░░  0.178
   month               ▓▓▓▓░░░░░░░░░░░░░░░░░░░░░░░░░  0.102
   day_of_week         ▓▓▓░░░░░░░░░░░░░░░░░░░░░░░░░░  0.090

87.4% accuracy — a solid model trained entirely in the cloud in your OCI Notebook Session! The "product type" and "order amount" are the most important factors in predicting whether a discount was applied. 📊

Feature: AI Quick Actions — Deploy LLMs Without Code! 

What Are AI Quick Actions?

AI Quick Actions (AQUA) is one of the most exciting features added to OCI Data Science in 2025. It gives you a no-code interface inside JupyterLab to:

  • 🔍 Browse a curated catalogue of state-of-the-art open-source models
  • 🚀 Deploy any model as a live HTTP endpoint — click, wait, done
  • ⚙️ Fine-tune models on your own data — no PyTorch code required
  • 📊 Evaluate fine-tuned models and compare performance
  • 🔗 Query deployed models via API in your own code

Models available include: Meta Llama 4, Mistral, IBM Granite, OpenAI gpt-oss-120b, and dozens more. GPU shapes available: A10, A100, H100, H200.

✅ You only need a CPU notebook to ACCESS AI Quick Actions! Create a cheap CPU notebook session to open the AI Quick Actions interface. When you click "Deploy" or "Fine-tune", AI Quick Actions automatically provisions the required GPU compute separately — you're not paying GPU prices for the small notebook you use to access the interface. Smart architecture saves you money! 💰

How to Access AI Quick Actions

  • Step 1: Open any active Notebook Session in JupyterLab
  • Step 2: In the left sidebar, click the AI Quick Actions icon (rocketship 🚀)
  • Step 3: Click Model Explorer — browse all available foundation models
  • Step 4: Find a model you want (e.g., Llama 4 Scout) → click Deploy
  • Step 5: Choose a compute shape (A10, A100, etc.), set logging options → click Deploy
  • Step 6: Wait for deployment to complete (5-20 minutes) → get your HTTPS endpoint URL
  • Step 7: Query your deployed model from any application! 🎉
📌 What the code below does — in plain English:
After deploying a model via AI Quick Actions, you get an endpoint URL. This code sends a chat-style question to your deployed model and prints the response. It uses the OpenAI-compatible API format — so any code written for OpenAI's GPT models works with your OCI-deployed model with just a URL change!
import ads
import ads.aqua
from openai import OpenAI

# ── STEP 1: Authenticate with OCI ────────────────────────────────────────────
# Using resource principal since we're inside a notebook session
ads.set_auth(auth="resource_principal")

# ── STEP 2: Your deployed model's endpoint URL ────────────────────────────────
# Find this in: OCI Console → Data Science → Model Deployments → Your Deployment
ENDPOINT = "https://modeldeployment.us-ashburn-1.oci.customer-oci.com/ocid1.datasciencemodeldeployment.oc1..."

# ── STEP 3: Create an OpenAI-compatible client pointing to your OCI model ─────
# ads.aqua.get_httpx_client() handles OCI authentication for the OpenAI client
client = OpenAI(
    api_key="OCI",                         # Required field but OCI auth handles the actual security
    base_url=f"{ENDPOINT}/predict/v1/",    # Your deployed model's endpoint
    http_client=ads.aqua.get_httpx_client()  # OCI-authenticated HTTP client
)

# ── STEP 4: Send a question to your deployed model ────────────────────────────
print(" Querying my deployed Llama 4 Scout model via AI Quick Actions...")
print("-" * 60)

response = client.chat.completions.create(
    model="odsc-llm",    # Always use "odsc-llm" for single-model deployments
    messages=[
        {
            "role": "system",
            "content": "You are a helpful data science assistant. Be concise and clear."
        },
        {
            "role": "user",
            "content": "What are the top 3 signs that a machine learning model is overfitting?"
        }
    ],
    max_tokens=400,
    temperature=0.5
)

print(response.choices[0].message.content)
print(f"\n📊 Tokens used — Prompt: {response.usage.prompt_tokens}, Response: {response.usage.completion_tokens}")

Example output:

Querying my deployed Llama 4 Scout model via AI Quick Actions...
------------------------------------------------------------
The top 3 signs your machine learning model is overfitting:

1. **Large gap between training and validation accuracy**
   Your model scores 97% on training data but only 71% on validation data.
   The gap shows the model memorized training examples instead of learning general patterns.

2. **Performance drops significantly on new, unseen data**
   The model performs brilliantly in the notebook but poorly when tested on real-world data it has never encountered.

3. **Overly complex model for the dataset size**
   A 500-node neural network trained on 200 examples is almost guaranteed to overfit.
   A good rule: more complex models need much more training data to generalise well.

📊 Tokens used — Prompt: 67, Response: 142

Custom Conda Environments — Saving and Sharing Your Setup 🔧

Sometimes the pre-built OCI conda environments don't have a specific library you need. No problem — you can install additional libraries and save your customised environment to OCI Object Storage. Then your teammates can install the exact same environment with one command!

📌 What the commands below do — in plain English:
These terminal commands show how to: install extra libraries into an existing conda environment, save your customised environment to Object Storage so it's not lost when the session is deactivated, and then reinstall it from Object Storage in any other notebook session. Think of it like saving your custom recipe — anyone with access to the storage bucket can recreate your exact cooking environment!
# ── Inside JupyterLab Terminal ─────────────────────────────────────────────

# Step 1: First install the pre-built general ML conda environment
odsc conda install -s generalml_p311_cpu_x86_64_v1

# Step 2: Activate that conda environment to work inside it
conda activate /home/datascience/conda/generalml_p311_cpu_x86_64_v1

# Step 3: Install additional libraries that aren't in the pre-built env
pip install langchain-oci
pip install prophet         # Facebook's time series forecasting library
pip install shap            # Model explainability library
pip install plotly          # Interactive charts

# Step 4: Publish your customised environment to Object Storage
# This saves the ENTIRE environment so anyone can reinstall it exactly
odsc conda publish \
    -s my-custom-ml-env \                   # Give your custom env a name
    --bucket my-conda-bucket \              # Your Object Storage bucket
    --prefix conda-envs/                    # Folder inside the bucket
    # This creates: oci://my-conda-bucket@namespace/conda-envs/my-custom-ml-env_v1.tar.gz

echo "✅ Custom conda environment saved to Object Storage!"

# Step 5: To install this custom env in any future notebook session, use:
# odsc conda install -e oci://my-conda-bucket@namespace/conda-envs/my-custom-ml-env_v1.tar.gz
✅ Best Practice for Teams: Create ONE shared Object Storage bucket for your team's conda environments. When you create a custom environment, publish it there. All team members can install the exact same environment — guaranteed consistency! This eliminates the classic "but it works on my machine!" problem.
❌ DON'T: Don't install libraries using pip install directly into the base system Python (outside of a conda environment). Changes to the base environment are not persistent across session deactivations. Always activate a conda environment first, then install into it. And always publish the customised conda environment to Object Storage before deactivating the session — otherwise your custom installs are lost!

Notebook Session Lifecycle Scripts — Automate Your Setup 🔄

Every time you create or activate a notebook session, you might need to run the same setup steps: install a conda environment, clone a git repo, set environment variables. Doing this manually every time is tedious!

Lifecycle Scripts solve this — they are bash scripts that run automatically at specific points in the notebook session lifecycle: CREATE, ACTIVATE, DEACTIVATE, and DELETE.

📌 What the script below does — in plain English:
This is an ACTIVATE lifecycle script — it runs automatically every time you activate (start) a notebook session. It does four things: installs your team's custom conda environment, clones your project's Git repository, sets an environment variable (like saving a setting that all your notebooks can read), and logs a message. Think of it like a robot assistant that automatically sets up your workspace every morning before you arrive!
#!/bin/bash
# ACTIVATE lifecycle script — runs every time the notebook session starts
# Save this as activate.sh and provide it when creating/editing the notebook session

set -e   # Exit immediately if any command fails
echo "🚀 Starting notebook session setup (ACTIVATE lifecycle script)..."

# ── Install the team's custom conda environment from Object Storage ───────────
echo "📦 Installing team conda environment..."
odsc conda install \
    -e "oci://team-bucket@abcde12345/conda-envs/team_ml_env_v3.tar.gz" \
    --overwrite false   # Don't reinstall if already exists (saves time!)

echo "✅ Conda environment ready!"

# ── Clone the project Git repository ─────────────────────────────────────────
PROJECT_DIR="/home/datascience/projects/sales-prediction"
if [ ! -d "$PROJECT_DIR" ]; then
    echo "📥 Cloning project repository..."
    git clone https://github.com/your-org/sales-prediction.git "$PROJECT_DIR"
    echo "✅ Repository cloned!"
else
    echo "🔄 Pulling latest changes from repository..."
    cd "$PROJECT_DIR" && git pull
    echo "✅ Repository updated!"
fi

# ── Set environment variables ─────────────────────────────────────────────────
echo "⚙️ Setting environment variables..."
echo "export PROJECT_BUCKET=my-project-bucket" >> ~/.bashrc
echo "export PROJECT_NAMESPACE=abcde12345" >> ~/.bashrc
echo "export MLFLOW_TRACKING_URI=http://my-mlflow-server:5000" >> ~/.bashrc

echo ""
echo "🎉 Notebook session is ready! Happy coding!"

To attach a lifecycle script to your notebook session: In OCI Console → When creating or editing a notebook session → look for Advanced Options → Lifecycle Scripts → upload your script or paste its Object Storage path.

Saving Your Model to the OCI Model Catalog 🗃️

Once you've trained a great model in your notebook, you want to save it somewhere safe and organised. The OCI Model Catalog is a managed repository specifically for ML models. Think of it like a library for your trained models — versioned, searchable, and ready to be deployed!

📌 What the code below does — in plain English:
After training the Random Forest model earlier, this code packages it up and saves it to the OCI Model Catalog. It creates a "model artifact" (think: a ZIP file with your model and all information needed to run it) and uploads it to OCI. Once saved to the catalog, the model can be deployed as a live REST API endpoint with one more command — or deployed by your DevOps team without them needing to understand how the model was trained. It's like putting your finished work in a proper filing system instead of leaving it on your desk!
import ads
from ads.model import SklearnModel
from ads.common.auth import default_signer

# ── Set authentication ────────────────────────────────────────────────────────
ads.set_auth(auth="resource_principal")

# ── Wrap the trained model with ADS framework ────────────────────────────────
# ADS provides wrappers for all common ML frameworks: Sklearn, TensorFlow, PyTorch, etc.
# This wrapper helps prepare everything needed to save and deploy the model.
ads_model = SklearnModel(
    estimator=model,            # Your trained RandomForestClassifier
    kind="model",               # Type of ADS model object
    artifact_dir="./model_artifact"   # Local folder to prepare the artifact
)

# ── Prepare the model artifact ───────────────────────────────────────────────
# This creates a score.py file (for predictions) and a runtime.yaml (environment info)
# These are needed for deployment. ADS generates them automatically!
ads_model.prepare(
    inference_conda_env="oci://my-conda-bucket@namespace/conda-envs/team_ml_env_v3.tar.gz",
    # ↑ Tell OCI which conda env to use when serving predictions
    X_sample=X_test.head(5),    # Provide sample input so ADS can infer the schema
    force_overwrite=True
)

# ── Verify the artifact looks correct ────────────────────────────────────────
print("📦 Model artifact contents:")
ads_model.introspect()      # Checks the artifact for common issues

# ── Save to the OCI Model Catalog ────────────────────────────────────────────
COMPARTMENT_ID = "ocid1.compartment.oc1..aaaaaaaa...your-compartment-id..."
PROJECT_ID     = "ocid1.datascienceproject.oc1..aaaaaaaa...your-project-id..."

catalog_entry = ads_model.save(
    display_name="sales_discount_predictor_v1",
    description=(
        "Random Forest classifier predicting whether a sales order received a discount. "
        "Trained on Q3 2026 sales data. Accuracy: 87.4%. "
        "Features: product, region, amount, day_of_week, month."
    ),
    project_id=PROJECT_ID,
    compartment_id=COMPARTMENT_ID,
    ignore_pending_changes=True
)

print(f"\n✅ Model saved to OCI Model Catalog!")
print(f"   Model OCID: {catalog_entry.id}")
print(f"   Display Name: {catalog_entry.display_name}")
print(f"   Status: {catalog_entry.lifecycle_state}")
print(f"\n💡 You can now deploy this model from OCI Console →")
print(f"   Data Science → Projects → Your Project → Models → {catalog_entry.display_name}")

Example output:

📦 Model artifact contents:
   ✅ score.py         — prediction function found
   ✅ runtime.yaml     — conda environment reference found
   ✅ model.pkl        — model binary found
   ✅ input_schema.json — input schema found

✅ Model saved to OCI Model Catalog!
   Model OCID: ocid1.datasciencemodel.oc1.iad.aaaaaaaa...
   Display Name: sales_discount_predictor_v1
   Status: ACTIVE

💡 You can now deploy this model from OCI Console →
   Data Science → Projects → Your Project → Models → sales_discount_predictor_v1

Common Mistakes and How to Avoid Them 🚨

❌ Mistake 1: Leaving GPU sessions running when not in use
GPU notebooks cost significantly per hour. Always deactivate when you stop working. Set a calendar reminder if needed — "Deactivate GPU session before leaving office!"
❌ Mistake 2: Installing libraries with pip outside a conda environment
Changes to the base Python environment disappear after deactivation. Always activate a conda environment first, then pip install inside it, then publish the environment to Object Storage.
❌ Mistake 3: Storing sensitive credentials directly in notebook cells
Never hardcode passwords, API keys, or secrets in your notebooks. Use OCI Vault to store secrets securely. In notebooks, retrieve them via ADS: from ads.secrets import Secret and call Secret(secret_id="ocid1.vaultsecret...").get()
❌ Mistake 4: Using a GPU notebook session just to access AI Quick Actions
You need a CPU notebook to access the AI Quick Actions interface — NOT a GPU notebook. If you create a GPU notebook for AI Quick Actions access, that GPU counts against your GPU quota and can't be used for actual model deployments or fine-tuning inside AI Quick Actions. Save your GPU quota for actual GPU work!
❌ Mistake 5: Not setting up IAM policies before creating resources
The most common beginner issue: "I created a notebook but can't access Object Storage!" This is almost always because the Resource Principal Dynamic Group policies were not set up. Always configure IAM policies before creating your first notebook session. Talk to your OCI administrator or follow the policy setup documentation first!

Quick Summary 📝

What we learned about OCI Data Science Notebook Sessions:

  • OCI Data Science → Fully managed ML platform. Oracle handles all infrastructure. You focus on data and models.
  • Projects → Folders that organise your notebook sessions and models by use case. Shared across your team.
  • Notebook Sessions → JupyterLab environments running on OCI compute (CPU or GPU). Ready in 5 minutes.
  • Lifecycle States → Active (running, billed), Inactive (stopped, files safe, only storage billed), Deleted (everything gone). Always deactivate, never leave GPU sessions running!
  • Compute Shapes → CPU for data prep and light training. GPU (A10/A100/H100/H200) for deep learning. GPU limits are zero by default — request an increase first!
  • Conda Environments → Pre-built tool boxes with all the right libraries. Install with odsc conda install. Publish custom ones to Object Storage with odsc conda publish.
  • ADS SDK → Oracle's Python library for OCI Data Science. Simplifies reading data, training, evaluation, and saving models.
  • Resource Principal → The recommended authentication method inside notebook sessions. No keys to manage — the session authenticates itself.
  • AI Quick Actions (AQUA) → No-code interface to deploy and fine-tune SOTA LLMs (Llama 4, Mistral, GPT-OSS). Use a CPU notebook to access it!
  • Lifecycle Scripts → Bash scripts that auto-run at CREATE, ACTIVATE, DEACTIVATE, DELETE. Automate your setup so every session starts ready!
  • Model Catalog → Managed repository for saving and versioning trained models. From catalog to live REST API endpoint with one command!

Comments