Imagine you are baking a birthday cake. 🎂 You don't do everything at once — you mix the batter first, then bake it, then add frosting, one step at a time, in the right order!
That's exactly what an OCI Pipeline does — but for machine learning (AI) models. Instead of cake steps, you have AI steps: load data → clean data → train model → save model. And OCI (Oracle Cloud Infrastructure) runs all those steps for you, automatically, on powerful cloud computers!
💡 Think of it like: A smart robot kitchen that follows your recipe, step by step, without you having to stand there pressing buttons!
📋 What You Will Learn
- What is OCI and OCI Data Science?
- What is a Pipeline and why do we need it?
- The key building blocks of a Pipeline (Pipeline, Steps, DAG, Runs)
- What types of Steps exist and what each one does
- How Pipelines are configured — infrastructure, logs, environment variables
- Step-by-step: Create your first Pipeline using the OCI Console (no code needed!)
- Step-by-step: Create a Pipeline using Python + ADS SDK (with full code + explanation)
- How to run, monitor, and read logs of a Pipeline
- How to schedule a Pipeline to run automatically
- Advanced: Parallel steps, BYOC (Docker containers), and Data Flow steps
- Best practices — DOs and DON'Ts every beginner must know
Let's go! 🎯
1. 🏗️ What is OCI? (Plain English)
OCI stands for Oracle Cloud Infrastructure. Think of it as a giant warehouse of incredibly powerful computers that Oracle owns and rents out to you over the internet.
Instead of buying your own expensive server (which can cost thousands of dollars), you just borrow Oracle's computers for as long as you need — paying only for the time you actually use them. The moment you are done, they stop charging you. It's like renting a taxi instead of buying a car!
🧒 Kid-friendly version: Imagine a giant toy library. Instead of buying every toy, you just borrow what you need for the day and return it. That's OCI — but the "toys" are supercomputers!
Why is OCI special for AI?
- 🖥️ It has GPU machines — super fast processors that train AI models 5–10x faster than regular computers
- 🧪 It has OCI Data Science — a full platform just for building and running AI/ML models
- 🔗 It connects easily to Object Storage (to store your data), Autonomous Database (Oracle's smart database), and other services
- 🔒 It is enterprise-grade — your data and models are kept safe and secure
2. 🔬 OCI Data Science — Your AI Lab in the Cloud
OCI Data Science is a fully managed, serverless platform for building, training, and managing machine learning models. "Fully managed" means Oracle handles all the boring server setup work — you just focus on your AI code!
🧒 Think of it like this: OCI Data Science is a fully equipped science lab. All the equipment is already there — microscopes, beakers, computers. You just walk in, do your experiments (AI training), and walk out. No setup needed!
Here are the main parts of OCI Data Science:
- Projects: A folder/container to organise all your work. Like a school binder that holds all your notebooks and assignments together.
- Notebook Sessions: An interactive coding environment (like Jupyter Notebook) where you write and test Python code. Think of it as your coding desk in the cloud lab.
- Jobs: A way to run a single Python script on a cloud machine. Like sending your homework to a superfast printer — it runs without you sitting there watching.
- Pipelines: Chain multiple Jobs together into a workflow that runs automatically, step by step.
- Model Catalog: A library to save, version, and manage your trained AI models. Like a bookshelf for your AI brains.
- Model Deployment: Makes your trained model available as an API — so other apps can send data to it and get predictions back. Like opening a shop where customers can ask your AI questions.
In 2024–2025, Oracle added huge upgrades including AI Quick Actions (deploy large language models like Llama 3 with zero code!), Data Flow steps in Pipelines (run Apache Spark jobs inside your pipeline), and support for NVIDIA A100, H100, and L40S GPUs in all Data Science resources.
3. 🍜 What is a Pipeline and Why Do We Need One?
Before Pipelines existed, a data scientist had to run each AI step manually. They would run "load data" script. Wait. Then run "clean data" script. Wait. Then run "train model" script. And so on. If they forgot a step or ran them in the wrong order — the whole thing broke! 😩
A Pipeline solves this problem. You tell it: "Here are my steps, here is the order, here are the rules." Then you press one button — and OCI runs the whole thing for you, handles failures, logs everything, and even retries if something goes wrong.
🧒 Cake analogy:
- Step 1 → Mix batter (cannot skip!)
- Step 2 → Pour into pan and bake (cannot do before Step 1!)
- Step 3 → Add frosting (cannot do before cake is baked!)
- Step 4 → Decorate and serve! 🎂
An OCI ML Pipeline works the same way — but for AI:
- Step 1 → Load raw data from Object Storage
- Step 2 → Clean and transform the data
- Step 3 → Train the machine learning model (on a GPU if needed!)
- Step 4 → Evaluate the model accuracy
- Step 5 → Save the model to the Model Catalog
Why is a Pipeline better than running scripts manually?
- ✅ Automated: Press one button — everything runs in the right order
- ✅ Repeatable: Run it again next week with new data — same exact result
- ✅ Parallel steps: Steps that don't depend on each other can run at the SAME time (saves hours!)
- ✅ Full logs: Every step's output is recorded so you can debug easily
- ✅ Cost-efficient: Each step gets its own machine that starts and stops automatically — no wasted cloud spend
- ✅ Schedulable: Set it to run every Monday at 8am automatically — zero manual work
4. 🧬 The Key Building Blocks of a Pipeline
Before writing any code or clicking any buttons, you need to understand 6 key terms. Don't worry — we'll explain each one simply!
📦 Pipeline
The Pipeline is the overall recipe — it holds all your steps and their dependency rules. It also stores default settings (like which computer to use and where to send logs) that apply to all steps unless you override them for a specific step.
You can edit a pipeline after creating it — change the name, log settings, environment variables, etc. Once a pipeline has all its steps with code uploaded, it moves to ACTIVE state and is ready to run!
🔩 Pipeline Step
A Pipeline Step is one single task inside the pipeline. It contains the code to run (called the "artifact"), the compute machine to use, log settings, environment variables, and timeout settings.
Each step is completely independent — it runs on its own machine, has its own environment, and even supports different programming languages (Python, Bash, Java are all supported)! This means Step 2 can use TensorFlow while Step 3 uses PyTorch — no conflicts!
📄 Step Artifact
The artifact is the actual code that runs inside the step.
It must be a single file — either a .py script, a .sh bash script, a .java file, or a .zip archive (which can contain multiple files, but you tell OCI which file is the starting point).
All script-type steps MUST have an artifact uploaded before the pipeline can move to ACTIVE state and be run.
🗺️ DAG (Directed Acyclic Graph)
The DAG is the map of which step depends on which other step. "Acyclic" just means: no loops allowed! A step cannot depend on itself, and you cannot create a circular dependency (Step A → Step B → Step A would be a loop — not allowed!).
🧒 Think of it like: A family tree. A child can have parents, but a parent cannot also be a child of their own child. No loops!
OCI automatically runs steps in parallel when they don't depend on each other. For example, if Step 2 and Step 3 both only depend on Step 1, OCI will run Step 2 and Step 3 at the same time — saving you time!
▶️ Pipeline Run
A Pipeline Run is one actual execution of the pipeline. You can think of the Pipeline as the recipe, and the Pipeline Run as actually cooking the meal. You can run the same pipeline many times — each run is tracked separately with its own logs and status.
When creating a pipeline run, you can also override settings — for example, use different environment variables for this specific run without changing the pipeline definition itself.
🔄 Pipeline Step Run
Inside each Pipeline Run, there is a Pipeline Step Run for each step. This is the actual execution of that one step during that one run. You can see the status and logs of each individual step run separately.
5. 🔧 Types of Pipeline Steps
OCI Pipelines support three main types of steps. Let's understand each one:
Type 1: Script Step (Most Common for Beginners! ⭐)
A Script Step runs a Python, Bash, or Java script that you upload as the artifact. You also choose which Conda environment to use (Oracle provides many pre-built ones with TensorFlow, PyTorch, scikit-learn, etc. already installed).
🧒 Analogy: You write a recipe on a card, give it to the robot, tell it which tools (Conda environment) to use, and it cooks your dish!
Best for: data processing, model training, evaluation, saving models — any standard Python ML task.
Type 2: Job Step (Reference an Existing OCI Job)
A Job Step points to an existing Data Science Job that you've already created and tested. Instead of uploading code again, you just give it the OCID (the unique ID) of that job.
🧒 Analogy: Instead of writing the recipe again, you just say "use recipe #47 from my recipe book." Much faster!
This is the recommended approach for production pipelines because it lets you test each job independently before plugging them into the pipeline. If one job breaks, you only need to fix and re-test that job — not the whole pipeline.
Type 3: Custom Container Step / BYOC (Bring Your Own Container)
A BYOC Step lets you package your entire environment in a Docker container and run that inside the pipeline step. You build the container image, push it to OCI Container Registry (OCIR), and reference it in the step.
🧒 Analogy: Instead of using the robot kitchen's built-in tools, you bring your own complete kitchen-in-a-box with all the special tools you need!
Best for: non-Python code, very specific library versions, complex system dependencies, or teams already using Docker.
Type 4: Data Flow Step (Run Apache Spark Jobs) — New in 2024!
A Data Flow Step integrates OCI Data Flow (Oracle's Apache Spark as a Service) directly into your pipeline. This is perfect when you need to process huge datasets (millions of rows) that don't fit on a single machine — Spark spreads the work across many machines automatically.
Best for: big data processing, large-scale ETL (Extract, Transform, Load) operations before training.
6. ⚙️ How Pipelines Are Configured (Under the Hood)
Every OCI Pipeline has two levels of configuration. Think of it like a workplace: the company has general rules (pipeline-level), but individual departments can have their own specific rules (step-level) that override the general ones.
Level 1: Pipeline-Level Default Configuration
These settings apply to all steps in the pipeline unless a step overrides them:
- Compute Shape: The type of machine to use (e.g.,
VM.Standard.E4.Flexfor CPU work,VM.GPU.A10.1for GPU training) - OCPUs and Memory: How many CPU cores and how much RAM each step gets
- Block Storage: How much disk space (between 50 GB and 10,240 GB). Default is 100 GB.
- Log Group: Where to send all log output from all steps
- Environment Variables: Key-value pairs your scripts can read (like settings or file paths)
- Maximum Runtime: How long a step is allowed to run before being automatically killed (prevents runaway costs!)
- Network: Default networking (managed by Oracle, easiest for beginners) or Custom VCN (your own private network, for advanced security)
Level 2: Step-Level Configuration (Overrides Level 1)
Each individual step can override any of the pipeline-level settings. For example, your data-prep step might use a small CPU machine, but your training step needs a big GPU machine. You simply set a different compute shape for the training step — it overrides the pipeline default just for that step!
💡 Important: The Step Run configuration (what you set when actually running the pipeline) takes priority over both the step definition config AND the pipeline-level config. So you can even override settings at run time without changing the pipeline definition!
What is a Compute Shape?
A compute shape determines what kind of machine your step runs on. Here are the most common ones:
VM.Standard.E4.Flex— A flexible CPU machine. Great for data processing and lightweight tasks. You choose how many OCPUs (1–64) and how much RAM you need.VM.Standard3.Flex— Another flexible CPU option, based on Intel chips.VM.GPU.A10.1— Has 1 NVIDIA A10 GPU. Great for training medium-sized deep learning models.BM.GPU.A100-v2.8— A bare metal server with 8 NVIDIA A100 GPUs. For training very large models. (Very powerful and expensive — use wisely!)
What are Environment Variables?
Environment variables are like sticky notes that you attach to your pipeline step.
Your Python script can read them using os.environ.get("VARIABLE_NAME").
They are perfect for passing settings like model names, bucket names, or thresholds to your script without hardcoding them inside the script itself. This way, you can reuse the same script with different settings in different pipeline runs!
Never put passwords, API keys, or sensitive data directly in your Python file. Use environment variables for non-sensitive settings, and OCI Vault for sensitive secrets like API keys. Your code file might end up on GitHub one day — secrets should never be in code files!
7. 🖱️ Step-by-Step: Create Your First Pipeline Using OCI Console
We'll start with the OCI Console — this is the website interface, no code needed! Perfect for absolute beginners to understand what's happening visually.
Prerequisites — Do These First!
Step A: Create an OCI Free Account
Go to oracle.com/cloud/free and sign up. Oracle gives you $300 free credits and always-free services to get started. You'll need a credit card for verification but you won't be charged during the free tier.
Step B: Create a Compartment
A Compartment is like a folder that organises all your OCI resources together.
Navigate to: Identity & Security → Compartments → Create Compartment.
Give it a name like my-ai-project.
All your pipeline resources will live inside this compartment.
Step C: Set Up Required Policies
OCI uses Policies (permission rules) to control what can access what. Your pipeline steps need permission to access Object Storage, Logging, and other services. Ask your OCI administrator (or do it yourself if you're the admin) to create a Policy with these statements:
allow dynamic-group <your-dynamic-group-name> to manage data-science-family
in compartment <your-compartment-name>
allow dynamic-group <your-dynamic-group-name> to manage object-family
in compartment <your-compartment-name>
allow dynamic-group <your-dynamic-group-name> to use log-groups
in compartment <your-compartment-name>
allow dynamic-group <your-dynamic-group-name> to use log-content
in compartment <your-compartment-name>
Step D: Create a Data Science Project
Navigate to: Analytics & AI → Data Science → Projects → Create Project.
A Project is a workspace that holds all your notebooks, jobs, and pipelines together.
Give it a meaningful name like customer-churn-prediction.
Step E: Create a Log Group in OCI Logging
Navigate to: Observability & Management → Logging → Log Groups → Create Log Group.
Name it pipeline-logs.
This is where all your pipeline step outputs and errors will be stored so you can read them later.
Now — Create the Pipeline!
Step 1: Open Pipelines in Your Project
Inside your Data Science Project, look at the left-side menu. Click Pipelines. Then click the blue button "Create Pipeline".
Step 2: Fill in Pipeline Basics
Enter a name for your pipeline (e.g., my-first-ml-pipeline).
Optionally add a description.
Under Logging, select your Log Group (pipeline-logs) and enable auto-log creation.
OCI will automatically create separate logs for each step run.
Step 3: Configure Default Infrastructure
Set the default compute shape for all steps.
For beginners, choose VM.Standard.E4.Flex with 1 OCPU and 16 GB RAM.
Set Block Storage to 100 GB (the default).
For networking, select "Default networking" — Oracle manages this for you, easiest option!
Step 4: Add Your First Step
Click "Add pipeline step".
Choose "Script" (not a job or container — we want to upload our own script).
Give the step a name: data-preparation.
Upload your Python script file as the artifact.
Under Environment, choose a Conda environment — select generalml_p311_cpu_x86_64_v1 (this has scikit-learn, pandas, and other common ML libraries pre-installed).
Step 5: Add Your Second Step
Click "Add pipeline step" again.
Name this one model-training.
Upload your training script.
Under "Step dependencies", select data-preparation.
This tells OCI: only start this step AFTER the data-preparation step has finished successfully!
Step 6: Review and Create
Click "Review" to see a summary of your pipeline configuration. If everything looks correct, click "Create".
OCI will now create the pipeline. It will first be in CREATING state. Once your step artifacts are uploaded and verified, it moves to ACTIVE state. When ACTIVE, you can run it!
Step 7: Run Your Pipeline!
On the pipeline detail page, click "Run pipeline".
Give this run a name like first-run-2026.
You can also override environment variables just for this run if needed.
Click "Create" to start the run.
8. 💻 Create a Pipeline Using Python Code (ADS SDK)
The ADS (Accelerated Data Science) SDK is Oracle's official Python library for OCI Data Science. It makes creating pipelines with Python code clean, readable, and fast. Think of it as a friendly remote control for OCI — instead of clicking buttons on the website, you write a few lines of Python!
First — Install ADS
Open your terminal (or OCI Notebook Session) and run this command to install ADS:
This installs the ADS Python library onto your machine (or notebook session). ADS is the toolkit that lets you talk to OCI Data Science using Python code instead of clicking through the website.
pip install oracle-ads
If you're working inside an OCI Notebook Session, ADS is already pre-installed! You just need to import it.
Step 1 — Authenticate with OCI
Before doing anything, you need to tell ADS who you are so it can talk to OCI on your behalf. There are two ways to authenticate:
- API Key auth: For running code on your local laptop. Reads your
~/.oci/configfile. - Resource Principal: For running code inside OCI (inside a Notebook Session or Job). No config file needed — OCI knows who you are automatically.
This tells the ADS library how to log into OCI. It's like showing your ID badge at the door before entering the building. Use "api_key" when working on your laptop. Use "resource_principal" when working inside an OCI Notebook.
import ads # Use this when running on your local laptop ads.set_auth(auth="api_key") # OR use this when running inside an OCI Notebook Session # ads.set_auth(auth="resource_principal")
Step 2 — Write Your Step Scripts
Each pipeline step needs a Python script (the "artifact").
Let's write two simple scripts — one for data preparation and one for model training.
Save each one as a separate .py file on your computer.
This is the script for Step 1 of our pipeline — Data Preparation. It pretends to load data, cleans it up, and saves the cleaned version to a file. In a real project, you would replace the fake data here with real code to load your actual dataset. The key thing to notice: it uses os.environ.get() to read settings from environment variables — no hardcoded values!
# data_prep.py — This is the script for Step 1: Data Preparation
# OCI will run this file automatically inside Step 1 of the pipeline
import os
import json
import logging
# Set up good logging so we can see what's happening in OCI Logging
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s [%(levelname)s] %(message)s'
)
log = logging.getLogger("data-prep")
# Read settings from environment variables (not hardcoded!)
# This means we can change these settings in the pipeline config without touching the code
data_source = os.environ.get("DATA_SOURCE", "sample_data.csv")
output_path = os.environ.get("OUTPUT_PATH", "/home/datascience/clean_data.json")
log.info(f"Starting data preparation...")
log.info(f"Reading data from: {data_source}")
# Simulate loading and cleaning data
# In a real project, replace this with: df = pd.read_csv(data_source)
raw_data = [
{"customer_id": 1, "age": 25, "spending": 1200, "churned": 0},
{"customer_id": 2, "age": 45, "spending": 300, "churned": 1},
{"customer_id": 3, "age": 33, "spending": 850, "churned": 0},
{"customer_id": 4, "age": 52, "spending": 100, "churned": 1},
]
# Simulate cleaning: remove rows where spending is less than 0 (bad data)
clean_data = [row for row in raw_data if row["spending"] >= 0]
log.info(f"Cleaned data: {len(clean_data)} records ready for training")
# Save the cleaned data so the NEXT step can read it
with open(output_path, "w") as f:
json.dump(clean_data, f)
log.info(f"Clean data saved to: {output_path}")
log.info("Data preparation complete! ✅")
This is the script for Step 2 — Model Training. It reads the cleaned data that Step 1 saved, trains a simple machine learning model on it, checks how accurate the model is, and saves the result to a file. This is a toy example — in a real project you would train a TensorFlow or PyTorch model here.
# train.py — This is the script for Step 2: Model Training
# OCI runs this ONLY after data_prep.py has finished successfully
import os
import json
import logging
import sys
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s [%(levelname)s] %(message)s'
)
log = logging.getLogger("model-training")
# Read settings from environment variables
data_path = os.environ.get("CLEAN_DATA_PATH", "/home/datascience/clean_data.json")
model_name = os.environ.get("MODEL_NAME", "churn_model_v1")
epochs = int(os.environ.get("EPOCHS", "10"))
min_accuracy = float(os.environ.get("MIN_ACCURACY", "0.75"))
log.info(f"Starting model training for: {model_name}")
log.info(f"Training for {epochs} epochs")
# Load the clean data from the previous step
try:
with open(data_path, "r") as f:
data = json.load(f)
log.info(f"Loaded {len(data)} training records")
except FileNotFoundError:
log.error(f"Clean data not found at {data_path} — did Step 1 run successfully?")
sys.exit(1) # Exit code 1 = FAILURE — tells OCI this step FAILED
# Simulate model training (replace with real ML code in your project)
# Example: model = RandomForestClassifier(); model.fit(X_train, y_train)
import time
for epoch in range(1, epochs + 1):
time.sleep(0.5) # Simulating training work
fake_accuracy = 0.60 + (epoch * 0.025) # Accuracy improves each epoch
log.info(f" Epoch {epoch}/{epochs} — Accuracy: {fake_accuracy:.3f}")
final_accuracy = fake_accuracy
log.info(f"Training complete. Final accuracy: {final_accuracy:.3f}")
# Check if accuracy is good enough
if final_accuracy < min_accuracy:
log.error(f"Accuracy {final_accuracy:.3f} is below minimum {min_accuracy}. NOT saving model!")
sys.exit(1) # FAIL the step — don't deploy a bad model!
# Save model results
result = {
"model_name": model_name,
"final_accuracy": final_accuracy,
"status": "success"
}
with open("/home/datascience/model_result.json", "w") as f:
json.dump(result, f)
log.info(f"Model saved successfully! ✅")
If your script encounters a serious error and calls
sys.exit(1),
OCI marks that step as FAILED and automatically stops all dependent steps from running.
If you don't do this, OCI might think your step succeeded even when it crashed — and will run the next steps on broken data!
Step 3 — Create the Pipeline with ADS SDK
This is the main code that creates your pipeline in OCI! Think of it as filling in a very detailed form — but using Python instead of clicking website buttons. It defines two steps (data-prep and training), sets their compute shape, links the scripts to each step, tells OCI which step depends on which other step, and then actually creates the whole pipeline in OCI Data Science. After creating it, it immediately runs the pipeline and shows you a live status update!
# create_and_run_pipeline.py
# Run this from your local laptop or OCI Notebook Session
# This will create the pipeline in OCI Data Science and then run it!
import ads
import os
from ads.pipeline import Pipeline, PipelineStep, CustomScriptStep
from ads.jobs import ScriptRuntime
# ──────────────────────────────────────────────────────────────
# STEP A: Authenticate with OCI
# ──────────────────────────────────────────────────────────────
ads.set_auth(auth="api_key") # Change to "resource_principal" inside OCI Notebook
# ──────────────────────────────────────────────────────────────
# STEP B: Set your OCI IDs — replace these with your actual values!
# You find these in the OCI Console on the respective resource pages
# ──────────────────────────────────────────────────────────────
compartment_id = "ocid1.compartment.oc1..your_compartment_id_here"
project_id = "ocid1.datascienceproject.oc1..your_project_id_here"
log_group_id = "ocid1.loggroup.oc1..your_log_group_id_here"
# ──────────────────────────────────────────────────────────────
# STEP C: Define the infrastructure (what kind of machine each step uses)
# CustomScriptStep = the machine spec that runs our Python scripts
# ──────────────────────────────────────────────────────────────
infrastructure = (
CustomScriptStep()
.with_shape_name("VM.Standard.E4.Flex") # Flexible CPU machine
.with_shape_config_details(ocpus=1, memory_in_gbs=16) # 1 CPU core, 16 GB RAM
.with_block_storage_size(100) # 100 GB disk space
)
# ──────────────────────────────────────────────────────────────
# STEP D: Define the runtime for each step (what code runs + which conda env)
# ScriptRuntime = run a Python/Bash script using a Conda environment
# ──────────────────────────────────────────────────────────────
data_prep_runtime = (
ScriptRuntime()
.with_source("data_prep.py") # Our data prep script
.with_service_conda("generalml_p311_cpu_x86_64_v1") # Pre-built conda with ML packages
.with_environment_variable(
DATA_SOURCE="oci://my-bucket@namespace/raw_data.csv",
OUTPUT_PATH="/home/datascience/clean_data.json"
)
.with_maximum_runtime_in_minutes(30) # Kill if it runs over 30 minutes!
)
training_runtime = (
ScriptRuntime()
.with_source("train.py") # Our training script
.with_service_conda("generalml_p311_cpu_x86_64_v1")
.with_environment_variable(
CLEAN_DATA_PATH="/home/datascience/clean_data.json",
MODEL_NAME="churn_model_v1",
EPOCHS="20",
MIN_ACCURACY="0.75"
)
.with_maximum_runtime_in_minutes(120) # Training can take up to 2 hours
)
# ──────────────────────────────────────────────────────────────
# STEP E: Create PipelineStep objects — each step gets a name,
# its infrastructure, its runtime, and optionally environment variables
# ──────────────────────────────────────────────────────────────
step_data_prep = (
PipelineStep("data-preparation")
.with_description("Load raw data and clean it for training")
.with_infrastructure(infrastructure)
.with_runtime(data_prep_runtime)
)
step_training = (
PipelineStep("model-training")
.with_description("Train the churn prediction model")
.with_infrastructure(infrastructure)
.with_runtime(training_runtime)
)
# ──────────────────────────────────────────────────────────────
# STEP F: Assemble the Pipeline!
# .with_dag() defines the order — "step1 >> step2" means step2 runs after step1
# ──────────────────────────────────────────────────────────────
pipeline = (
Pipeline("customer-churn-ml-pipeline")
.with_compartment_id(compartment_id)
.with_project_id(project_id)
.with_log_group_id(log_group_id)
.with_step_details([step_data_prep, step_training])
.with_dag(["data-preparation >> model-training"]) # Step 2 runs after Step 1
)
# ──────────────────────────────────────────────────────────────
# STEP G: Create the pipeline in OCI (saves the definition)
# ──────────────────────────────────────────────────────────────
pipeline.create()
print(f"✅ Pipeline created! ID: {pipeline.id}")
# Optional: visualize the pipeline structure in your notebook
pipeline.show()
# ──────────────────────────────────────────────────────────────
# STEP H: Run the pipeline!
# This triggers one Pipeline Run — OCI will spin up machines and execute the steps
# ──────────────────────────────────────────────────────────────
pipeline_run = pipeline.run()
print(f"🚀 Pipeline run started! Run ID: {pipeline_run.id}")
# This streams live logs to your terminal/notebook as the pipeline runs
# It blocks until the pipeline finishes (succeeds, fails, or is cancelled)
pipeline_run.watch()
# After .watch() returns, check the final status
print(f"🏁 Final status: {pipeline_run.status}")
What happens when you run this code:
- Python connects to OCI using your credentials
- OCI creates the pipeline definition (like saving a recipe card)
- You press "run" — OCI spins up a VM, runs
data_prep.pyon it, then stops that VM - Then OCI spins up a NEW VM, runs
train.pyon it, then stops that VM too - You see live status updates in your terminal as it runs
- All logs are also available in OCI Logging for later review
Step 4 — Build a Pipeline from Existing Jobs (Recommended for Production!)
As mentioned earlier, the best practice for real projects is to first create and test each step as a standalone Job, then reference those tested jobs in your pipeline. Here is how to build a pipeline from already-existing jobs:
Instead of attaching scripts directly to pipeline steps, this code references existing OCI Jobs by their OCID (their unique ID). This is powerful because you can test each job individually first to make sure it works, and THEN plug the tested jobs into the pipeline for the final automated workflow. It's like testing each recipe step on its own before combining them into the full meal!
from ads.pipeline import Pipeline, PipelineStep
import ads
ads.set_auth(auth="api_key")
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
project_id = "ocid1.datascienceproject.oc1..your_project_id"
log_group_id = "ocid1.loggroup.oc1..your_log_group_id"
# Reference existing jobs by their OCID
# You find these OCIDs on the Jobs page inside your OCI Data Science Project
step_one = (
PipelineStep("data-preparation")
.with_description("Load and clean raw data")
.with_job_id("ocid1.datasciencejob.oc1..your_data_prep_job_id")
)
step_two = (
PipelineStep("model-training")
.with_description("Train the churn prediction model")
.with_job_id("ocid1.datasciencejob.oc1..your_training_job_id")
)
step_three = (
PipelineStep("model-evaluation")
.with_description("Evaluate model and save to model catalog")
.with_job_id("ocid1.datasciencejob.oc1..your_evaluation_job_id")
)
# Build and run the pipeline
# with_dag uses ">>" to define order: step1 must finish before step2 starts
pipeline = (
Pipeline("production-churn-pipeline")
.with_compartment_id(compartment_id)
.with_project_id(project_id)
.with_log_group_id(log_group_id)
.with_step_details([step_one, step_two, step_three])
.with_dag(["data-preparation >> model-training >> model-evaluation"])
)
pipeline.create()
print("✅ Pipeline created from existing jobs!")
pipeline_run = pipeline.run()
pipeline_run.watch()
9. ⚡ Running Steps in Parallel (Speed Up Your Pipeline!)
One of the most powerful Pipeline features is parallel execution. If two steps don't depend on each other, OCI will run them at the same time — cutting your total pipeline time dramatically!
🧒 Analogy: Imagine you are making breakfast. You don't have to toast the bread first, THEN brew coffee, THEN fry eggs one by one. You can do all three at the same time! The pipeline does the same thing — runs independent steps simultaneously.
Here is a real example: training 3 different models in parallel to find the best one:
Step 1: data-preparation
↓
┌────────┼────────┐
↓ ↓ ↓
Step 2a Step 2b Step 2c
Train RF Train XGB Train NN
└────────┼────────┘
↓
Step 3: pick-best-model-and-deploy
This creates a pipeline where Steps 2a, 2b, and 2c all run at the SAME TIME after Step 1 finishes. Then Step 3 waits for ALL THREE of them to finish before running. The magic is in the
with_dag() line — the parentheses (step2a, step2b, step2c)
mean "run all of these in parallel."
pipeline = (
Pipeline("parallel-model-training-pipeline")
.with_compartment_id(compartment_id)
.with_project_id(project_id)
.with_log_group_id(log_group_id)
.with_step_details([
step_data_prep,
step_train_random_forest, # These 3 steps have no dependency on each other
step_train_xgboost, # So OCI will run all 3 in PARALLEL after data-prep
step_train_neural_net,
step_pick_best_model
])
# DAG syntax:
# "A >> B" means B runs after A
# "(A, B, C) >> D" means D runs after ALL of A, B, and C finish
.with_dag([
"data-preparation >> (train-random-forest, train-xgboost, train-neural-net)",
"(train-random-forest, train-xgboost, train-neural-net) >> pick-best-model"
])
)
pipeline.create()
pipeline_run = pipeline.run()
pipeline_run.watch()
10. 📊 Running, Monitoring, and Reading Pipeline Logs
Once your pipeline is running, you need to know how to check what's happening and read logs when something goes wrong.
How to Monitor in the OCI Console
- Navigate to your Data Science Project → Pipelines → Click your pipeline → Pipeline Runs
- Click on a specific run to see the step-by-step status — each step shows ACCEPTED, IN_PROGRESS, SUCCEEDED, or FAILED
- The console shows a visual graph of your DAG with real-time coloured status for each step (green = success, red = failed, blue = running)
- Click on any individual step to see its logs directly in the console
Pipeline Run Lifecycle States
A pipeline run goes through these states:
- ACCEPTED → OCI has received your run request and is getting ready
- IN_PROGRESS → At least one step is currently running
- SUCCEEDED → All steps completed successfully! 🎉
- FAILED → At least one step failed. Check the logs to find out why.
- CANCELING / CANCELED → You (or someone) cancelled the run
- DELETING / DELETED → The run is being deleted
How to Read Logs in OCI Logging
All your print() statements and log.info() messages go here:
- Navigate to: Observability & Management → Logging → Log Groups
- Click your Log Group → click the Log name → you'll see all log messages with timestamps
- Use the search bar to filter by step name, error keyword (like "ERROR"), or time range
- You can also click "Explore with Log Explorer" for advanced filtering and charts
Monitor from Python Code
This shows how to check the status of a pipeline run from Python code. Useful when you want to trigger a pipeline from your own application and wait for it to complete. The loop checks every 15 seconds and prints the current status until the pipeline finishes.
import ads
import oci
import time
ads.set_auth(auth="api_key")
config = oci.config.from_file()
ds_client = oci.data_science.DataScienceClient(config)
pipeline_run_id = "ocid1.datasciencepipelinerun.oc1..your_run_id"
# Poll until pipeline finishes
while True:
run = ds_client.get_pipeline_run(pipeline_run_id).data
status = run.lifecycle_state
print(f" Pipeline status: {status}")
# These states mean the pipeline has stopped (either success or failure)
if status in ["SUCCEEDED", "FAILED", "CANCELED"]:
break
time.sleep(15) # Check again in 15 seconds
if status == "SUCCEEDED":
print("🎉 Pipeline completed successfully!")
else:
print(f"❌ Pipeline ended with status: {status}. Check your logs!")
11. ⏰ Schedule a Pipeline to Run Automatically
What if you want your pipeline to run every Monday at 6am to retrain your model on fresh data? OCI Data Science now has a built-in Scheduler that lets you do exactly that!
Schedule via OCI Console
- Go to your Data Science Project → click "Schedules" in the left menu
- Click "Create Schedule"
- Choose resource type: Pipeline, select your pipeline
- Set the frequency: choose from Hourly, Daily, Weekly, Monthly, or enter a custom Cron expression
- Set a start date/time, and optionally an end date
- Click Create
That's it! OCI will now automatically trigger a pipeline run on your chosen schedule — no extra code needed.
Cron Expression Examples
A cron expression is a compact text format for describing a schedule.
It looks like: minute hour day month weekday
0 6 * * 1→ Every Monday at 6:00 AM0 0 * * *→ Every day at midnight0 8 1 * *→ First day of every month at 8:00 AM30 12 * * 1-5→ Every weekday (Mon–Fri) at 12:30 PM
12. 🐳 Advanced: Bring Your Own Container (BYOC)
Sometimes the built-in Conda environments don't have the exact libraries you need. Maybe you need a specific version of a library, or you want to use R instead of Python, or your team already uses Docker containers. That's when BYOC (Bring Your Own Container) is perfect!
🧒 Analogy: Instead of using Oracle's pre-set kitchen (Conda), you bring your own packed lunch box (Docker container) with exactly the food you want inside!
How BYOC Works — 3 Simple Steps
Step A: Write a Dockerfile
A Dockerfile is a recipe for building a Docker container. This one starts from a base Python image, installs some special ML libraries (XGBoost and LightGBM), copies your training script into the container, and sets it to run automatically when the container starts. Once built, this container has everything your step needs — no Conda environment required!
# Dockerfile — defines the custom environment for our pipeline step
# Start from Python 3.11 base image
FROM python:3.11-slim
# Set working directory inside the container
WORKDIR /app
# Install the exact ML libraries we need
RUN pip install --no-cache-dir \
xgboost==2.0.3 \
lightgbm==4.3.0 \
scikit-learn==1.4.0 \
pandas==2.2.0
# Copy our training script into the container
COPY train_xgboost.py /app/train_xgboost.py
# Set the command to run when the container starts
CMD ["python", "/app/train_xgboost.py"]
Step B: Build and Push to OCI Container Registry (OCIR)
These terminal commands build the Docker container from your Dockerfile, log in to Oracle's container registry (OCIR), tag the container with your OCIR address, and push (upload) it there so OCI can pull it when running your pipeline step. Replace <region> with your OCI region (e.g.,
fra for Frankfurt) and <namespace> with your OCI tenancy namespace.
# Build the container docker build -t my-xgboost-trainer:v1 . # Log in to OCI Container Registry docker login <region>.ocir.io --username <namespace>/<your-username> # Tag the image with the OCIR address docker tag my-xgboost-trainer:v1 <region>.ocir.io/<namespace>/my-xgboost-trainer:v1 # Push to OCIR docker push <region>.ocir.io/<namespace>/my-xgboost-trainer:v1
Step C: Reference the Container in Your Pipeline Step
Instead of using ScriptRuntime (which uses a Conda environment), we use ContainerRuntime, which tells OCI to pull our custom Docker image from OCIR and run it as the pipeline step. The step will use our container with its exact library versions instead of any Conda environment.
from ads.pipeline import Pipeline, PipelineStep, CustomScriptStep
from ads.jobs import ContainerRuntime
import ads
ads.set_auth(auth="api_key")
byoc_image = "<region>.ocir.io/<namespace>/my-xgboost-trainer:v1"
infrastructure = (
CustomScriptStep()
.with_shape_name("VM.Standard.E4.Flex")
.with_shape_config_details(ocpus=2, memory_in_gbs=32)
.with_block_storage_size(100)
)
# Use ContainerRuntime instead of ScriptRuntime for BYOC
container_runtime = (
ContainerRuntime()
.with_image(byoc_image) # Point to our custom Docker image in OCIR
.with_replica(1) # Run on a single machine (not distributed)
)
step_xgboost = (
PipelineStep("xgboost-training")
.with_description("Train XGBoost model using custom container")
.with_infrastructure(infrastructure)
.with_runtime(container_runtime)
)
# Then assemble the pipeline as usual...
13. ✅❌ Best Practices — DOs and DON'Ts for Beginners
Create and test each Job individually first. Once each job works perfectly on its own, then connect them into a pipeline. This saves you hours of debugging — it's much easier to fix one small job than to debug an entire pipeline!
Never put sensitive values directly in your
.py files. Use environment variables for non-sensitive settings,
and OCI Vault for secrets. Your code might end up on GitHub or be shared — secrets should never be in code files!
Always configure
with_maximum_runtime_in_minutes() for each step.
If your step gets stuck in an infinite loop or hangs, without a timeout it will keep running (and billing you) forever.
Set a realistic maximum — 30 minutes for data prep, 120 minutes for training, etc.
Relying only on
print() makes it hard to filter messages later. Use Python's logging library instead.
With logging.info(), logging.error(), etc., every message gets a timestamp and severity level —
making it much easier to search and debug in OCI Logging later.
This is critical! Calling
sys.exit(1) tells OCI that the step FAILED, which automatically prevents downstream steps from running.
Without this, OCI may think the step succeeded even if your script crashed midway and wrote broken output data.
Resist the temptation to write one huge Python script that does everything (data prep + training + evaluation + deployment all in one). Break your work into meaningful steps. This way, if evaluation fails, you don't have to re-run data prep and training again — saving time and money!
If you are comparing multiple models (Random Forest vs XGBoost vs Neural Net), train them in parallel! Set up parallel DAG steps so they run simultaneously — this can cut your pipeline runtime from 3 hours to 1 hour.
Data cleaning and feature engineering don't need a GPU — they run fine on a cheap CPU machine. Only use GPU shapes (like
VM.GPU.A10.1) for the actual model training step.
Assign the right compute shape to each step individually and save a LOT of money!
Name your script files with versions — e.g.,
train_v2.py, data_prep_v3.py.
If you update a script and the pipeline breaks, you can easily roll back to the previous version.
14. 📝 Quick Reference — Key ADS SDK Methods
Here is a handy cheat sheet of the most important ADS Pipeline commands:
ads.set_auth(auth="api_key")→ Authenticate using your OCI config file (for local laptop)ads.set_auth(auth="resource_principal")→ Authenticate inside OCI Notebook/Job (no config file needed)Pipeline("name")→ Create a new pipeline object.with_compartment_id("ocid1...")→ Set which compartment the pipeline lives in.with_project_id("ocid1...")→ Set which Data Science project it belongs to.with_log_group_id("ocid1...")→ Set where logs go.with_step_details([step1, step2, ...])→ Add all your steps.with_dag(["step1 >> step2"])→ Define step order and dependenciespipeline.create()→ Actually create the pipeline in OCI (saves the definition)pipeline.show()→ Display a visual diagram of the pipeline in your notebookpipeline.run()→ Start a pipeline run and return a PipelineRun objectpipeline_run.watch()→ Stream live logs until the run finishespipeline_run.status→ Get the current status string of the runPipelineStep("name")→ Create a step object.with_job_id("ocid1...")→ Reference an existing OCI Job as a step.with_infrastructure(CustomScriptStep())→ Set compute shape for the step.with_runtime(ScriptRuntime())→ Set script and Conda environment for the stepScriptRuntime().with_source("train.py")→ Set the script to run.with_service_conda("generalml_p311_cpu_x86_64_v1")→ Use a built-in Conda env.with_environment_variable(KEY="value")→ Pass env vars to the step.with_maximum_runtime_in_minutes(60)→ Set a timeout for the step
Summary 📝
What we learned — complete OCI Pipeline journey:
- OCI → Oracle's cloud platform — rented supercomputers on the internet
- OCI Data Science → A fully managed AI lab with Notebooks, Jobs, Pipelines, and Model Catalog
- Pipeline → An automated, multi-step ML workflow that runs in the right order every time
- Pipeline Step → One individual task (data prep, training, evaluation) — each runs on its own machine
- Step Artifact → The actual script (
.py,.sh,.zip) that runs inside the step - DAG → The dependency map — defines which steps must run before others, and which can run in parallel
- Pipeline Run → One actual execution of the pipeline. Same pipeline can be run many times.
- Step Types → Script step, Job step, BYOC (Docker) step, Data Flow (Spark) step
- Two config levels → Pipeline defaults apply to all steps; step-level config overrides them for that step only
- ADS SDK → Oracle's Python library that makes building pipelines clean and easy with
.with_*()methods - Parallel steps → Use DAG syntax
(stepA, stepB) >> stepCto run A and B simultaneously - Scheduler → Set your pipeline to run automatically on a cron schedule (e.g., every Monday at 6am)
- BYOC → Package your entire environment in a Docker container and run it as a pipeline step
- Logging → All output goes to OCI Logging — use Python's
logginglibrary and alwayssys.exit(1)on errors
Happy building! ☁️🤖✨
Comments
Post a Comment