Imagine you are a scientist with a brand-new laboratory. 🔬
You need a workbench to run experiments, the right chemicals (libraries) on your shelf,
a notebook to record every experiment, and a filing cabinet to save your work safely.
OCI Data Science gives you all of this — in the cloud —
fully managed, secure, and ready in minutes!
🗺️ Big Picture — What is OCI Data Science?
OCI Data Science is a fully managed platform on Oracle Cloud where data scientists and developers can build, train, and deploy AI/ML models without worrying about managing servers, installing software, or setting up environments.
┌──────────────────────────────────────────────────────────────────────────┐ │ OCI DATA SCIENCE PLATFORM │ │ │ │ ┌────────────┐ ┌──────────────────┐ ┌───────────────────────────┐ │ │ │ PROJECT │ │ NOTEBOOK SESSION │ │ CONDA ENVIRONMENT │ │ │ │ │ │ │ │ │ │ │ │ Team │──►│ Your JupyterLab │──►│ Python + Libraries │ │ │ │ workspace │ │ Cloud Workbench │ │ (scikit-learn, torch etc) │ │ │ └────────────┘ └──────────────────┘ └───────────────────────────┘ │ │ │ │ │ │ │ ▼ ▼ ▼ │ │ ┌─────────────┐ ┌─────────────────┐ ┌───────────────────────────┐ │ │ │ IAM Policies│ │ Git Repository │ │ Object Storage Backup │ │ │ │ (Access │ │ (Code Version │ │ (Custom Conda Env Save) │ │ │ │ Control) │ │ Control) │ │ │ │ │ └─────────────┘ └─────────────────┘ └───────────────────────────┘ │ └──────────────────────────────────────────────────────────────────────────┘
Every concept above is something you will master in this guide. Let's go one layer at a time! 🎯
📚 Key Concepts — Explained
- 🏢 Project — A Project is like a school subject folder. All your notebooks, models, and experiments for one topic live together inside it. E.g. one project for "Customer Churn Prediction", another for "Image Classifier".
- 💻 Notebook Session — This is your actual working computer in the cloud — a JupyterLab environment where you write and run Python code. You can switch it on when you need it and switch it off to save money!
- 🧪 Conda Environment — Conda is a package manager that creates isolated Python environments. Think of it like a separate lunchbox for each project — so libraries from one project never mess up another project!
- 🌿 Git Repository — Git tracks every change you make to your code — like a magical "Undo" history that goes back months! Multiple team members can work on the same code safely.
- ☁️ Object Storage Backup — When you install custom Python libraries not in OCI's pre-built list, you save your entire Conda environment to Object Storage so you (and your team) can restore it on any new notebook instantly.
🔒 Step 1 — Set Up IAM Policies (The Permission System)
Before anyone can use OCI Data Science, the right IAM (Identity and Access Management) policies must be in place. Think of IAM policies as the rules posted on a school notice board: "Only students in Class 10 may use the science lab." 🏫
These IAM policy statements tell OCI who is allowed to use the Data Science service, who can create/manage notebook sessions, and what resources the service itself can access. Without these, notebooks cannot start and models cannot be saved. 🔑
-- OCI Data Science IAM Policies -- Add these in OCI Console → Identity → Policies → Create Policy -- 1. Allow your data science team to use the service Allow group DataScienceTeam to use data-science-family in compartment DataScienceCompartment -- 2. Allow the service itself to read/write to Object Storage Allow service datascience to use object-family in tenancy -- 3. Allow notebook sessions to call other OCI services (e.g. Vision, Language) Allow dynamic-group DataScienceNotebooks to use ai-service-family in compartment DataScienceCompartment -- 4. Allow notebooks to read/write model artifacts in Object Storage Allow dynamic-group DataScienceNotebooks to manage objects in compartment DataScienceCompartment where target.bucket.name = 'data-science-models' -- 5. Allow team to manage models and jobs Allow group DataScienceTeam to manage data-science-model-deployments in compartment DataScienceCompartment
OCI provides a Policy Builder in the console that can generate these policies automatically! Go to Identity → Policies → Create Policy → Use Policy Builder and select "Data Science" from the service dropdown. 🪄
🏗️ Step 2 — Create a Data Science Project
A Project is the top-level container in OCI Data Science. Every notebook session, model, and job you create will belong to a project. It helps you keep work organised and control team access at the project level.
Option A: Create via OCI Console (No Code)
- Log into OCI Console at cloud.oracle.com
- Go to Analytics & AI → Data Science
- Click "Create Project"
- Enter a name:
Customer-Churn-Prediction - Select your compartment
- Add a description: "Predict which customers will cancel their subscription"
- Click "Create" ✅
Option B: Create via Python SDK
This code creates a new Data Science Project on OCI using Python. Think of it as telling Oracle: "Please create a new science lab folder called 'Customer-Churn-Prediction' where my team will store all their AI experiments." 📁
import oci
# Load OCI config from ~/.oci/config
config = oci.config.from_file()
ds_client = oci.data_science.DataScienceClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
# Create the project
response = ds_client.create_project(
oci.data_science.models.CreateProjectDetails(
compartment_id=compartment_id,
display_name="Customer-Churn-Prediction",
description="Predict which telecom customers will cancel their subscription next month."
)
)
project_id = response.data.id
print(f"✅ Project created successfully!")
print(f" Project ID : {project_id}")
print(f" Display Name : {response.data.display_name}")
print(f" Status : {response.data.lifecycle_state}")
Output:
✅ Project created successfully! Project ID : ocid1.datascienceproject.oc1.ap-mumbai-1.examplexxxxxxx Display Name : Customer-Churn-Prediction Status : ACTIVE
💻 Step 3 — Create a Notebook Session
A Notebook Session is a managed JupyterLab server running in the cloud. When you open a notebook session, you get a full Python environment — just like running Jupyter on your own laptop, but on powerful Oracle Cloud hardware! 🖥️
A Notebook Session is like renting a fully equipped laboratory for the day. When you are done experimenting, you return the lab (stop the session) and stop paying. The next time you rent it, everything is exactly where you left it! 🔬
Step 3a: Choose the Right Compute Shape
OCI NOTEBOOK COMPUTE SHAPES — CHOOSE WISELY ───────────────────────────────────────────────────────────────── SHAPE CPU RAM GPU BEST FOR ───────────────────────────────────────────────────────────────── VM.Standard.E4.Flex(2,32) 2 32GB No Exploration, EDA VM.Standard.E4.Flex(8,64) 8 64GB No Training small ML VM.GPU.A10.1 16 240GB 1xA10 Deep Learning, LLM VM.GPU.A10.2 32 480GB 2xA10 Large model training BM.GPU.A100-v2.8 128 2TB 8xA100 Foundation models ───────────────────────────────────────────────────────────────── 💡 Start small (E4.Flex) — you can always resize later!
Step 3b: Create Notebook Session via Console
- Inside your Project, click "Create Notebook Session"
- Enter a name:
churn-exploration-nb - Select compute shape:
VM.Standard.E4.Flex— set 4 OCPUs, 64GB RAM - Set Block Volume size:
100 GB(this is your local disk space) - Select a VCN and subnet (use the default if unsure)
- Click "Create" — it takes 3–5 minutes to start ✅
- Once status is ACTIVE, click "Open" to launch JupyterLab! 🎉
Step 3c: Create Notebook Session via Python SDK
This code creates and starts a cloud notebook server on OCI automatically. It sets the machine type (like choosing a laptop model), the disk size (like choosing how much hard drive space), and links it to our project. Once done, you can open JupyterLab in a browser! 🌐
import oci
import time
config = oci.config.from_file()
ds_client = oci.data_science.DataScienceClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
project_id = "ocid1.datascienceproject.oc1.ap-mumbai-1.examplexxxxxxx"
subnet_id = "ocid1.subnet.oc1.ap-mumbai-1.your_subnet_id"
# Create the notebook session
print("🚀 Creating Notebook Session...")
response = ds_client.create_notebook_session(
oci.data_science.models.CreateNotebookSessionDetails(
compartment_id=compartment_id,
project_id=project_id,
display_name="churn-exploration-nb",
# Compute shape — 4 CPU cores, 64 GB RAM
notebook_session_config_details=oci.data_science.models.NotebookSessionConfigDetails(
shape="VM.Standard.E4.Flex",
block_storage_size_in_gbs=100, # 100 GB local storage
notebook_session_shape_config_details=oci.data_science.models.NotebookSessionShapeConfigDetails(
ocpus=4,
memory_in_gbs=64
),
subnet_id=subnet_id
)
)
)
notebook_id = response.data.id
print(f" Notebook Session ID : {notebook_id}")
print(f" Status : {response.data.lifecycle_state}")
# Poll until the notebook is ready
print("\n⏳ Waiting for notebook to become ACTIVE...")
while True:
nb = ds_client.get_notebook_session(notebook_id)
state = nb.data.lifecycle_state
print(f" Status: {state}")
if state == "ACTIVE":
print(f"\n✅ Notebook is READY!")
print(f" Open URL: {nb.data.notebook_session_url}")
break
elif state == "FAILED":
print("❌ Notebook creation failed. Check shape availability and subnet.")
break
time.sleep(20)
Output:
🚀 Creating Notebook Session... Notebook Session ID : ocid1.datasciencenotebooksession.oc1..examplexxxxxxx Status : CREATING ⏳ Waiting for notebook to become ACTIVE... Status: CREATING Status: CREATING Status: ACTIVE ✅ Notebook is READY! Open URL: https://notebook.ap-mumbai-1.oci.oraclecloud.com/?token=xxxxx
A running notebook session keeps charging you every hour, even if you are not using it. Like leaving a taxi meter running while you sleep — always click "Deactivate" when you are done for the day! Your notebooks and files are safely preserved on the block volume. 💸
🧪 Step 4 — Understanding and Using Conda Environments
When your notebook opens in JupyterLab, you need to choose which Python environment (Conda environment) to use. Think of Conda environments like different toolboxes — each one has exactly the tools needed for a specific job! 🧰
OCI Pre-Built Conda Environments
OCI provides a rich set of ready-made Conda environments — no setup required! Just choose one from the Environment Explorer and it installs in minutes.
OCI PRE-BUILT CONDA ENVIRONMENTS
──────────────────────────────────────────────────────────────────
ENVIRONMENT NAME KEY LIBRARIES INCLUDED
──────────────────────────────────────────────────────────────────
General Machine Learning scikit-learn, XGBoost, LightGBM,
pandas, numpy, matplotlib
TensorFlow 2.x for CPU/GPU TensorFlow, Keras, numpy
PyTorch for CPU/GPU PyTorch, torchvision, torchaudio
Natural Language Processing HuggingFace Transformers, spaCy,
NLTK, sentence-transformers
PySpark Apache Spark, Delta Lake, Koalas
Oracle AutoML Oracle AutoML, ads (OCI ADS SDK)
Computer Vision OpenCV, Pillow, detectron2
Time Series Prophet, statsmodels, sktime
──────────────────────────────────────────────────────────────────
Inside JupyterLab:
- Click the Environment Explorer icon in the left sidebar
- Browse the list of available environments
- Click "Install" next to the one you want
- Wait 3–5 minutes for it to install ☕
- Select it as your kernel in any new notebook — done! ✅
Step 4a: Install Conda Environment via Terminal
This command, run in the JupyterLab Terminal, installs a pre-built OCI Conda environment on your notebook server. It tells the OCI environment system: "Please download and set up the General Machine Learning environment for me." Like installing an app on your phone! 📱
# Open JupyterLab Terminal (File → New → Terminal) and run: # Install the General Machine Learning conda environment odsc conda install -s mlcpuv1 # Or install PyTorch GPU environment odsc conda install -s pytorch21_p39_gpu_v1 # List all installed environments on this notebook odsc conda list # Activate an environment manually in terminal conda activate /home/datascience/conda/mlcpuv1
Output:
Installing conda environment: mlcpuv1 Downloading... ████████████████████ 100% Extracting... ████████████████████ 100% Setting up kernel... ✅ Environment 'mlcpuv1' installed successfully! Kernel name: 'Python [conda env:mlcpuv1]' You can now select this kernel in your notebooks.
Step 4b: Create a Custom Conda Environment
Sometimes you need libraries that are not in any pre-built environment — like a niche finance library, a custom internal package, or a very specific version of a tool. In that case, you create your own custom Conda environment! 🔧
These commands, run in the JupyterLab Terminal, create a completely fresh Python 3.11 environment from scratch, install your custom libraries into it, and then register it as a Jupyter kernel so you can select it in notebooks. Like building your own custom toolbox instead of using a pre-made one! 🧰🔨
# In JupyterLab Terminal:
# Step 1: Create a brand new conda environment with Python 3.11
conda create -n my_custom_env python=3.11 -y
# Step 2: Activate it
conda activate my_custom_env
# Step 3: Install your standard ML libraries
pip install pandas numpy scikit-learn matplotlib seaborn
# Step 4: Install custom/niche libraries not in pre-built envs
pip install ta-lib # Technical Analysis Library for finance
pip install prophet # Facebook Prophet for time series
pip install my-company-internal-sdk # Your private company package
# Step 5: Register the environment as a Jupyter kernel
# So it appears in the kernel dropdown inside notebooks
python -m ipykernel install \
--user \
--name my_custom_env \
--display-name "Python (My Custom ML Env)"
echo "✅ Custom Conda environment ready!"
echo " Select 'Python (My Custom ML Env)' from the kernel menu in JupyterLab"
Output:
Collecting package metadata... Solving environment: done Preparing transaction: done Verifying transaction: done Executing transaction: done Successfully installed ta-lib-0.4.28 prophet-1.1.5 ... Installed kernelspec my_custom_env in /home/datascience/.local/share/jupyter/kernels/ ✅ Custom Conda environment ready! Select 'Python (My Custom ML Env)' from the kernel menu in JupyterLab
☁️ Step 5 — Save Custom Conda to Object Storage (CRITICAL!)
Here is the most important thing beginners miss:
custom Conda environments are stored on your notebook's block volume —
and that block volume is tied to your notebook session.
If you delete the notebook or create a new one,
your custom environment is gone! 😱
The solution: publish your custom Conda environment to OCI Object Storage so anyone (or any new notebook) can restore it instantly.
Imagine you spent a whole day setting up your workbench perfectly — tools sorted, labels on everything, exactly how you like it. Saving to Object Storage is like taking a photo of your workbench so you can recreate it perfectly on a new workbench tomorrow! 📸→🔧
Step 5a: Publish Conda to Object Storage (One Command!)
This command takes your entire custom Conda environment — all Python packages, all settings — packages them into a compressed file, and uploads that file to your OCI Object Storage bucket. Like zipping up your whole toolbox and putting it in cloud storage! 📦☁️
# In JupyterLab Terminal:
# Publish your custom conda environment to Object Storage
# Replace the bucket name and namespace with your values
odsc conda publish \
-s my_custom_env \
--uri oci://data-science-models@your_namespace/conda_envs/my_custom_env/
# You will see a progress bar as it uploads
Output:
Publishing conda environment: my_custom_env Packaging environment... ████████████████████ 100% Uploading to Object Storage... oci://data-science-models@your_namespace/conda_envs/my_custom_env/ ████████████████████████████████████████ 100% (1.2 GB) ✅ Conda environment published successfully! URI: oci://data-science-models@your_namespace/conda_envs/my_custom_env/ Share this URI with your team — they can install it on any notebook! 🎉
Step 5b: Restore Conda Environment on a New Notebook
This command downloads and installs your saved Conda environment from Object Storage onto a brand new notebook session. Any team member can run this and get the exact same Python environment you set up — like handing someone a copy of your perfectly arranged toolbox! 🔑
# On any new notebook session, in the Terminal:
# Install the saved custom environment from Object Storage
odsc conda install \
-uri oci://data-science-models@your_namespace/conda_envs/my_custom_env/
# Activate it after installation
conda activate my_custom_env
echo "✅ Custom environment restored from Object Storage!"
echo " All your packages are available exactly as you saved them."
Step 5c: Automate Conda Backup Using Python
This Python script exports a frozen list of all packages in your Conda environment (name + exact version number) and saves it to Object Storage as a text file. This "requirements snapshot" lets you recreate the exact same environment at any point in the future — like saving the exact recipe of a dish you cooked perfectly! 🍲📄
import oci
import subprocess
import datetime
import io
config = oci.config.from_file()
os_client = oci.object_storage.ObjectStorageClient(config)
namespace = os_client.get_namespace().data
bucket = "data-science-models"
env_name = "my_custom_env"
# Step 1: Export the exact list of installed packages from the conda environment
print(f"📋 Exporting package list from '{env_name}'...")
result = subprocess.run(
["conda", "run", "-n", env_name, "pip", "freeze"],
capture_output=True,
text=True
)
requirements_text = result.stdout
# Step 2: Also export conda-specific packages
conda_result = subprocess.run(
["conda", "env", "export", "-n", env_name, "--no-builds"],
capture_output=True,
text=True
)
conda_yaml_text = conda_result.stdout
# Step 3: Upload both files to Object Storage
timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
req_key = f"conda_envs/backups/{env_name}_requirements_{timestamp}.txt"
yaml_key = f"conda_envs/backups/{env_name}_environment_{timestamp}.yaml"
os_client.put_object(
namespace, bucket, req_key,
requirements_text.encode("utf-8")
)
print(f"✅ requirements.txt uploaded: {req_key}")
os_client.put_object(
namespace, bucket, yaml_key,
conda_yaml_text.encode("utf-8")
)
print(f"✅ environment.yaml uploaded: {yaml_key}")
print(f"\n🎉 Conda environment backup complete!")
print(f" Timestamp : {timestamp}")
print(f" Packages : {len(requirements_text.splitlines())} packages saved")
print(f"\n💡 To restore: run 'conda env create -f {yaml_key}' on any machine")
Output:
📋 Exporting package list from 'my_custom_env'... ✅ requirements.txt uploaded: conda_envs/backups/my_custom_env_requirements_20260411_143022.txt ✅ environment.yaml uploaded: conda_envs/backups/my_custom_env_environment_20260411_143022.yaml 🎉 Conda environment backup complete! Timestamp : 20260411_143022 Packages : 187 packages saved 💡 To restore: run 'conda env create -f environment_20260411_143022.yaml' on any machine
- After installing any new package you want to keep long-term
- Before upgrading any library (backup first, then upgrade)
- Every Friday as part of your weekly routine 📅
- Whenever you share your notebook with a colleague
🌿 Step 6 — Git Integration in OCI Data Science
Code without version control is like writing a document without ever saving it.
One bad edit and everything is lost! 😱
Git tracks every change, lets you go back in time,
and lets your entire team work on the same codebase safely.
Imagine writing a book where you save a new version every time you edit a page. You can go back to any previous version — from two minutes ago or two years ago. You can also let a friend edit their own copy and then merge both copies together cleanly. That is Git! 📖🕰️
Step 6a: Configure Git Inside JupyterLab
These commands set up your identity in Git — your name and email. Every commit (save point) you create will show your name, so your team knows who made which change. This only needs to be done once per notebook session! 👤
# In JupyterLab Terminal: # Set your Git identity (do this once on each new notebook) git config --global user.name "Priya Sharma" git config --global user.email "priya.sharma@mycompany.com" # Confirm settings were saved git config --global --list
Output:
user.name=Priya Sharma user.email=priya.sharma@mycompany.com
Step 6b: Clone a Repository from GitHub / GitLab / OCI DevOps
The
git clone command downloads a copy of an entire code repository from GitHub
(or any other Git server) onto your notebook server.
Like downloading a whole project folder from a shared drive to your local machine! 📥
# Clone from GitHub using HTTPS git clone https://github.com/your-org/churn-prediction-project.git cd churn-prediction-project # Clone from OCI DevOps (Oracle's built-in Git service) git clone https://devops.scmservice.ap-mumbai-1.oci.oraclecloud.com/namespaces/your_namespace/projects/ChurnProject/repositories/churn-models # Check what was downloaded ls -la
Step 6c: Set Up SSH Key for Secure Git Access
Using a password every time you push to GitHub is tedious. SSH keys let you connect securely without typing a password! 🔐
This generates a pair of secret keys (like a lock and a key). You add the public key to GitHub. GitHub then recognises your notebook automatically — no password needed, ever again! 🗝️
# Step 1: Generate an SSH key pair on your notebook ssh-keygen -t ed25519 -C "priya.sharma@mycompany.com" -f ~/.ssh/id_ed25519 -N "" # Step 2: Print your PUBLIC key (copy this and paste it into GitHub → Settings → SSH Keys) cat ~/.ssh/id_ed25519.pub # Step 3: Test the connection to GitHub ssh -T git@github.com
Output:
Generating public/private ed25519 key pair. Your identification has been saved in /home/datascience/.ssh/id_ed25519 Your public key has been saved in /home/datascience/.ssh/id_ed25519.pub --- PUBLIC KEY (copy this to GitHub) --- ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIGrxyz123...example priya.sharma@mycompany.com --- GitHub Connection Test --- Hi priya-sharma! You've successfully authenticated, but GitHub does not provide shell access.
Step 6d: Full Git Workflow Inside JupyterLab
This is the standard daily Git workflow for a data scientist. Think of it as:
Pull latest work from the team → Write your code → Mark your changes → Save a snapshot → Share with the team! Like checking out library books, making notes, and returning them so others can use your notes too! 📚
# ── The Daily Data Scientist Git Workflow ────────────────────────── # Step 1: Always pull the latest work from your team first! git pull origin main # Step 2: Create a new branch for YOUR work (never work directly on main!) git checkout -b feature/add-feature-engineering # Step 3: Do your work — write notebooks, edit scripts, etc. # ... (you write code here) ... # Step 4: Check what files you changed git status # Step 5: Stage the files you want to save (add them to the "staging area") git add notebooks/feature_engineering.ipynb git add src/features.py # Or stage ALL changed files git add . # Step 6: Commit (save a snapshot) with a clear, descriptive message git commit -m "feat: add RFM features for customer churn model - Added Recency, Frequency, Monetary value features - Added 30-day rolling averages for transaction count - Tested on 10K sample — no data leakage detected" # Step 7: Push your branch to GitHub so your team can see it git push origin feature/add-feature-engineering # Step 8: On GitHub, open a Pull Request to merge your branch into main # (do this in the GitHub website — click 'Compare & Pull Request')
- Never commit large data files (CSVs, Parquet, images) to Git — use Object Storage for data!
- Never commit your
~/.oci/configfile or API keys — they are secrets! - Never work directly on the
mainbranch — always create a feature branch - Never use vague commit messages like "fixed stuff" or "changes"
Step 6e: Create a .gitignore for Data Science Projects
A
.gitignore file tells Git:
"Never track these files — even if I accidentally type 'git add .'."
It protects you from accidentally uploading large data files or secret credentials! 🛡️
# .gitignore for OCI Data Science Projects # Create this file in the root of your project: touch .gitignore # ── Data files (store in Object Storage instead!) ── *.csv *.parquet *.json *.xlsx data/ datasets/ raw/ # ── Model files (store as OCI Model Artifacts instead!) ── *.pkl *.joblib *.h5 *.onnx *.pt *.pth models/ # ── OCI Credentials (NEVER commit these!) ── ~/.oci/config *.pem *.key *.env .env* # ── Python junk files ── __pycache__/ *.pyc *.pyo .ipynb_checkpoints/ .pytest_cache/ # ── Conda environment folders (use odsc conda publish instead!) ── conda-envs/ envs/ *.tar.gz # ── Jupyter notebook outputs (optional — keep notebooks clean) ── # *.ipynb # Uncomment this to not track notebooks at all
🏛️ Step 7 — Using OCI DevOps Git (Built-in Oracle Git)
If you prefer to keep everything inside Oracle Cloud (no GitHub, no GitLab), OCI has its own built-in Git service called OCI DevOps — Code Repositories. It is fully private, integrated with IAM, and completely free for your team! 🏛️
Create an OCI DevOps Repository (Console Steps)
- Go to OCI Console → Developer Services → DevOps
- Create a DevOps Project (if you don't have one): name it
DataScienceMLProject - Inside the project, click "Code Repositories"
- Click "Create Repository"
- Name it:
churn-prediction-code - Click "Create" ✅
- Copy the HTTPS clone URL provided — use it with
git clonein your notebook terminal
This code uses the OCI SDK to create a new Git repository in OCI DevOps — all programmatically, no clicking in the console needed. It is like creating a new private folder on Oracle Cloud's internal Git server! 🌿
import oci
config = oci.config.from_file()
devops_client = oci.devops.DevopsClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
# First, get (or create) a DevOps project
devops_project_id = "ocid1.devopsproject.oc1..your_devops_project_id"
# Create a new code repository inside that project
print("🌿 Creating OCI DevOps Git repository...")
repo_response = devops_client.create_repository(
oci.devops.models.CreateRepositoryDetails(
name="churn-prediction-code",
project_id=devops_project_id,
repository_type="HOSTED", # HOSTED = stored in OCI; MIRRORED = synced from GitHub
description="ML notebooks and scripts for customer churn prediction project"
)
)
repo = repo_response.data
print(f"✅ Repository created!")
print(f" Name : {repo.name}")
print(f" SSH URL : {repo.ssh_url}")
print(f" HTTPS URL : {repo.http_url}")
print(f"\n📋 To clone, run:")
print(f" git clone {repo.http_url}")
Output:
🌿 Creating OCI DevOps Git repository... ✅ Repository created! Name : churn-prediction-code SSH URL : ssh://devops.scmservice.ap-mumbai-1.oci.oraclecloud.com/... HTTPS URL : https://devops.scmservice.ap-mumbai-1.oci.oraclecloud.com/... 📋 To clone, run: git clone https://devops.scmservice.ap-mumbai-1.oci.oraclecloud.com/namespaces/...
🏭 Step 8 — Full End-to-End Workflow Example
Let's now put everything together! Here is what a real data science team's daily workflow looks like when using OCI Data Science properly:
DAY IN THE LIFE OF A DATA SCIENTIST ON OCI
─────────────────────────────────────────────────────────────────────────
9:00 AM ── Activate Notebook Session
(OCI Console → Project → Notebook → Activate)
9:05 AM ── Open JupyterLab in browser
Select kernel: "Python (My Custom ML Env)"
9:10 AM ── Git pull — get latest team changes
$ git pull origin main
$ git checkout -b feature/model-v3-hypertuning
9:15 AM ── Load data from Object Storage
df = pd.read_parquet("oci://datalake-curated@ns/sales/")
10:30 AM ── Write / run notebook experiments
(feature engineering, model training, evaluation)
12:00 PM ── Lunch Break!
⚠️ DEACTIVATE the notebook to save costs! 💸
1:00 PM ── Reactivate notebook — resumes exactly where you left off
3:00 PM ── Results look good! Commit and push code to Git
$ git add notebooks/model_v3.ipynb src/train.py
$ git commit -m "feat: hyperparameter tuning improved AUC to 0.94"
$ git push origin feature/model-v3-hypertuning
4:00 PM ── New library needed for SHAP explainability
$ pip install shap
After testing: save updated conda env to Object Storage
$ odsc conda publish -s my_custom_env --uri oci://...
5:00 PM ── Deactivate notebook for the night
✅ Code safe in Git
✅ Data safe in Object Storage
✅ Environment saved in Object Storage
✅ No charges overnight! 🎉
🏆 Step 9 — Best Practices for OCI Data Science
- 📁 One Project per Use Case — Create a separate OCI Data Science Project for each distinct AI use case. Never mix churn prediction and image classification in the same project.
-
💡 Use Tags on Notebook Sessions —
Tag notebooks with
team=datascience,project=churn,env=devso finance can track costs per team and project accurately. - 🔁 Restart Kernel Regularly — Restart your Jupyter kernel and run the notebook from top to bottom regularly. This catches bugs where cells were run out of order.
- 📦 Store Large Data in Object Storage, NOT the notebook block volume — The block volume is for code only. Data goes in Object Storage and is read via OCI ADS or boto3.
- 🔐 Use OCI Vault for secrets — Never hardcode passwords, API keys, or connection strings in your notebooks. Use OCI Vault to store and retrieve secrets safely.
-
📊 Use OCI ADS (Accelerated Data Science SDK) —
OCI provides a high-level Python SDK called
adsthat simplifies loading data from Object Storage, saving models, connecting to ADW, and more.
Quick OCI ADS Example
OCI ADS (Accelerated Data Science) is Oracle's powerful Python helper library pre-installed in every notebook session. This code uses ADS to load data directly from Object Storage into a Pandas DataFrame in just two lines — much easier than writing all the OCI SDK code manually! 🚀
import ads
import pandas as pd
# Authenticate using your notebook session's resource principal
# (No config file needed inside a notebook session!)
ads.set_auth(auth="resource_principal")
# Load data from Object Storage directly into Pandas — 2 lines!
df = pd.read_parquet(
"oci://datalake-curated@your_namespace/sales/year=2026/month=04/",
storage_options={"config": {}}
)
print(f"✅ Loaded {len(df):,} rows from Object Storage")
print(f" Columns: {list(df.columns)}")
df.head()
Output:
✅ Loaded 124,901 rows from Object Storage Columns: ['order_id', 'customer_id', 'product_name', 'amount', 'order_date']
📝 Quick Summary — What We Learned
- Projects → Top-level containers that organise all notebooks, models, and jobs for one use case
- Notebook Sessions → Managed JupyterLab servers — turn on when working, turn off when done to save cost
- Conda Environments → Isolated Python toolboxes — pick pre-built ones or build custom ones
- Save Conda to Object Storage → Use
odsc conda publishto backup custom environments so any notebook can restore them - Git Integration → Version control for all your code — clone, branch, commit, push — daily!
- OCI DevOps Git → Oracle's own private Git service, integrated with IAM — no GitHub account needed
- OCI ADS SDK → High-level Python library inside every notebook to read data, save models, and connect to OCI services easily
Projects keep you organised. Notebooks give you power. Conda keeps your tools tidy.
Git keeps your code safe. Object Storage keeps your data and environments backed up.
Put it all together and prepare a world-class AI lab in the cloud! 🔬✨
Comments
Post a Comment