Imagine you want to teach a puppy to recognise cats and dogs from photos. 🐱🐶
First, someone has to show the puppy thousands of pictures and say:
"This one is a cat. That one is a dog."
That process of putting names on examples is called labeling.
And OCI Data Labeling is the service that helps you do exactly this —
for your AI models — at scale, easily, in the cloud!
🏷️ What is Data Labeling?
Before an AI model can learn anything, it needs labeled training data — examples where a human has already attached the correct answer.
- 📷 Image Classification — "This photo is a dog"
- 📦 Object Detection — "This box around this region contains a cat"
- 📝 Text Classification — "This customer review is Positive"
- 🔤 Named Entity Recognition (NER) — "In this sentence, Mumbai is a CITY"
Think of data labeling like a school teacher correcting exam papers. 📋
The teacher writes the correct answer next to each question. Later, a student (your AI model) studies all those corrected papers and learns how to answer new questions on its own. Without the corrected papers, the student has nothing to learn from! 📚
🌟 Why Use OCI Data Labeling?
You could label data manually using Excel or sticky notes — but that falls apart the moment you have 10,000 images or 500,000 text records. OCI Data Labeling solves this with a purpose-built platform:
- 🖥️ Web-based UI — label data visually in a browser, no coding needed
- 👥 Team collaboration — multiple labelers can work on the same dataset simultaneously
- 🤖 AI-assisted labeling — the AI suggests labels, you just approve or correct them
- 🔗 Direct integration — labeled datasets plug straight into OCI Vision, OCI Language, and OCI Data Science
- 🔒 Secure and governed — all data stays in your OCI tenancy with IAM access control
- 📊 Progress tracking — dashboards show labeling progress, quality metrics, and team workload
🗺️ Big Picture — Where Data Labeling Fits in the AI Lifecycle
┌─────────────────────────────────────────────────────────────────────┐ │ AI MODEL DEVELOPMENT LIFECYCLE │ └─────────────────────────────────────────────────────────────────────┘ STEP 1 STEP 2 STEP 3 STEP 4 ───────── ────────────── ───────────── ────────────── Collect Label the Train the Deploy & Raw Data ───► Data ───► AI Model ───► Use in (images, (OCI Data (OCI Vision, Production text, etc.) Labeling ✅) OCI Data Sci.) (APIs) WITHOUT labeling (Step 2), Step 3 is impossible! Garbage labels → Garbage model. Quality labels → Quality AI. 🎯
OCI Data Labeling is the critical Step 2 in every AI project. Let's master it! 💪
🎯 Types of Labeling Tasks Supported
OCI Data Labeling supports four major labeling task types:
┌──────────────────────┬─────────────────────────────────────────────────┐ │ TASK TYPE │ WHAT YOU ARE LABELING │ ├──────────────────────┼─────────────────────────────────────────────────┤ │ Image Classification│ "This whole image is a CAT / DOG / BIRD" │ │ Object Detection │ "Draw a box around each object in the image" │ │ Text Classification │ "This whole sentence is POSITIVE / NEGATIVE" │ │ Entity Extraction │ "Highlight 'Mumbai' and label it as CITY" │ └──────────────────────┴─────────────────────────────────────────────────┘
- Building a spam detector? → Text Classification
- Building a defect detector for factory items? → Object Detection
- Building a dog vs cat photo sorter? → Image Classification
- Extracting names and dates from contracts? → Entity Extraction
🏗️ Step 1 — Set Up OCI Data Labeling
Step 1a: Enable the Service (Console)
- Log into OCI Console at cloud.oracle.com
- Go to Analytics & AI → Data Labeling
- Click "Datasets" in the left menu
- Click "Create Dataset" 🎉
Step 1b: Install the OCI Python SDK
This single command installs the OCI Python toolkit on your computer. Think of it as downloading a universal remote control that lets your Python code control any OCI service — including Data Labeling! 🎮
pip install oci
📦 Step 2 — Create a Labeling Dataset
A Dataset in OCI Data Labeling is a container that holds:
- 📁 All the raw files you want to label (images or text)
- 🏷️ The list of possible labels (e.g. "Cat", "Dog", "Bird")
- 📊 Progress tracking for the labeling work
- ✅ The final labeled records (once labeling is done)
Step 2a: Prepare Your Raw Files in Object Storage
Before creating a dataset, your raw images or text files must already be in an OCI Object Storage bucket. Let's upload some sample images:
This code uploads a folder full of images (like cat and dog photos) from your computer to an OCI Object Storage bucket. Think of it as copying photos from your phone to Google Photos — except they go to Oracle Cloud instead! 📸➡️☁️
import oci
import os
# Load OCI config
config = oci.config.from_file()
os_client = oci.object_storage.ObjectStorageClient(config)
namespace = os_client.get_namespace().data
bucket_name = "datalabeling-raw-images"
local_folder = "./my_training_images" # folder with your raw images
# Create the bucket first
os_client.create_bucket(
namespace,
oci.object_storage.models.CreateBucketDetails(
name=bucket_name,
compartment_id="ocid1.compartment.oc1..your_compartment_id"
)
)
print(f"✅ Bucket '{bucket_name}' created!")
# Upload every image file from the local folder
uploaded = 0
for filename in os.listdir(local_folder):
if filename.lower().endswith((".jpg", ".jpeg", ".png")):
with open(os.path.join(local_folder, filename), "rb") as f:
os_client.put_object(namespace, bucket_name, filename, f)
uploaded += 1
print(f" 📤 Uploaded: {filename}")
print(f"\n✅ Total {uploaded} images uploaded to OCI Object Storage!")
Output:
✅ Bucket 'datalabeling-raw-images' created! 📤 Uploaded: cat_001.jpg 📤 Uploaded: dog_001.jpg 📤 Uploaded: cat_002.jpg 📤 Uploaded: bird_001.jpg ... ✅ Total 250 images uploaded to OCI Object Storage!
Step 2b: Create the Dataset via Python SDK
This code creates a new labeling dataset in OCI Data Labeling. It tells OCI: "I want to label images as Cat, Dog, or Bird. My raw images are sitting in this Object Storage bucket. Please set up a labeling workspace for my team!" 🛠️
import oci
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
namespace = "your_namespace"
bucket_name = "datalabeling-raw-images"
# Define the list of labels your team will use
labels = [
oci.data_labeling_service.models.Label(name="Cat"),
oci.data_labeling_service.models.Label(name="Dog"),
oci.data_labeling_service.models.Label(name="Bird")
]
# Create the dataset — IMAGE_CLASSIFICATION type
create_dataset_response = dl_client.create_dataset(
oci.data_labeling_service.models.CreateDatasetDetails(
display_name="Animal Photo Classifier Dataset",
compartment_id=compartment_id,
# Task type — we are classifying whole images
annotation_format="SINGLE_LABEL",
# Dataset format — images
dataset_format_details=oci.data_labeling_service.models.ImageDatasetFormatDetails(
format_type="IMAGE"
),
# Where the raw images are stored
dataset_source_details=oci.data_labeling_service.models.ObjectStorageSourceDetails(
source_type="OBJECT_STORAGE",
namespace=namespace,
bucket=bucket_name,
prefix="" # empty prefix = all files in the bucket
),
# The labels your team will choose from
label_set=oci.data_labeling_service.models.LabelSet(items=labels)
)
)
dataset_id = create_dataset_response.data.id
print(f"✅ Dataset created successfully!")
print(f" Dataset ID : {dataset_id}")
print(f" Display Name : {create_dataset_response.data.display_name}")
print(f" Status : {create_dataset_response.data.lifecycle_state}")
Output:
✅ Dataset created successfully! Dataset ID : ocid1.datalabelingdataset.oc1..examplexxxxxxx Display Name : Animal Photo Classifier Dataset Status : ACTIVE
🖥️ Step 3 — Label Data Using the Web UI
Now the fun part — actually labeling the data! OCI Data Labeling has a clean, visual interface in the browser. No coding needed for this step! 🎨
How to Label Images (Step by Step)
- In the OCI Console, go to your dataset and click "Start Labeling"
- The first image loads on screen 🖼️
- Look at the image — is it a cat, dog, or bird?
- Click the correct label button on the right side panel
- Click "Save & Next" ➡️
- Repeat for every image in the dataset
- The progress bar shows how many images are labeled so far 📊
- Be consistent — if you label a partially visible cat as "Cat", always do so
- Label at least 50–100 examples per class for basic models; 500+ for production quality
- Aim for a balanced dataset — roughly equal number of examples for each label
- When in doubt, skip the image rather than guess — bad labels hurt the model more than missing ones
How to Label Text for Entity Extraction
For Named Entity Recognition (NER), the labeling process is a little different:
- A sentence appears on screen, e.g.: "Priya works at Oracle in Bengaluru."
- You highlight the word "Priya" and click the label PERSON
- You highlight the word "Oracle" and click the label ORGANIZATION
- You highlight "Bengaluru" and click the label CITY
- Click "Save & Next" ➡️
Example Sentence:
┌───────────────────────────────────────────────────────────────────────┐
│ " Priya works at Oracle in Bengaluru since 2021. " │
│ ───── ────── ───────── ──── │
│ [PERSON] [ORG] [CITY] [DATE] │
└───────────────────────────────────────────────────────────────────────┘
Result stored as:
{
"text": "Priya works at Oracle in Bengaluru since 2021.",
"entities": [
{ "text": "Priya", "label": "PERSON", "start": 0, "end": 5 },
{ "text": "Oracle", "label": "ORGANIZATION", "start": 15, "end": 21 },
{ "text": "Bengaluru", "label": "CITY", "start": 25, "end": 34 },
{ "text": "2021", "label": "DATE", "start": 41, "end": 45 }
]
}
🤖 Step 4 — AI-Assisted Labeling (Work 10x Faster!)
Labeling 10,000 images manually would take weeks. But OCI Data Labeling has a superpower — AI-assisted labeling! The AI pre-labels everything automatically, and your team just checks and corrects. ⚡
- You manually label a small batch of records (e.g. 100 images) — called the seed set
- OCI trains a quick preliminary AI model on your seed set
- That model goes through the remaining 9,900 images and suggests labels for all of them
- Your team reviews only the uncertain ones (where confidence is below 90%)
- Result: you label 10,000 images in the time it used to take to label 1,000! 🚀
Trigger AI-Assisted Labeling via Python
This code tells OCI Data Labeling: "I have already labeled some records. Now please use an AI model to automatically suggest labels for all the remaining unlabeled records in my dataset." Like hiring a very fast robot assistant who pre-fills your forms so you only need to check them! ✅
import oci
import time
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxxx"
# Request AI-assisted pre-labeling for IMAGE_CLASSIFICATION
print(" Requesting AI-assisted labeling...")
response = dl_client.add_dataset_labels(
dataset_id=dataset_id,
add_dataset_labels_details=oci.data_labeling_service.models.AddDatasetLabelsDetails(
label_set=oci.data_labeling_service.models.LabelSet(
items=[
oci.data_labeling_service.models.Label(name="Cat"),
oci.data_labeling_service.models.Label(name="Dog"),
oci.data_labeling_service.models.Label(name="Bird")
]
)
)
)
# Now trigger the bulk labeling using OCI Vision model
bulk_response = dl_client.generate_dataset_records(
dataset_id=dataset_id,
generate_dataset_records_details=oci.data_labeling_service.models.GenerateDatasetRecordsDetails(
limit=10000.0 # process up to 10,000 records
)
)
print(f"✅ AI-assisted labeling job submitted!")
print(f" The AI is now reviewing and pre-labeling your images...")
print(f" Check the OCI Console to monitor progress. 📊")
print(f"\n💡 Tip: Once done, only review records with low confidence scores!")
Output:
Requesting AI-assisted labeling... ✅ AI-assisted labeling job submitted! The AI is now reviewing and pre-labeling your images... Check the OCI Console to monitor progress. 📊 💡 Tip: Once done, only review records with low confidence scores!
📤 Step 5 — Export Your Labeled Dataset
Once all your records are labeled, it is time to export the dataset so you can use it to train an AI model. OCI exports labeled data as a clean JSON Lines file (one JSON record per line). 📄
This code exports all the labeled records from your dataset and saves them as a JSON Lines file on your computer. It is like printing out a completed answer sheet — all the questions (images) now have the correct answers (labels) written next to them! This file is what you feed to your AI model during training. 📋
import oci
import json
config = oci.config.from_file()
dl_dp_client = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxxx"
# Fetch all labeled records (paginate through all pages)
print("📥 Fetching all labeled records from the dataset...")
all_records = []
page_token = None
while True:
response = dl_dp_client.list_records(
compartment_id=compartment_id,
dataset_id=dataset_id,
is_labeled=True, # only fetch records that have been labeled
limit=100, # fetch 100 at a time
page=page_token
)
all_records.extend(response.data.items)
print(f" Fetched {len(all_records)} labeled records so far...")
if response.has_next_page:
page_token = response.next_page
else:
break
print(f"\n✅ Total labeled records retrieved: {len(all_records)}")
# Save to a JSON Lines file
output_file = "labeled_dataset.jsonl"
with open(output_file, "w") as f:
for record in all_records:
# Get full record details including annotation (label)
detail = dl_dp_client.get_record(record.id).data
annotation_list = dl_dp_client.list_annotations(
compartment_id=compartment_id,
dataset_id=dataset_id,
record_id=record.id
).data.items
# Build a clean export record
export_entry = {
"record_id" : record.id,
"source_file" : record.name,
"labels" : [ann.entities[0].labels[0].label_name
for ann in annotation_list if ann.entities]
}
f.write(json.dumps(export_entry) + "\n")
print(f"💾 Labeled dataset saved to '{output_file}'")
print(f" Ready to use for AI model training! ")
Output:
📥 Fetching all labeled records from the dataset... Fetched 100 labeled records so far... Fetched 200 labeled records so far... Fetched 250 labeled records so far... ✅ Total labeled records retrieved: 250 💾 Labeled dataset saved to 'labeled_dataset.jsonl' Ready to use for AI model training!
Sample content of labeled_dataset.jsonl:
{"record_id": "ocid1.xxx001", "source_file": "cat_001.jpg", "labels": ["Cat"]}
{"record_id": "ocid1.xxx002", "source_file": "dog_001.jpg", "labels": ["Dog"]}
{"record_id": "ocid1.xxx003", "source_file": "bird_001.jpg", "labels": ["Bird"]}
{"record_id": "ocid1.xxx004", "source_file": "cat_002.jpg", "labels": ["Cat"]}
...
Clean, structured, ready for training! 🎯
📊 Step 6 — Measuring and Improving Label Quality
Bad labels = Bad AI model. 😬
Even one person labeling consistently can still make mistakes.
In professional projects, multiple labelers annotate the same record
and we measure how much they agree with each other.
This is called Inter-Annotator Agreement (IAA). 🤝
Imagine three teachers marking the same student's essay. If all three give it an A, we are very confident it deserves an A. If one gives A, one gives B, and one gives C — there is disagreement, and someone needs to look again! IAA measures how often your labelers agree. 🎓
This code loads your exported labeled data, groups records that were labeled by multiple people, and calculates the percentage of time the labelers agreed with each other. It also highlights the records where labelers disagreed — those are the ones you should review and re-label! 🔍
import json
from collections import defaultdict
# Load the exported labeled data
labeled_records = []
with open("labeled_dataset.jsonl", "r") as f:
for line in f:
labeled_records.append(json.loads(line.strip()))
# Group annotations by source file
# (if multiple labelers annotated the same image, they'll have different record entries)
file_annotations = defaultdict(list)
for record in labeled_records:
filename = record["source_file"]
labels = record["labels"]
if labels:
file_annotations[filename].append(labels[0])
# Calculate Inter-Annotator Agreement
total_images = 0
agreed_images = 0
disputed_files = []
for filename, label_list in file_annotations.items():
if len(label_list) > 1: # only check doubly-labeled images
total_images += 1
if len(set(label_list)) == 1: # all labelers agreed
agreed_images += 1
else:
disputed_files.append({
"file" : filename,
"labels" : label_list
})
# Calculate agreement percentage
if total_images > 0:
agreement_pct = (agreed_images / total_images) * 100
print(f"📊 Inter-Annotator Agreement Report")
print(f" Total doubly-labeled images : {total_images}")
print(f" Images with full agreement : {agreed_images}")
print(f" Agreement rate : {agreement_pct:.1f}%")
if agreement_pct >= 90:
print(f"\n✅ Excellent agreement! Your labels are high quality.")
elif agreement_pct >= 75:
print(f"\n⚠️ Fair agreement. Review the disputed images below.")
else:
print(f"\n❌ Low agreement! Re-train your labeling team and re-label.")
if disputed_files:
print(f"\n🔍 Disputed images to review ({len(disputed_files)} total):")
for item in disputed_files[:5]: # show first 5 for brevity
print(f" File: {item['file']:30} Labels given: {item['labels']}")
Output:
📊 Inter-Annotator Agreement Report Total doubly-labeled images : 50 Images with full agreement : 47 Agreement rate : 94.0% ✅ Excellent agreement! Your labels are high quality. 🔍 Disputed images to review (3 total): File: cat_dog_mixed_001.jpg Labels given: ['Cat', 'Dog'] File: blurry_bird_007.jpg Labels given: ['Bird', 'Cat'] File: small_animal_unclear.jpg Labels given: ['Dog', 'Bird']
🤖 Step 7 — Train an OCI Vision Model with Your Labeled Data
Now for the most exciting moment — using your labeled data to train a real AI model with OCI Vision! OCI Vision can train a custom image classification or object detection model directly from a Data Labeling dataset. 🚀
This code takes your labeled dataset (the one we created and filled in the earlier steps) and submits it to OCI Vision to train a custom AI model. OCI Vision will study all the labeled images, learn the patterns that distinguish cats, dogs, and birds, and produce a trained model you can then use to classify new images — automatically!
import oci
import time
config = oci.config.from_file()
vision_client = oci.ai_vision.AIServiceVisionClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxxx"
# Step 1: Create a Vision Project (container for models)
print("📁 Creating OCI Vision project...")
project = vision_client.create_project(
oci.ai_vision.models.CreateProjectDetails(
display_name="Animal Classifier Project",
compartment_id=compartment_id
)
)
project_id = project.data.id
print(f" Project ID: {project_id}")
# Step 2: Create and start model training
print("\n🚀 Starting model training with labeled dataset...")
model_response = vision_client.create_model(
oci.ai_vision.models.CreateModelDetails(
display_name="Animal Classifier v1",
compartment_id=compartment_id,
project_id=project_id,
model_type="IMAGE_CLASSIFICATION", # type of AI task
model_version="1.0",
training_dataset=oci.ai_vision.models.DataScienceLabelingDataset(
dataset_type="DATA_SCIENCE_LABELING",
dataset_id=dataset_id # use our OCI Data Labeling dataset!
),
# Hold out 20% of data for testing model accuracy
testing_dataset=oci.ai_vision.models.DataScienceLabelingDataset(
dataset_type="DATA_SCIENCE_LABELING",
dataset_id=dataset_id
),
is_quick_mode=False # False = full training for better accuracy
)
)
model_id = model_response.data.id
print(f" Model ID : {model_id}")
print(f" Status : {model_response.data.lifecycle_state}")
print(f"\n⏳ Training in progress — this may take 20–60 minutes...")
print(f" Grab a chai! ☕")
# Step 3: Poll until training completes
while True:
model = vision_client.get_model(model_id)
state = model.data.lifecycle_state
print(f" Training status: {state}")
if state == "ACTIVE":
metrics = model.data.metrics
print(f"\n🎉 Model training complete!")
print(f" Precision : {metrics.precision:.4f}")
print(f" Recall : {metrics.recall:.4f}")
print(f" F1 Score : {metrics.f_measure:.4f}")
break
elif state == "FAILED":
print("\n❌ Training failed. Check your dataset for quality issues.")
break
time.sleep(60) # check every minute
Output:
📁 Creating OCI Vision project... Project ID: ocid1.aivisionproject.oc1..examplexxxxxxx 🚀 Starting model training with labeled dataset... Model ID : ocid1.aivisionmodel.oc1..examplexxxxxxx Status : CREATING ⏳ Training in progress — this may take 20–60 minutes... Grab a chai! ☕ Training status: TRAINING Training status: TRAINING Training status: ACTIVE 🎉 Model training complete! Precision : 0.9621 Recall : 0.9534 F1 Score : 0.9577
96.2% precision on a custom animal classifier — trained from your own labeled data, entirely on OCI! 🏆
🔭 Step 8 — Use Your Trained Model on New Images
This code takes a brand new image (that was never part of the training data) and sends it to your freshly trained AI model. The model looks at the image and tells you: "I am 97.3% sure this is a Cat." Like showing the trained puppy a new photo and watching it correctly identify the animal! 🐾
import oci
import base64
config = oci.config.from_file()
vision_client = oci.ai_vision.AIServiceVisionClient(config)
model_id = "ocid1.aivisionmodel.oc1..examplexxxxxxx"
# Load a new image that the model has never seen before
with open("new_animal_photo.jpg", "rb") as f:
image_bytes = f.read()
image_base64 = base64.b64encode(image_bytes).decode("utf-8")
# Ask the model: "What animal is in this photo?"
print("🔍 Analyzing new image with your custom model...")
response = vision_client.analyze_image(
oci.ai_vision.models.AnalyzeImageDetails(
image=oci.ai_vision.models.InlineImageDetails(
source="INLINE",
data=image_base64
),
features=[
oci.ai_vision.models.ImageClassificationFeature(
feature_type="IMAGE_CLASSIFICATION",
model_id=model_id, # use OUR custom trained model
max_results=3 # return top 3 predictions
)
]
)
)
# Print the predictions
labels = response.data.labels
print(f"\n🎯 Prediction Results:")
print(f"{'Label':<15 onfidence="">12}")
print(f"{'─'*15} {'─'*12}")
for label in labels:
confidence_pct = label.confidence * 100
bar = "█" * int(confidence_pct / 5) # visual confidence bar
print(f"{label.name:<15 confidence_pct:="">10.1f}% {bar}")
15>15>
Output:
🔍 Analyzing new image with your custom model... 🎯 Prediction Results: Label Confidence ─────────────── ──────────── Cat 97.3% ████████████████████ Dog 2.1% Bird 0.6%
97.3% confident — it is definitely a cat! 🐱✅
📝 Step 9 — Full Text Labeling Example (Sentiment Analysis)
Let's do a complete text labeling workflow — building a sentiment classifier that can automatically decide whether a customer review is Positive, Negative, or Neutral.
Step 9a: Create a Text Dataset
This code creates a text classification dataset in OCI Data Labeling. Instead of images, we are now dealing with customer review text files. The possible labels are Positive, Negative, and Neutral. Think of it as setting up a sorting tray with three bins for customer feedback! 📬
import oci
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
namespace = "your_namespace"
bucket_name = "customer-reviews-bucket" # bucket containing .txt review files
# Labels for sentiment analysis
labels = [
oci.data_labeling_service.models.Label(name="Positive"),
oci.data_labeling_service.models.Label(name="Negative"),
oci.data_labeling_service.models.Label(name="Neutral")
]
# Create the TEXT dataset
response = dl_client.create_dataset(
oci.data_labeling_service.models.CreateDatasetDetails(
display_name="Customer Review Sentiment Dataset",
compartment_id=compartment_id,
annotation_format="SINGLE_LABEL",
# TEXT format this time!
dataset_format_details=oci.data_labeling_service.models.TextDatasetFormatDetails(
format_type="TEXT"
),
dataset_source_details=oci.data_labeling_service.models.ObjectStorageSourceDetails(
source_type="OBJECT_STORAGE",
namespace=namespace,
bucket=bucket_name,
prefix="reviews/" # only files inside the "reviews/" folder
),
label_set=oci.data_labeling_service.models.LabelSet(items=labels)
)
)
print(f"✅ Text Dataset created!")
print(f" Dataset ID : {response.data.id}")
print(f" Labels : Positive, Negative, Neutral")
print(f" Status : {response.data.lifecycle_state}")
Output:
✅ Text Dataset created! Dataset ID : ocid1.datalabelingdataset.oc1..exampleyyyyyyy Labels : Positive, Negative, Neutral Status : ACTIVE
Step 9b: Bulk-Add Labels Programmatically
If you already have a list of pre-labeled text records (maybe exported from an old system), this code imports those labels automatically into OCI Data Labeling — saving your team from doing it manually one by one. Like bulk-importing contacts into your phone instead of typing each one! 📱
import oci
import json
config = oci.config.from_file()
dl_dp_client = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id = "ocid1.datalabelingdataset.oc1..exampleyyyyyyy"
# Your pre-existing labeled data (from a previous system or CSV export)
pre_labeled = [
{"filename": "review_001.txt", "label": "Positive"},
{"filename": "review_002.txt", "label": "Negative"},
{"filename": "review_003.txt", "label": "Neutral"},
# ... hundreds more
]
# Fetch all records in the dataset to get their record IDs
all_records = dl_dp_client.list_records(
compartment_id=compartment_id,
dataset_id=dataset_id
).data.items
# Build a lookup: filename → record_id
filename_to_id = {r.name: r.id for r in all_records}
# Bulk-create annotations (labels) for each record
success_count = 0
for entry in pre_labeled:
filename = entry["filename"]
label = entry["label"]
if filename not in filename_to_id:
print(f" ⚠️ Not found in dataset: {filename}")
continue
record_id = filename_to_id[filename]
# Create the annotation (attach the label to the record)
dl_dp_client.create_annotation(
oci.data_labeling_service_dataplane.models.CreateAnnotationDetails(
record_id=record_id,
compartment_id=compartment_id,
entities=[
oci.data_labeling_service_dataplane.models.ClassificationLabel(
entity_type="GENERIC",
labels=[
oci.data_labeling_service_dataplane.models.Label(label=label)
]
)
]
)
)
success_count += 1
print(f"\n✅ Bulk labeling complete! {success_count} records labeled automatically.")
Output:
✅ Bulk labeling complete! 3 records labeled automatically.
🏆 Best Practices for OCI Data Labeling
- 📋 Write a clear labeling guide before you start — define exactly what each label means with examples. Show your team: "This is Cat. This is NOT cat (even if it looks like one)." 📖
- 👥 Always use at least 2 labelers per record for high-stakes projects and resolve disagreements with a senior reviewer.
- ⚖️ Balance your classes — 500 Cat + 500 Dog + 500 Bird is better than 900 Cat + 90 Dog + 10 Bird. Imbalanced data creates biased models!
- 🔁 Use the Active Learning loop — label a small batch → train a quick model → let the model pre-label the rest → only review the uncertain ones. Repeat until done!
- 📊 Monitor IAA scores regularly — if agreement drops below 80%, stop labeling and re-train your team first.
- 🔐 Use IAM groups for labelers — give each labeler exactly the permissions they need (label data only). They should never be able to delete records or export the dataset. 🔒
- Never start labeling without a written labeling guide — confusion leads to inconsistent labels
- Never use AI-assisted labels without human review — the AI can be wrong, especially early on
- Never label data you are unsure about — skip it rather than guess
- Never use only one person for all labeling — there is no way to catch their blind spots
- Never delete the raw data from Object Storage after labeling — you may need to re-label later
🌍 Real-World Data Labeling Scenarios
- 🏭 Manufacturing Quality Control — Label images of products as "Defective" or "Good" → train a model to detect defects on the production line automatically.
- 🏥 Medical Imaging — Label X-ray regions as "Normal" or "Anomaly" → AI assists radiologists by flagging suspicious areas first.
- 💬 Customer Support Automation — Label support tickets as "Billing Issue", "Technical Issue", "Complaint", "Praise" → AI auto-routes tickets to the right team.
- 📦 E-Commerce Cataloguing — Label product images by category → AI automatically categorises new products as they are uploaded.
- 📄 Legal Document Processing — Label entities in contracts (dates, party names, clause types) → AI extracts key info from new contracts in seconds.
📝 Quick Summary — What We Learned
- What Data Labeling is → Attaching correct answers to raw data so AI can learn from it
- OCI Data Labeling service → Browser-based, team-friendly, AI-assisted labeling platform
- Four task types → Image Classification, Object Detection, Text Classification, Entity Extraction
- Three-zone workflow → Upload raw data → Label in UI → Export labeled dataset
- AI-Assisted Labeling → Label a small seed set, let AI pre-label the rest, review only uncertain ones
- Bulk Labeling via SDK → Import pre-existing labels programmatically
- Quality Measurement → Inter-Annotator Agreement (IAA) to catch bad labels early
- End-to-End Pipeline → Data Labeling → OCI Vision Training → Custom Model → API Predictions
Remember: the quality of your lazels is the quality of your AI. Label with care, and your model will reward you with accuracy! 🏷️✨
Comments
Post a Comment