Skip to main content

OCI Data Labeling

Calculating read time…

Imagine you want to teach a puppy to recognise cats and dogs from photos. 🐱🐶
First, someone has to show the puppy thousands of pictures and say: "This one is a cat. That one is a dog."


That process of putting names on examples is called labeling.
And OCI Data Labeling is the service that helps you do exactly this — for your AI models — at scale, easily, in the cloud!




🏷️ What is Data Labeling?

Before an AI model can learn anything, it needs labeled training data — examples where a human has already attached the correct answer.

  • 📷 Image Classification — "This photo is a dog"
  • 📦 Object Detection — "This box around this region contains a cat"
  • 📝 Text Classification — "This customer review is Positive"
  • 🔤 Named Entity Recognition (NER) — "In this sentence, Mumbai is a CITY"
💡 Real-World Analogy:

Think of data labeling like a school teacher correcting exam papers. 📋
The teacher writes the correct answer next to each question. Later, a student (your AI model) studies all those corrected papers and learns how to answer new questions on its own. Without the corrected papers, the student has nothing to learn from! 📚

🌟 Why Use OCI Data Labeling?

You could label data manually using Excel or sticky notes — but that falls apart the moment you have 10,000 images or 500,000 text records. OCI Data Labeling solves this with a purpose-built platform:

  • 🖥️ Web-based UI — label data visually in a browser, no coding needed
  • 👥 Team collaboration — multiple labelers can work on the same dataset simultaneously
  • 🤖 AI-assisted labeling — the AI suggests labels, you just approve or correct them
  • 🔗 Direct integration — labeled datasets plug straight into OCI Vision, OCI Language, and OCI Data Science
  • 🔒 Secure and governed — all data stays in your OCI tenancy with IAM access control
  • 📊 Progress tracking — dashboards show labeling progress, quality metrics, and team workload

🗺️ Big Picture — Where Data Labeling Fits in the AI Lifecycle

  ┌─────────────────────────────────────────────────────────────────────┐
  │                    AI MODEL DEVELOPMENT LIFECYCLE                   │
  └─────────────────────────────────────────────────────────────────────┘

  STEP 1           STEP 2              STEP 3             STEP 4
  ─────────        ──────────────      ─────────────      ──────────────
  Collect          Label the           Train the          Deploy &
  Raw Data   ───►  Data          ───►  AI Model    ───►  Use in
  (images,         (OCI Data           (OCI Vision,       Production
  text, etc.)      Labeling ✅)        OCI Data Sci.)     (APIs)

  WITHOUT labeling (Step 2), Step 3 is impossible!
  Garbage labels → Garbage model. Quality labels → Quality AI. 🎯

OCI Data Labeling is the critical Step 2 in every AI project. Let's master it! 💪

🎯 Types of Labeling Tasks Supported

OCI Data Labeling supports four major labeling task types:

  ┌──────────────────────┬─────────────────────────────────────────────────┐
  │  TASK TYPE           │  WHAT YOU ARE LABELING                         │
  ├──────────────────────┼─────────────────────────────────────────────────┤
  │  Image Classification│  "This whole image is a CAT / DOG / BIRD"      │
  │  Object Detection    │  "Draw a box around each object in the image"   │
  │  Text Classification │  "This whole sentence is POSITIVE / NEGATIVE"  │
  │  Entity Extraction   │  "Highlight 'Mumbai' and label it as CITY"      │
  └──────────────────────┴─────────────────────────────────────────────────┘
💡 Which task type should I choose?

  • Building a spam detector? → Text Classification
  • Building a defect detector for factory items? → Object Detection
  • Building a dog vs cat photo sorter? → Image Classification
  • Extracting names and dates from contracts? → Entity Extraction

🏗️ Step 1 — Set Up OCI Data Labeling

Step 1a: Enable the Service (Console)

  • Log into OCI Console at cloud.oracle.com
  • Go to Analytics & AI → Data Labeling
  • Click "Datasets" in the left menu
  • Click "Create Dataset" 🎉

Step 1b: Install the OCI Python SDK

📌 What does this command do?

This single command installs the OCI Python toolkit on your computer. Think of it as downloading a universal remote control that lets your Python code control any OCI service — including Data Labeling! 🎮
pip install oci

📦 Step 2 — Create a Labeling Dataset

A Dataset in OCI Data Labeling is a container that holds:

  • 📁 All the raw files you want to label (images or text)
  • 🏷️ The list of possible labels (e.g. "Cat", "Dog", "Bird")
  • 📊 Progress tracking for the labeling work
  • ✅ The final labeled records (once labeling is done)

Step 2a: Prepare Your Raw Files in Object Storage

Before creating a dataset, your raw images or text files must already be in an OCI Object Storage bucket. Let's upload some sample images:

📌 What does the code below do?

This code uploads a folder full of images (like cat and dog photos) from your computer to an OCI Object Storage bucket. Think of it as copying photos from your phone to Google Photos — except they go to Oracle Cloud instead! 📸➡️☁️
import oci
import os

# Load OCI config
config    = oci.config.from_file()
os_client = oci.object_storage.ObjectStorageClient(config)
namespace = os_client.get_namespace().data

bucket_name  = "datalabeling-raw-images"
local_folder = "./my_training_images"       # folder with your raw images

# Create the bucket first
os_client.create_bucket(
    namespace,
    oci.object_storage.models.CreateBucketDetails(
        name=bucket_name,
        compartment_id="ocid1.compartment.oc1..your_compartment_id"
    )
)
print(f"✅ Bucket '{bucket_name}' created!")

# Upload every image file from the local folder
uploaded = 0
for filename in os.listdir(local_folder):
    if filename.lower().endswith((".jpg", ".jpeg", ".png")):
        with open(os.path.join(local_folder, filename), "rb") as f:
            os_client.put_object(namespace, bucket_name, filename, f)
        uploaded += 1
        print(f"   📤 Uploaded: {filename}")

print(f"\n✅ Total {uploaded} images uploaded to OCI Object Storage!")

Output:

✅ Bucket 'datalabeling-raw-images' created!
   📤 Uploaded: cat_001.jpg
   📤 Uploaded: dog_001.jpg
   📤 Uploaded: cat_002.jpg
   📤 Uploaded: bird_001.jpg
   ...

✅ Total 250 images uploaded to OCI Object Storage!

Step 2b: Create the Dataset via Python SDK

📌 What does the code below do?

This code creates a new labeling dataset in OCI Data Labeling. It tells OCI: "I want to label images as Cat, Dog, or Bird. My raw images are sitting in this Object Storage bucket. Please set up a labeling workspace for my team!" 🛠️
import oci

config          = oci.config.from_file()
dl_client       = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id  = "ocid1.compartment.oc1..your_compartment_id"
namespace       = "your_namespace"
bucket_name     = "datalabeling-raw-images"

# Define the list of labels your team will use
labels = [
    oci.data_labeling_service.models.Label(name="Cat"),
    oci.data_labeling_service.models.Label(name="Dog"),
    oci.data_labeling_service.models.Label(name="Bird")
]

# Create the dataset — IMAGE_CLASSIFICATION type
create_dataset_response = dl_client.create_dataset(
    oci.data_labeling_service.models.CreateDatasetDetails(
        display_name="Animal Photo Classifier Dataset",
        compartment_id=compartment_id,

        # Task type — we are classifying whole images
        annotation_format="SINGLE_LABEL",

        # Dataset format — images
        dataset_format_details=oci.data_labeling_service.models.ImageDatasetFormatDetails(
            format_type="IMAGE"
        ),

        # Where the raw images are stored
        dataset_source_details=oci.data_labeling_service.models.ObjectStorageSourceDetails(
            source_type="OBJECT_STORAGE",
            namespace=namespace,
            bucket=bucket_name,
            prefix=""          # empty prefix = all files in the bucket
        ),

        # The labels your team will choose from
        label_set=oci.data_labeling_service.models.LabelSet(items=labels)
    )
)

dataset_id = create_dataset_response.data.id
print(f"✅ Dataset created successfully!")
print(f"   Dataset ID   : {dataset_id}")
print(f"   Display Name : {create_dataset_response.data.display_name}")
print(f"   Status       : {create_dataset_response.data.lifecycle_state}")

Output:

✅ Dataset created successfully!
   Dataset ID   : ocid1.datalabelingdataset.oc1..examplexxxxxxx
   Display Name : Animal Photo Classifier Dataset
   Status       : ACTIVE

🖥️ Step 3 — Label Data Using the Web UI

Now the fun part — actually labeling the data! OCI Data Labeling has a clean, visual interface in the browser. No coding needed for this step! 🎨

How to Label Images (Step by Step)

  • In the OCI Console, go to your dataset and click "Start Labeling"
  • The first image loads on screen 🖼️
  • Look at the image — is it a cat, dog, or bird?
  • Click the correct label button on the right side panel
  • Click "Save & Next" ➡️
  • Repeat for every image in the dataset
  • The progress bar shows how many images are labeled so far 📊
✅ Labeling Tips for Better AI Models:

  • Be consistent — if you label a partially visible cat as "Cat", always do so
  • Label at least 50–100 examples per class for basic models; 500+ for production quality
  • Aim for a balanced dataset — roughly equal number of examples for each label
  • When in doubt, skip the image rather than guess — bad labels hurt the model more than missing ones

How to Label Text for Entity Extraction

For Named Entity Recognition (NER), the labeling process is a little different:

  • A sentence appears on screen, e.g.: "Priya works at Oracle in Bengaluru."
  • You highlight the word "Priya" and click the label PERSON
  • You highlight the word "Oracle" and click the label ORGANIZATION
  • You highlight "Bengaluru" and click the label CITY
  • Click "Save & Next" ➡️
  Example Sentence:
  ┌───────────────────────────────────────────────────────────────────────┐
  │  " Priya works at Oracle in Bengaluru since 2021. "                  │
  │     ─────              ──────    ─────────        ────               │
  │    [PERSON]          [ORG]      [CITY]           [DATE]              │
  └───────────────────────────────────────────────────────────────────────┘

  Result stored as:
  {
    "text": "Priya works at Oracle in Bengaluru since 2021.",
    "entities": [
      { "text": "Priya",     "label": "PERSON",       "start": 0,  "end": 5  },
      { "text": "Oracle",    "label": "ORGANIZATION", "start": 15, "end": 21 },
      { "text": "Bengaluru", "label": "CITY",         "start": 25, "end": 34 },
      { "text": "2021",      "label": "DATE",         "start": 41, "end": 45 }
    ]
  }

🤖 Step 4 — AI-Assisted Labeling (Work 10x Faster!)

Labeling 10,000 images manually would take weeks. But OCI Data Labeling has a superpower — AI-assisted labeling! The AI pre-labels everything automatically, and your team just checks and corrects. ⚡

💡 How AI-Assisted Labeling Works:

  1. You manually label a small batch of records (e.g. 100 images) — called the seed set
  2. OCI trains a quick preliminary AI model on your seed set
  3. That model goes through the remaining 9,900 images and suggests labels for all of them
  4. Your team reviews only the uncertain ones (where confidence is below 90%)
  5. Result: you label 10,000 images in the time it used to take to label 1,000! 🚀

Trigger AI-Assisted Labeling via Python

📌 What does the code below do?

This code tells OCI Data Labeling: "I have already labeled some records. Now please use an AI model to automatically suggest labels for all the remaining unlabeled records in my dataset." Like hiring a very fast robot assistant who pre-fills your forms so you only need to check them! ✅
import oci
import time

config         = oci.config.from_file()
dl_client      = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id     = "ocid1.datalabelingdataset.oc1..examplexxxxxxx"

# Request AI-assisted pre-labeling for IMAGE_CLASSIFICATION
print(" Requesting AI-assisted labeling...")
response = dl_client.add_dataset_labels(
    dataset_id=dataset_id,
    add_dataset_labels_details=oci.data_labeling_service.models.AddDatasetLabelsDetails(
        label_set=oci.data_labeling_service.models.LabelSet(
            items=[
                oci.data_labeling_service.models.Label(name="Cat"),
                oci.data_labeling_service.models.Label(name="Dog"),
                oci.data_labeling_service.models.Label(name="Bird")
            ]
        )
    )
)

# Now trigger the bulk labeling using OCI Vision model
bulk_response = dl_client.generate_dataset_records(
    dataset_id=dataset_id,
    generate_dataset_records_details=oci.data_labeling_service.models.GenerateDatasetRecordsDetails(
        limit=10000.0    # process up to 10,000 records
    )
)

print(f"✅ AI-assisted labeling job submitted!")
print(f"   The AI is now reviewing and pre-labeling your images...")
print(f"   Check the OCI Console to monitor progress. 📊")
print(f"\n💡 Tip: Once done, only review records with low confidence scores!")

Output:

 Requesting AI-assisted labeling...
✅ AI-assisted labeling job submitted!
   The AI is now reviewing and pre-labeling your images...
   Check the OCI Console to monitor progress. 📊

💡 Tip: Once done, only review records with low confidence scores!

📤 Step 5 — Export Your Labeled Dataset

Once all your records are labeled, it is time to export the dataset so you can use it to train an AI model. OCI exports labeled data as a clean JSON Lines file (one JSON record per line). 📄

📌 What does the code below do?

This code exports all the labeled records from your dataset and saves them as a JSON Lines file on your computer. It is like printing out a completed answer sheet — all the questions (images) now have the correct answers (labels) written next to them! This file is what you feed to your AI model during training. 📋
import oci
import json

config          = oci.config.from_file()
dl_dp_client    = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id  = "ocid1.compartment.oc1..your_compartment_id"
dataset_id      = "ocid1.datalabelingdataset.oc1..examplexxxxxxx"

# Fetch all labeled records (paginate through all pages)
print("📥 Fetching all labeled records from the dataset...")
all_records = []
page_token  = None

while True:
    response = dl_dp_client.list_records(
        compartment_id=compartment_id,
        dataset_id=dataset_id,
        is_labeled=True,           # only fetch records that have been labeled
        limit=100,                 # fetch 100 at a time
        page=page_token
    )

    all_records.extend(response.data.items)
    print(f"   Fetched {len(all_records)} labeled records so far...")

    if response.has_next_page:
        page_token = response.next_page
    else:
        break

print(f"\n✅ Total labeled records retrieved: {len(all_records)}")

# Save to a JSON Lines file
output_file = "labeled_dataset.jsonl"
with open(output_file, "w") as f:
    for record in all_records:
        # Get full record details including annotation (label)
        detail = dl_dp_client.get_record(record.id).data
        annotation_list = dl_dp_client.list_annotations(
            compartment_id=compartment_id,
            dataset_id=dataset_id,
            record_id=record.id
        ).data.items

        # Build a clean export record
        export_entry = {
            "record_id"   : record.id,
            "source_file" : record.name,
            "labels"      : [ann.entities[0].labels[0].label_name
                             for ann in annotation_list if ann.entities]
        }
        f.write(json.dumps(export_entry) + "\n")

print(f"💾 Labeled dataset saved to '{output_file}'")
print(f"   Ready to use for AI model training! ")

Output:

📥 Fetching all labeled records from the dataset...
   Fetched 100 labeled records so far...
   Fetched 200 labeled records so far...
   Fetched 250 labeled records so far...

✅ Total labeled records retrieved: 250

💾 Labeled dataset saved to 'labeled_dataset.jsonl'
   Ready to use for AI model training! 

Sample content of labeled_dataset.jsonl:

{"record_id": "ocid1.xxx001", "source_file": "cat_001.jpg",  "labels": ["Cat"]}
{"record_id": "ocid1.xxx002", "source_file": "dog_001.jpg",  "labels": ["Dog"]}
{"record_id": "ocid1.xxx003", "source_file": "bird_001.jpg", "labels": ["Bird"]}
{"record_id": "ocid1.xxx004", "source_file": "cat_002.jpg",  "labels": ["Cat"]}
...

Clean, structured, ready for training! 🎯

📊 Step 6 — Measuring and Improving Label Quality

Bad labels = Bad AI model. 😬
Even one person labeling consistently can still make mistakes. In professional projects, multiple labelers annotate the same record and we measure how much they agree with each other. This is called Inter-Annotator Agreement (IAA). 🤝

💡 Analogy:

Imagine three teachers marking the same student's essay. If all three give it an A, we are very confident it deserves an A. If one gives A, one gives B, and one gives C — there is disagreement, and someone needs to look again! IAA measures how often your labelers agree. 🎓
📌 What does the code below do?

This code loads your exported labeled data, groups records that were labeled by multiple people, and calculates the percentage of time the labelers agreed with each other. It also highlights the records where labelers disagreed — those are the ones you should review and re-label! 🔍
import json
from collections import defaultdict

# Load the exported labeled data
labeled_records = []
with open("labeled_dataset.jsonl", "r") as f:
    for line in f:
        labeled_records.append(json.loads(line.strip()))

# Group annotations by source file
# (if multiple labelers annotated the same image, they'll have different record entries)
file_annotations = defaultdict(list)
for record in labeled_records:
    filename = record["source_file"]
    labels   = record["labels"]
    if labels:
        file_annotations[filename].append(labels[0])

# Calculate Inter-Annotator Agreement
total_images   = 0
agreed_images  = 0
disputed_files = []

for filename, label_list in file_annotations.items():
    if len(label_list) > 1:                        # only check doubly-labeled images
        total_images += 1
        if len(set(label_list)) == 1:              # all labelers agreed
            agreed_images += 1
        else:
            disputed_files.append({
                "file"   : filename,
                "labels" : label_list
            })

# Calculate agreement percentage
if total_images > 0:
    agreement_pct = (agreed_images / total_images) * 100
    print(f"📊 Inter-Annotator Agreement Report")
    print(f"   Total doubly-labeled images : {total_images}")
    print(f"   Images with full agreement  : {agreed_images}")
    print(f"   Agreement rate              : {agreement_pct:.1f}%")

    if agreement_pct >= 90:
        print(f"\n✅ Excellent agreement! Your labels are high quality.")
    elif agreement_pct >= 75:
        print(f"\n⚠️  Fair agreement. Review the disputed images below.")
    else:
        print(f"\n❌ Low agreement! Re-train your labeling team and re-label.")

    if disputed_files:
        print(f"\n🔍 Disputed images to review ({len(disputed_files)} total):")
        for item in disputed_files[:5]:           # show first 5 for brevity
            print(f"   File: {item['file']:30}  Labels given: {item['labels']}")

Output:

📊 Inter-Annotator Agreement Report
   Total doubly-labeled images : 50
   Images with full agreement  : 47
   Agreement rate              : 94.0%

✅ Excellent agreement! Your labels are high quality.

🔍 Disputed images to review (3 total):
   File: cat_dog_mixed_001.jpg        Labels given: ['Cat', 'Dog']
   File: blurry_bird_007.jpg          Labels given: ['Bird', 'Cat']
   File: small_animal_unclear.jpg     Labels given: ['Dog', 'Bird']

🤖 Step 7 — Train an OCI Vision Model with Your Labeled Data

Now for the most exciting moment — using your labeled data to train a real AI model with OCI Vision! OCI Vision can train a custom image classification or object detection model directly from a Data Labeling dataset. 🚀

📌 What does the code below do?

This code takes your labeled dataset (the one we created and filled in the earlier steps) and submits it to OCI Vision to train a custom AI model. OCI Vision will study all the labeled images, learn the patterns that distinguish cats, dogs, and birds, and produce a trained model you can then use to classify new images — automatically! 
import oci
import time

config          = oci.config.from_file()
vision_client   = oci.ai_vision.AIServiceVisionClient(config)
compartment_id  = "ocid1.compartment.oc1..your_compartment_id"
dataset_id      = "ocid1.datalabelingdataset.oc1..examplexxxxxxx"

# Step 1: Create a Vision Project (container for models)
print("📁 Creating OCI Vision project...")
project = vision_client.create_project(
    oci.ai_vision.models.CreateProjectDetails(
        display_name="Animal Classifier Project",
        compartment_id=compartment_id
    )
)
project_id = project.data.id
print(f"   Project ID: {project_id}")

# Step 2: Create and start model training
print("\n🚀 Starting model training with labeled dataset...")
model_response = vision_client.create_model(
    oci.ai_vision.models.CreateModelDetails(
        display_name="Animal Classifier v1",
        compartment_id=compartment_id,
        project_id=project_id,
        model_type="IMAGE_CLASSIFICATION",     # type of AI task
        model_version="1.0",
        training_dataset=oci.ai_vision.models.DataScienceLabelingDataset(
            dataset_type="DATA_SCIENCE_LABELING",
            dataset_id=dataset_id               # use our OCI Data Labeling dataset!
        ),
        # Hold out 20% of data for testing model accuracy
        testing_dataset=oci.ai_vision.models.DataScienceLabelingDataset(
            dataset_type="DATA_SCIENCE_LABELING",
            dataset_id=dataset_id
        ),
        is_quick_mode=False                    # False = full training for better accuracy
    )
)
model_id = model_response.data.id
print(f"   Model ID : {model_id}")
print(f"   Status   : {model_response.data.lifecycle_state}")
print(f"\n⏳ Training in progress — this may take 20–60 minutes...")
print(f"   Grab a chai! ☕")

# Step 3: Poll until training completes
while True:
    model = vision_client.get_model(model_id)
    state = model.data.lifecycle_state
    print(f"   Training status: {state}")

    if state == "ACTIVE":
        metrics = model.data.metrics
        print(f"\n🎉 Model training complete!")
        print(f"   Precision : {metrics.precision:.4f}")
        print(f"   Recall    : {metrics.recall:.4f}")
        print(f"   F1 Score  : {metrics.f_measure:.4f}")
        break
    elif state == "FAILED":
        print("\n❌ Training failed. Check your dataset for quality issues.")
        break

    time.sleep(60)    # check every minute

Output:

📁 Creating OCI Vision project...
   Project ID: ocid1.aivisionproject.oc1..examplexxxxxxx

🚀 Starting model training with labeled dataset...
   Model ID : ocid1.aivisionmodel.oc1..examplexxxxxxx
   Status   : CREATING

⏳ Training in progress — this may take 20–60 minutes...
   Grab a chai! ☕
   Training status: TRAINING
   Training status: TRAINING
   Training status: ACTIVE

🎉 Model training complete!
   Precision : 0.9621
   Recall    : 0.9534
   F1 Score  : 0.9577

96.2% precision on a custom animal classifier — trained from your own labeled data, entirely on OCI! 🏆

🔭 Step 8 — Use Your Trained Model on New Images

📌 What does the code below do?

This code takes a brand new image (that was never part of the training data) and sends it to your freshly trained AI model. The model looks at the image and tells you: "I am 97.3% sure this is a Cat." Like showing the trained puppy a new photo and watching it correctly identify the animal! 🐾
import oci
import base64

config        = oci.config.from_file()
vision_client = oci.ai_vision.AIServiceVisionClient(config)
model_id      = "ocid1.aivisionmodel.oc1..examplexxxxxxx"

# Load a new image that the model has never seen before
with open("new_animal_photo.jpg", "rb") as f:
    image_bytes  = f.read()
    image_base64 = base64.b64encode(image_bytes).decode("utf-8")

# Ask the model: "What animal is in this photo?"
print("🔍 Analyzing new image with your custom model...")
response = vision_client.analyze_image(
    oci.ai_vision.models.AnalyzeImageDetails(
        image=oci.ai_vision.models.InlineImageDetails(
            source="INLINE",
            data=image_base64
        ),
        features=[
            oci.ai_vision.models.ImageClassificationFeature(
                feature_type="IMAGE_CLASSIFICATION",
                model_id=model_id,    # use OUR custom trained model
                max_results=3         # return top 3 predictions
            )
        ]
    )
)

# Print the predictions
labels = response.data.labels
print(f"\n🎯 Prediction Results:")
print(f"{'Label':<15 onfidence="">12}")
print(f"{'─'*15}  {'─'*12}")
for label in labels:
    confidence_pct = label.confidence * 100
    bar = "█" * int(confidence_pct / 5)   # visual confidence bar
    print(f"{label.name:<15 confidence_pct:="">10.1f}%  {bar}")

Output:

🔍 Analyzing new image with your custom model...

🎯 Prediction Results:
Label            Confidence
───────────────  ────────────
Cat                    97.3%  ████████████████████
Dog                     2.1%  
Bird                    0.6%

97.3% confident — it is definitely a cat! 🐱✅

📝 Step 9 — Full Text Labeling Example (Sentiment Analysis)

Let's do a complete text labeling workflow — building a sentiment classifier that can automatically decide whether a customer review is Positive, Negative, or Neutral.

Step 9a: Create a Text Dataset

📌 What does the code below do?

This code creates a text classification dataset in OCI Data Labeling. Instead of images, we are now dealing with customer review text files. The possible labels are Positive, Negative, and Neutral. Think of it as setting up a sorting tray with three bins for customer feedback! 📬
import oci

config         = oci.config.from_file()
dl_client      = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
namespace      = "your_namespace"
bucket_name    = "customer-reviews-bucket"    # bucket containing .txt review files

# Labels for sentiment analysis
labels = [
    oci.data_labeling_service.models.Label(name="Positive"),
    oci.data_labeling_service.models.Label(name="Negative"),
    oci.data_labeling_service.models.Label(name="Neutral")
]

# Create the TEXT dataset
response = dl_client.create_dataset(
    oci.data_labeling_service.models.CreateDatasetDetails(
        display_name="Customer Review Sentiment Dataset",
        compartment_id=compartment_id,
        annotation_format="SINGLE_LABEL",

        # TEXT format this time!
        dataset_format_details=oci.data_labeling_service.models.TextDatasetFormatDetails(
            format_type="TEXT"
        ),

        dataset_source_details=oci.data_labeling_service.models.ObjectStorageSourceDetails(
            source_type="OBJECT_STORAGE",
            namespace=namespace,
            bucket=bucket_name,
            prefix="reviews/"     # only files inside the "reviews/" folder
        ),

        label_set=oci.data_labeling_service.models.LabelSet(items=labels)
    )
)

print(f"✅ Text Dataset created!")
print(f"   Dataset ID : {response.data.id}")
print(f"   Labels     : Positive, Negative, Neutral")
print(f"   Status     : {response.data.lifecycle_state}")

Output:

✅ Text Dataset created!
   Dataset ID : ocid1.datalabelingdataset.oc1..exampleyyyyyyy
   Labels     : Positive, Negative, Neutral
   Status     : ACTIVE

Step 9b: Bulk-Add Labels Programmatically

📌 What does the code below do?

If you already have a list of pre-labeled text records (maybe exported from an old system), this code imports those labels automatically into OCI Data Labeling — saving your team from doing it manually one by one. Like bulk-importing contacts into your phone instead of typing each one! 📱
import oci
import json

config          = oci.config.from_file()
dl_dp_client    = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id  = "ocid1.compartment.oc1..your_compartment_id"
dataset_id      = "ocid1.datalabelingdataset.oc1..exampleyyyyyyy"

# Your pre-existing labeled data (from a previous system or CSV export)
pre_labeled = [
    {"filename": "review_001.txt", "label": "Positive"},
    {"filename": "review_002.txt", "label": "Negative"},
    {"filename": "review_003.txt", "label": "Neutral"},
    # ... hundreds more
]

# Fetch all records in the dataset to get their record IDs
all_records = dl_dp_client.list_records(
    compartment_id=compartment_id,
    dataset_id=dataset_id
).data.items

# Build a lookup: filename → record_id
filename_to_id = {r.name: r.id for r in all_records}

# Bulk-create annotations (labels) for each record
success_count = 0
for entry in pre_labeled:
    filename = entry["filename"]
    label    = entry["label"]

    if filename not in filename_to_id:
        print(f"   ⚠️  Not found in dataset: {filename}")
        continue

    record_id = filename_to_id[filename]

    # Create the annotation (attach the label to the record)
    dl_dp_client.create_annotation(
        oci.data_labeling_service_dataplane.models.CreateAnnotationDetails(
            record_id=record_id,
            compartment_id=compartment_id,
            entities=[
                oci.data_labeling_service_dataplane.models.ClassificationLabel(
                    entity_type="GENERIC",
                    labels=[
                        oci.data_labeling_service_dataplane.models.Label(label=label)
                    ]
                )
            ]
        )
    )
    success_count += 1

print(f"\n✅ Bulk labeling complete! {success_count} records labeled automatically.")

Output:

✅ Bulk labeling complete! 3 records labeled automatically.

🏆 Best Practices for OCI Data Labeling

  • 📋 Write a clear labeling guide before you start — define exactly what each label means with examples. Show your team: "This is Cat. This is NOT cat (even if it looks like one)." 📖
  • 👥 Always use at least 2 labelers per record for high-stakes projects and resolve disagreements with a senior reviewer.
  • ⚖️ Balance your classes — 500 Cat + 500 Dog + 500 Bird is better than 900 Cat + 90 Dog + 10 Bird. Imbalanced data creates biased models!
  • 🔁 Use the Active Learning loop — label a small batch → train a quick model → let the model pre-label the rest → only review the uncertain ones. Repeat until done!
  • 📊 Monitor IAA scores regularly — if agreement drops below 80%, stop labeling and re-train your team first.
  • 🔐 Use IAM groups for labelers — give each labeler exactly the permissions they need (label data only). They should never be able to delete records or export the dataset. 🔒
❌ Common Mistakes to Avoid:

  • Never start labeling without a written labeling guide — confusion leads to inconsistent labels
  • Never use AI-assisted labels without human review — the AI can be wrong, especially early on
  • Never label data you are unsure about — skip it rather than guess
  • Never use only one person for all labeling — there is no way to catch their blind spots
  • Never delete the raw data from Object Storage after labeling — you may need to re-label later

🌍 Real-World Data Labeling Scenarios

  • 🏭 Manufacturing Quality Control — Label images of products as "Defective" or "Good" → train a model to detect defects on the production line automatically.
  • 🏥 Medical Imaging — Label X-ray regions as "Normal" or "Anomaly" → AI assists radiologists by flagging suspicious areas first.
  • 💬 Customer Support Automation — Label support tickets as "Billing Issue", "Technical Issue", "Complaint", "Praise" → AI auto-routes tickets to the right team.
  • 📦 E-Commerce Cataloguing — Label product images by category → AI automatically categorises new products as they are uploaded.
  • 📄 Legal Document Processing — Label entities in contracts (dates, party names, clause types) → AI extracts key info from new contracts in seconds.

📝 Quick Summary — What We Learned

  • What Data Labeling is → Attaching correct answers to raw data so AI can learn from it
  • OCI Data Labeling service → Browser-based, team-friendly, AI-assisted labeling platform
  • Four task types → Image Classification, Object Detection, Text Classification, Entity Extraction
  • Three-zone workflow → Upload raw data → Label in UI → Export labeled dataset
  • AI-Assisted Labeling → Label a small seed set, let AI pre-label the rest, review only uncertain ones
  • Bulk Labeling via SDK → Import pre-existing labels programmatically
  • Quality Measurement → Inter-Annotator Agreement (IAA) to catch bad labels early
  • End-to-End Pipeline → Data Labeling → OCI Vision Training → Custom Model → API Predictions

Remember: the quality of your lazels is the quality of your AI. Label with care, and your model will reward you with accuracy! 🏷️✨

Comments