Skip to main content

OCI Dataset Life Cycle: Create, Manage, Version, and Use Datasets

Calculating read time…

Imagine you are baking a cake🎂

You go through stages: gather ingredients → mix them → bake → decorate → serve → finally clean up the kitchen.

A Dataset in OCI Data Labeling goes through a very similar journey — from being created, to being actively labeled, to being completed, saved, and eventually retired.
This journey is called the Dataset Life Cycle. Let's master every stage, step by step! 




🔄 What is a Dataset Life Cycle?

Every OCI Data Labeling Dataset moves through a defined set of states during its lifetime — from the moment it is created to the moment it is deleted.

Understanding these states is critical because:

  • 📋 Each state controls what actions you can perform on the dataset
  • 🤖 Your code needs to check a dataset's state before calling APIs on it
  • 💰 Inactive datasets still consume Object Storage — knowing when to delete saves cost
  • 🔐 Security and access controls behave differently depending on the state

🗺️ The Complete Dataset Life Cycle — All States at a Glance

  ┌─────────────────────────────────────────────────────────────────────────┐
  │              OCI DATA LABELING — DATASET LIFE CYCLE                     │
  └─────────────────────────────────────────────────────────────────────────┘

         CREATE
            │
            ▼
      ┌───────────┐
      │  CREATING │  ◄── OCI is setting up the dataset in the background
      └─────┬─────┘
            │  (auto-transitions in seconds to minutes)
            ▼
      ┌───────────┐       ┌────────────┐
      │  ACTIVE   │ ────► │  NEEDS     │  ◄── Needs attention / data issue
      │  (normal  │       │  ATTENTION │
      │   state)  │ ◄──── └────────────┘
      └─────┬─────┘
            │
     ┌──────┼────────────────────┐
     │      │                    │
     ▼      ▼                    ▼
  ┌──────┐ ┌──────────────┐  ┌──────────┐
  │UPDATE│ │  DELETING    │  │ IMPORTING│  ◄── Bulk import in progress
  │ (add │ │  (permanent) │  └──────────┘
  │labels│ └──────────────┘
  │  or  │
  │records│
  └──────┘

  STATE SUMMARY:
  ─────────────────────────────────────────────────────────────────
  CREATING        → Dataset is being provisioned (read-only, wait)
  ACTIVE          → Fully operational — label, query, export, update
  NEEDS ATTENTION → Something needs fixing (e.g. source bucket issue)
  DELETING        → Permanent deletion in progress
  IMPORTING       → Bulk record import is running (partial read-only)
  ─────────────────────────────────────────────────────────────────
💡 Analogy — Dataset States are like a Library Book:

📗 CREATING = The book is being printed — not available yet
📘 ACTIVE = Book is on the shelf — anyone can read or borrow it
📙 NEEDS ATTENTION = Book is damaged — needs repair before use
📕 IMPORTING = New pages being added — can read but not fully complete
🗑️ DELETING = Book is being shredded — gone forever

1️⃣ State: CREATING

When you first call the Create Dataset API, the dataset enters the CREATING state immediately. OCI is doing background work:

  • 🔍 Scanning the source Object Storage bucket to find all files
  • 📋 Building the initial list of records (one record per file)
  • 🏷️ Setting up the label schema (the list of labels you defined)
  • 🔐 Applying access controls and metadata
📌 What does the code below do?

This code creates a new dataset and then watches it using a loop — checking every few seconds until the dataset finishes the CREATING stage and becomes ACTIVE. Think of it like watching the oven light — waiting for it to turn off so you know the cake is ready! 🎂⏳
import oci
import time

config         = oci.config.from_file()
dl_client      = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"

# Step 1: Submit the Create Dataset request
print("📦 Creating dataset...")
response = dl_client.create_dataset(
    oci.data_labeling_service.models.CreateDatasetDetails(
        display_name="Dog-Cat-Bird Image Classifier Dataset",
        compartment_id=compartment_id,
        annotation_format="SINGLE_LABEL",
        dataset_format_details=oci.data_labeling_service.models.ImageDatasetFormatDetails(
            format_type="IMAGE"
        ),
        dataset_source_details=oci.data_labeling_service.models.ObjectStorageSourceDetails(
            source_type="OBJECT_STORAGE",
            namespace="your_namespace",
            bucket="raw-training-images",
            prefix=""
        ),
        label_set=oci.data_labeling_service.models.LabelSet(items=[
            oci.data_labeling_service.models.Label(name="Dog"),
            oci.data_labeling_service.models.Label(name="Cat"),
            oci.data_labeling_service.models.Label(name="Bird")
        ])
    )
)

dataset_id    = response.data.id
current_state = response.data.lifecycle_state

print(f"   Dataset ID    : {dataset_id}")
print(f"   Initial State : {current_state}")

# Step 2: Poll until the dataset leaves CREATING state
print(f"\n⏳ Waiting for dataset to become ACTIVE...")
while current_state == "CREATING":
    time.sleep(5)
    dataset       = dl_client.get_dataset(dataset_id)
    current_state = dataset.data.lifecycle_state
    record_count  = dataset.data.dataset_statistics.record_count if dataset.data.dataset_statistics else "?"
    print(f"   State: {current_state:20}  |  Records found so far: {record_count}")

if current_state == "ACTIVE":
    final = dl_client.get_dataset(dataset_id).data
    print(f"\n✅ Dataset is now ACTIVE!")
    print(f"   Total records : {final.dataset_statistics.record_count}")
    print(f"   Labeled       : {final.dataset_statistics.labeled_record_count}")
    print(f"   Unlabeled     : {final.dataset_statistics.unlabeled_record_count}")
elif current_state == "NEEDS_ATTENTION":
    print(f"\n⚠️  Dataset needs attention — check your source bucket!")

Output:

📦 Creating dataset...
   Dataset ID    : ocid1.datalabelingdataset.oc1..examplexxxxxx
   Initial State : CREATING

⏳ Waiting for dataset to become ACTIVE...
   State: CREATING              |  Records found so far: 0
   State: CREATING              |  Records found so far: 87
   State: CREATING              |  Records found so far: 203
   State: ACTIVE                |  Records found so far: 250

✅ Dataset is now ACTIVE!
   Total records : 250
   Labeled       : 0
   Unlabeled     : 250

2️⃣ State: ACTIVE — The Working State

ACTIVE is the most important state — it is where your dataset spends most of its life. In this state, you can do everything:

  • 🏷️ Create annotations — add labels to individual records
  • 🔍 Query records — list labeled, unlabeled, or specific records
  • 📊 Check statistics — see progress, label distribution, completion percentage
  • ➕ Add new labels — extend the label schema with additional classes
  • 📥 Import records — bulk-add new files from Object Storage
  • 📤 Export — download the labeled dataset for model training
  • ✏️ Update annotations — fix incorrect labels
  • 🗑️ Delete records — remove specific files from the dataset
✅ Rule of thumb:

If a dataset is ACTIVE, you are free to work on it. All APIs are available. This is the green light state! 🟢

Reading Dataset Statistics (ACTIVE state)

📌 What does the code below do?

This code reads the current statistics of your ACTIVE dataset — how many records are there, how many are labeled, and what percentage is complete. Think of it as checking a progress bar on a download — you want to know how much is done and how much is left! 📊
import oci

config    = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
dl_dp     = oci.data_labeling_service_dataplane.DataLabelingClient(config)

compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id     = "ocid1.datalabelingdataset.oc1..examplexxxxxx"

# Fetch the dataset details
dataset = dl_client.get_dataset(dataset_id).data

# Confirm it is ACTIVE before we proceed
print(f"📋 Dataset Status Report")
print(f"{'='*45}")
print(f"  Name          : {dataset.display_name}")
print(f"  State         : {dataset.lifecycle_state}")
print(f"  Format        : {dataset.dataset_format_details.format_type}")
print(f"  Annotation    : {dataset.annotation_format}")

stats = dataset.dataset_statistics
if stats:
    total        = stats.record_count
    labeled      = stats.labeled_record_count
    unlabeled    = stats.unlabeled_record_count
    pct_complete = (labeled / total * 100) if total > 0 else 0

    print(f"\n  📊 Progress:")
    print(f"  Total records    : {total:,}")
    print(f"  Labeled          : {labeled:,}")
    print(f"  Unlabeled        : {unlabeled:,}")
    print(f"  Completion       : {pct_complete:.1f}%")

    # Draw a simple ASCII progress bar
    filled = int(pct_complete / 5)
    bar    = "█" * filled + "░" * (20 - filled)
    print(f"\n  [{bar}] {pct_complete:.1f}% complete")

# Also get per-label breakdown
print(f"\n  🏷️  Label Distribution:")
annotations_summary = dl_dp.summarize_annotation(
    dataset_id=dataset_id,
    compartment_id=compartment_id,
    request_type="COUNT"
)
# Count per label
label_counts = {}
for item in annotations_summary.data.items:
    for entity in item.entities:
        for label in entity.labels:
            lname = label.label_name
            label_counts[lname] = label_counts.get(lname, 0) + 1

for label, count in sorted(label_counts.items(), key=lambda x: -x[1]):
    bar = "▓" * min(count, 30)
    print(f"  {label:<15 :="" count:="">5}  {bar}")

Output:

📋 Dataset Status Report
=============================================
  Name          : Dog-Cat-Bird Image Classifier Dataset
  State         : ACTIVE
  Format        : IMAGE
  Annotation    : SINGLE_LABEL

  📊 Progress:
  Total records    : 250
  Labeled          : 187
  Unlabeled        : 63
  Completion       : 74.8%

  [███████████████░░░░░] 74.8% complete

  🏷️  Label Distribution:
  Dog             :    82  ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓
  Cat             :    71  ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓
  Bird            :    34  ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓

3️⃣ Life Cycle Operation: Adding New Labels to an ACTIVE Dataset

Imagine you started labeling images as "Dog" and "Cat" but then realised you also have rabbit photos in the dataset! In the ACTIVE state, you can add new labels to the existing schema without disturbing any records already labeled. 🐰

📌 What does the code below do?

This code adds a new label ("Rabbit") to an existing dataset that already has Dog, Cat, Bird labels. It is like adding a new shelf section in a library without disturbing the existing books — all old labels stay exactly as they were! 📚➕
import oci

config    = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)

dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"

# Add a new label to an ACTIVE dataset
print("➕ Adding new label 'Rabbit' to dataset...")
response = dl_client.add_dataset_labels(
    dataset_id=dataset_id,
    add_dataset_labels_details=oci.data_labeling_service.models.AddDatasetLabelsDetails(
        label_set=oci.data_labeling_service.models.LabelSet(
            items=[
                oci.data_labeling_service.models.Label(name="Rabbit")
            ]
        )
    )
)

print(f"✅ New label added!")
print(f"   Work Request ID : {response.headers.get('opc-work-request-id', 'N/A')}")
print(f"\n   Current label set now includes:")
updated_dataset = dl_client.get_dataset(dataset_id).data
for label in updated_dataset.label_set.items:
    print(f"   · {label.name}")

Output:

➕ Adding new label 'Rabbit' to dataset...
✅ New label added!
   Work Request ID : ocid1.workrequest.oc1..examplexxxxxx

   Current label set now includes:
   · Dog
   · Cat
   · Bird
   · Rabbit
❌ You CANNOT remove labels from a dataset once added!

OCI Data Labeling does not allow deleting a label that already exists in a dataset. If you made a typo (e.g. "Rabit" instead of "Rabbit"), you cannot fix it — you must either work around it or delete and recreate the dataset. Always double-check your label names before creating the dataset! 🔤

4️⃣ Life Cycle Operation: Importing Records (IMPORTING state)

When you need to add more files to an existing ACTIVE dataset (say, your team found 500 more training images), you trigger a bulk record import. During this process, the dataset enters a temporary IMPORTING sub-state.

  IMPORTING State Flow:
  ─────────────────────────────────────────────────────────────
  ACTIVE ──► [Trigger import] ──► IMPORTING ──► ACTIVE
                                  (background job runs,
                                   new records appear
                                   one by one as
                                   they are added)
  ─────────────────────────────────────────────────────────────
  ⚠️  During IMPORTING: you CAN still label existing records
                        you CANNOT modify dataset metadata
📌 What does the code below do?

This code triggers a bulk import of new image files from an Object Storage prefix into an existing dataset. It is like a delivery truck bringing new books to a library — the library stays open while the books are being unloaded and shelved, but the catalogue is not finalised until the delivery is complete! 🚚📚
import oci
import time

config    = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
dl_dp     = oci.data_labeling_service_dataplane.DataLabelingClient(config)

compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id     = "ocid1.datalabelingdataset.oc1..examplexxxxxx"

# Step 1: Trigger bulk record generation (import from Object Storage prefix)
print("📥 Triggering bulk record import from Object Storage...")
import_response = dl_client.generate_dataset_records(
    dataset_id=dataset_id,
    generate_dataset_records_details=oci.data_labeling_service.models.GenerateDatasetRecordsDetails(
        limit=500.0      # import up to 500 new files
    )
)

work_request_id = import_response.headers.get("opc-work-request-id")
print(f"   Import job submitted!")
print(f"   Work Request ID : {work_request_id}")

# Step 2: Poll the work request status
wrc = oci.work_requests.WorkRequestClient(config)
print(f"\n⏳ Monitoring import progress...")

while True:
    wr    = wrc.get_work_request(work_request_id)
    state = wr.data.status
    pct   = wr.data.percent_complete

    bar  = "█" * int(pct / 5) + "░" * (20 - int(pct / 5))
    print(f"   [{bar}] {pct:.0f}%  Status: {state}")

    if state == "SUCCEEDED":
        print(f"\n✅ Import complete!")
        break
    elif state in ["FAILED", "CANCELED"]:
        print(f"\n❌ Import failed with state: {state}")
        for err in wr.data.errors:
            print(f"   Error: {err.message}")
        break

    time.sleep(10)

# Step 3: Check updated record count
dataset = dl_client.get_dataset(dataset_id).data
print(f"\n   Updated Total Records : {dataset.dataset_statistics.record_count:,}")
print(f"   Newly Imported        : {dataset.dataset_statistics.unlabeled_record_count:,} unlabeled")

Output:

📥 Triggering bulk record import from Object Storage...
   Import job submitted!
   Work Request ID : ocid1.workrequest.oc1..examplexxxxxx

⏳ Monitoring import progress...
   [░░░░░░░░░░░░░░░░░░░░]  0%  Status: IN_PROGRESS
   [████░░░░░░░░░░░░░░░░] 20%  Status: IN_PROGRESS
   [████████░░░░░░░░░░░░] 40%  Status: IN_PROGRESS
   [████████████░░░░░░░░] 60%  Status: IN_PROGRESS
   [████████████████░░░░] 80%  Status: IN_PROGRESS
   [████████████████████]100%  Status: SUCCEEDED

✅ Import complete!

   Updated Total Records : 750
   Newly Imported        : 500 unlabeled

5️⃣ State: NEEDS_ATTENTION — When Something Goes Wrong

The NEEDS_ATTENTION state is OCI's way of tapping you on the shoulder and saying: "Hey, something is not right — please check this dataset!" 👋

Common reasons a dataset enters this state:

  • 🪣 The source Object Storage bucket was deleted or renamed after the dataset was created
  • 🔐 The IAM permissions changed and OCI can no longer access the source bucket
  • 📁 The source prefix no longer exists or was moved
  • 🔄 A background import job failed partway through
📌 What does the code below do?

This code checks if any datasets in your compartment have entered the NEEDS_ATTENTION state and prints a helpful diagnosis for each one — like a nurse checking all patients in a ward to see who needs immediate care! 🏥
import oci

config         = oci.config.from_file()
dl_client      = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"

print("🔍 Scanning all datasets for issues...")
print("=" * 55)

# List ALL datasets in the compartment
all_datasets = dl_client.list_datasets(
    compartment_id=compartment_id,
    limit=100
).data.items

needs_attention = [ds for ds in all_datasets if ds.lifecycle_state == "NEEDS_ATTENTION"]
active_count    = len([ds for ds in all_datasets if ds.lifecycle_state == "ACTIVE"])

print(f"  Total datasets     : {len(all_datasets)}")
print(f"  ACTIVE             : {active_count}   ✅")
print(f"  NEEDS_ATTENTION    : {len(needs_attention)}   ⚠️")

if needs_attention:
    print(f"\n⚠️  Datasets requiring your attention:")
    for ds in needs_attention:
        print(f"\n  Dataset : {ds.display_name}")
        print(f"  ID      : {ds.id}")
        print(f"  State   : {ds.lifecycle_state}")

        # Fetch detailed info for diagnosis
        detail = dl_client.get_dataset(ds.id).data
        src    = detail.dataset_source_details

        print(f"\n  📁 Source details:")
        print(f"     Bucket    : {src.bucket}")
        print(f"     Namespace : {src.namespace}")
        print(f"     Prefix    : '{src.prefix or '(all files)'}'")

        print(f"\n  🔧 Suggested fixes:")
        print(f"     1. Confirm bucket '{src.bucket}' still exists in Object Storage")
        print(f"     2. Confirm your IAM policies still allow access to this bucket")
        print(f"     3. Check for any recent changes to the bucket permissions")
        print(f"     4. Contact your OCI admin if the bucket was deleted")
else:
    print(f"\n✅ All datasets are healthy! No attention needed.")

Output (when a dataset has an issue):

🔍 Scanning all datasets for issues...
=======================================================
  Total datasets     : 4
  ACTIVE             : 3   ✅
  NEEDS_ATTENTION    : 1   ⚠️

⚠️  Datasets requiring your attention:

  Dataset : Product Defect Detection Dataset
  ID      : ocid1.datalabelingdataset.oc1..exampleyyyyyyy
  State   : NEEDS_ATTENTION

  📁 Source details:
     Bucket    : factory-raw-images-old
     Namespace : ocitenancyexample
     Prefix    : 'products/2025/'

  🔧 Suggested fixes:
     1. Confirm bucket 'factory-raw-images-old' still exists in Object Storage
     2. Confirm your IAM policies still allow access to this bucket
     3. Check for any recent changes to the bucket permissions
     4. Contact your OCI admin if the bucket was deleted

6️⃣ Record Life Cycle — Inside the Dataset

Within an ACTIVE dataset, each individual Record also has its own lifecycle! A record is one file (one image or one text snippet) in the dataset.

  RECORD LIFECYCLE (inside an ACTIVE dataset):
  ──────────────────────────────────────────────────────────────────
  State            │ Meaning
  ─────────────────┼────────────────────────────────────────────────
  ACTIVE           │ Record exists, ready to be labeled or queried
  FAILED           │ OCI could not access this file in Object Storage
                   │ (file was deleted or moved after dataset creation)
  ──────────────────────────────────────────────────────────────────

  And each record can have one of these ANNOTATION states:
  ──────────────────────────────────────────────────────────────────
  (no annotation)  │ Not yet labeled — shows as Unlabeled
  ACTIVE           │ Has at least one label attached — Labeled
  ──────────────────────────────────────────────────────────────────

Querying Records by State

📌 What does the code below do?

This code queries the dataset to find: all records that have been labeled (so you can review them), all records that still need labeling (so your team can focus on them), and all records that have FAILED (so you can investigate missing files). Like a school roll-call that shows who is present, absent, or has a problem! 📋
import oci

config         = oci.config.from_file()
dl_dp          = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id     = "ocid1.datalabelingdataset.oc1..examplexxxxxx"

def fetch_all_records(is_labeled, lifecycle_state=None):
    """Fetch all records matching the given filters (handles pagination)."""
    all_items  = []
    page_token = None
    kwargs     = dict(
        compartment_id=compartment_id,
        dataset_id=dataset_id,
        is_labeled=is_labeled,
        limit=100
    )
    if lifecycle_state:
        kwargs["lifecycle_state"] = lifecycle_state

    while True:
        if page_token:
            kwargs["page"] = page_token
        resp = dl_dp.list_records(**kwargs)
        all_items.extend(resp.data.items)
        if resp.has_next_page:
            page_token = resp.next_page
        else:
            break
    return all_items

# Query each category of records
print("📋 Record Status Report for Dataset")
print("=" * 50)

labeled_records   = fetch_all_records(is_labeled=True)
unlabeled_records = fetch_all_records(is_labeled=False)
failed_records    = fetch_all_records(is_labeled=False, lifecycle_state="FAILED")

total = len(labeled_records) + len(unlabeled_records)
pct   = len(labeled_records) / total * 100 if total > 0 else 0

print(f"\n  ✅ Labeled records   : {len(labeled_records):>5}")
print(f"  ⏳ Unlabeled records : {len(unlabeled_records):>5}")
print(f"  ❌ Failed records    : {len(failed_records):>5}")
print(f"  📊 Completion        : {pct:.1f}%")

if failed_records:
    print(f"\n  ❌ Failed records (files no longer accessible):")
    for rec in failed_records[:5]:        # show first 5
        print(f"     · {rec.name}  (ID: {rec.id})")
    if len(failed_records) > 5:
        print(f"     ... and {len(failed_records)-5} more")

# Show a sample of unlabeled records (next ones to label)
if unlabeled_records:
    print(f"\n  ⏳ Next records to label (sample of 5):")
    for rec in unlabeled_records[:5]:
        print(f"     · {rec.name}")

Output:

📋 Record Status Report for Dataset
==================================================

  ✅ Labeled records   :   187
  ⏳ Unlabeled records :    63
  ❌ Failed records    :     2
  📊 Completion        :  74.8%

  ❌ Failed records (files no longer accessible):
     · deleted_image_001.jpg  (ID: ocid1.datalabelingrecord...001)
     · moved_cat_photo.png    (ID: ocid1.datalabelingrecord...002)

  ⏳ Next records to label (sample of 5):
     · dog_batch2_001.jpg
     · dog_batch2_002.jpg
     · bird_outdoor_007.jpg
     · cat_indoor_012.jpg
     · unknown_animal_001.jpg

7️⃣ Annotation Life Cycle — The Label on Each Record

An Annotation is the actual label (or set of labels) attached to a record. Annotations also have their own lifecycle within a dataset!

  ANNOTATION LIFECYCLE:
  ──────────────────────────────────────────────────────────────────
  Action               │ Result
  ─────────────────────┼────────────────────────────────────────────
  Create annotation    │ Record becomes "Labeled"
  Update annotation    │ Label changed (old annotation deleted,
                       │ new one created in ACTIVE state)
  Delete annotation    │ Record returns to "Unlabeled"
  ──────────────────────────────────────────────────────────────────
  Annotation states:
  ACTIVE              = The currently valid label for this record
  DELETED             = A previous label that was corrected/removed
  ──────────────────────────────────────────────────────────────────
📌 What does the code below do?

This code shows the complete annotation lifecycle operations: create a label, update it (fix a mistake), and delete it (remove a label from a record). Think of it like writing, erasing, and rewriting answers on an exam paper! ✏️📝
import oci

config         = oci.config.from_file()
dl_dp          = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id     = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
record_id      = "ocid1.datalabelingrecord.oc1..examplerecord1"

# ── OPERATION 1: CREATE an annotation (label a record for the first time)
print("1️⃣  Creating annotation (labeling record as 'Dog')...")
create_resp = dl_dp.create_annotation(
    oci.data_labeling_service_dataplane.models.CreateAnnotationDetails(
        record_id=record_id,
        compartment_id=compartment_id,
        entities=[
            oci.data_labeling_service_dataplane.models.ClassificationLabel(
                entity_type="GENERIC",
                labels=[
                    oci.data_labeling_service_dataplane.models.Label(label="Dog")
                ]
            )
        ]
    )
)
annotation_id = create_resp.data.id
print(f"   ✅ Annotation created!")
print(f"   Annotation ID : {annotation_id}")
print(f"   Label         : Dog")
print(f"   State         : {create_resp.data.lifecycle_state}")

# ── OPERATION 2: UPDATE the annotation (oops — it was actually a Cat!)
print(f"\n2️⃣  Updating annotation (correcting label to 'Cat')...")
update_resp = dl_dp.update_annotation(
    annotation_id=annotation_id,
    update_annotation_details=oci.data_labeling_service_dataplane.models.UpdateAnnotationDetails(
        entities=[
            oci.data_labeling_service_dataplane.models.ClassificationLabel(
                entity_type="GENERIC",
                labels=[
                    oci.data_labeling_service_dataplane.models.Label(label="Cat")
                ]
            )
        ]
    )
)
print(f"   ✅ Annotation updated!")
print(f"   New label : Cat  (was: Dog)")
print(f"   State     : {update_resp.data.lifecycle_state}")

# ── OPERATION 3: DELETE the annotation (record is too blurry to label reliably)
print(f"\n3️⃣  Deleting annotation (image too blurry to label confidently)...")
dl_dp.delete_annotation(annotation_id=annotation_id)
print(f"   ✅ Annotation deleted!")
print(f"   Record is now back to: Unlabeled")
print(f"   It will appear in the unlabeled queue again for a fresh labeler.")

Output:

1️⃣  Creating annotation (labeling record as 'Dog')...
   ✅ Annotation created!
   Annotation ID : ocid1.datalabelingannotation.oc1..exampleanno1
   Label         : Dog
   State         : ACTIVE

2️⃣  Updating annotation (correcting label to 'Cat')...
   ✅ Annotation updated!
   New label : Cat  (was: Dog)
   State     : ACTIVE

3️⃣  Deleting annotation (image too blurry to label confidently)...
   ✅ Annotation deleted!
   Record is now back to: Unlabeled
   It will appear in the unlabeled queue again for a fresh labeler.

8️⃣ Life Cycle Operation: Removing a Label from the Schema

When a label is no longer needed in the dataset — for example, you merged "Puppy" and "Dog" into just "Dog" — you can remove a label from the label schema.

📌 What does the code below do?

This code removes a label from the dataset's label schema. OCI will also automatically delete all existing annotations that used this label — those records become "Unlabeled" again and need to be re-labeled. Think of it like removing a subject from the school curriculum: all students who were enrolled in that subject need to re-enrol in a new one! 🏫
import oci
import time

config    = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)

dataset_id     = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
label_to_remove = "Puppy"    # we are merging this into "Dog"

print(f"🗑️  Removing label '{label_to_remove}' from dataset...")
print(f"   ⚠️  All records labeled '{label_to_remove}' will become Unlabeled!")

response = dl_client.remove_dataset_labels(
    dataset_id=dataset_id,
    remove_dataset_labels_details=oci.data_labeling_service.models.RemoveDatasetLabelsDetails(
        label_set=oci.data_labeling_service.models.LabelSet(
            items=[oci.data_labeling_service.models.Label(name=label_to_remove)]
        )
    )
)

# Monitor the work request
work_request_id = response.headers.get("opc-work-request-id")
wrc             = oci.work_requests.WorkRequestClient(config)

print(f"\n⏳ Waiting for label removal to complete...")
while True:
    wr    = wrc.get_work_request(work_request_id)
    state = wr.data.status
    pct   = wr.data.percent_complete
    print(f"   {state}  {pct:.0f}%")
    if state == "SUCCEEDED":
        break
    elif state in ["FAILED", "CANCELED"]:
        print(f"❌ Label removal failed!")
        break
    time.sleep(5)

print(f"\n✅ Label '{label_to_remove}' removed!")
print(f"   Affected records are now Unlabeled.")
print(f"   Go to the labeling UI and re-label them as 'Dog'.")

9️⃣ State: DELETING — The Final Stage

When you no longer need a dataset — maybe the project is complete, or the data was moved elsewhere — you delete it. The dataset enters the DELETING state while OCI cleans up all its metadata.

📌 What does the code below do?

This code deletes a dataset from OCI Data Labeling. It first checks the dataset is in ACTIVE state (safe to delete), then submits the delete request and monitors until it is fully gone. Think of it like returning a rented lab space — you clean everything out before handing back the keys! 🔑🏢
import oci
import time

config    = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)

dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"

# Step 1: Confirm the dataset state before deleting
dataset = dl_client.get_dataset(dataset_id).data
print(f"📋 Dataset to delete:")
print(f"   Name  : {dataset.display_name}")
print(f"   State : {dataset.lifecycle_state}")
print(f"   Records: {dataset.dataset_statistics.record_count:,}")

if dataset.lifecycle_state != "ACTIVE":
    print(f"\n⚠️  Cannot delete — dataset is in '{dataset.lifecycle_state}' state.")
    print(f"   Only ACTIVE datasets can be deleted.")
    exit()

# Step 2: Safety confirmation (in a real script, ask for user input!)
print(f"\n⚠️  WARNING: This will permanently delete the dataset")
print(f"   and ALL its annotations. This cannot be undone!")
print(f"\n   (In production, add: input('Type CONFIRM to proceed: '))")

# Step 3: Delete the dataset
print(f"\n🗑️  Deleting dataset...")
response = dl_client.delete_dataset(dataset_id=dataset_id)
print(f"   Delete request submitted.")

# Step 4: Monitor until deletion completes
print(f"\n⏳ Waiting for deletion to complete...")
while True:
    try:
        ds    = dl_client.get_dataset(dataset_id)
        state = ds.data.lifecycle_state
        print(f"   State: {state}")
        if state == "DELETED":
            break
        time.sleep(5)
    except oci.exceptions.ServiceError as e:
        if e.status == 404:
            # 404 means the dataset no longer exists — deletion complete!
            print(f"\n✅ Dataset deleted successfully!")
            print(f"   The dataset and all its annotations are permanently removed.")
            print(f"\n   ℹ️  Note: The raw image files in Object Storage are NOT deleted.")
            print(f"   Only the labeling metadata was removed from OCI Data Labeling.")
            break
        else:
            raise

Output:

📋 Dataset to delete:
   Name  : Old Product Test Dataset
   State : ACTIVE
   Records: 250

⚠️  WARNING: This will permanently delete the dataset
   and ALL its annotations. This cannot be undone!

   (In production, add: input('Type CONFIRM to proceed: '))

🗑️  Deleting dataset...
   Delete request submitted.

⏳ Waiting for deletion to complete...
   State: DELETING
   State: DELETING

✅ Dataset deleted successfully!
   The dataset and all its annotations are permanently removed.

   ℹ️  Note: The raw image files in Object Storage are NOT deleted.
   Only the labeling metadata was removed from OCI Data Labeling.
❌ Deletion is PERMANENT and IRREVERSIBLE!

  • All annotations (labels) created on this dataset are permanently lost
  • Deletion cannot be undone — there is no "trash bin" or "soft delete"
  • The raw files in Object Storage are NOT deleted — only the labeling metadata
  • Always export your labeled dataset as JSON Lines first before deleting!

🏆 Best Practice: Always Export Before Deleting

📌 What does the code below do?

This code exports all labeled records from the dataset to a JSON Lines file and uploads it to Object Storage as a permanent backup — before any deletion operation. Like making a photocopy of every page before shredding a document! 📷📄
import oci
import json
import datetime

config         = oci.config.from_file()
dl_dp          = oci.data_labeling_service_dataplane.DataLabelingClient(config)
os_client      = oci.object_storage.ObjectStorageClient(config)
namespace      = os_client.get_namespace().data
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id     = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
backup_bucket  = "datalabeling-exports"

print("💾 Exporting all labeled records before any destructive operation...")

# Fetch all labeled records with pagination
all_records, page = [], None
while True:
    kwargs = dict(compartment_id=compartment_id, dataset_id=dataset_id,
                  is_labeled=True, limit=100)
    if page:
        kwargs["page"] = page
    resp = dl_dp.list_records(**kwargs)
    all_records.extend(resp.data.items)
    if resp.has_next_page:
        page = resp.next_page
    else:
        break

print(f"   Found {len(all_records)} labeled records to export")

# Build the JSONL export
export_lines = []
for record in all_records:
    annotations = dl_dp.list_annotations(
        compartment_id=compartment_id,
        dataset_id=dataset_id,
        record_id=record.id
    ).data.items

    labels = []
    for ann in annotations:
        for entity in ann.entities:
            for lbl in entity.labels:
                labels.append(lbl.label_name)

    export_lines.append(json.dumps({
        "record_id"   : record.id,
        "source_file" : record.name,
        "labels"      : labels
    }))

jsonl_content = "\n".join(export_lines).encode("utf-8")

# Upload to Object Storage as a backup
timestamp  = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
export_key = f"exports/dataset_{dataset_id[-8:]}_{timestamp}.jsonl"

os_client.put_object(namespace, backup_bucket, export_key, jsonl_content)

print(f"\n✅ Export complete!")
print(f"   Records exported : {len(all_records)}")
print(f"   Backup location  : oci://{backup_bucket}@{namespace}/{export_key}")
print(f"\n🔒 Safe to delete the dataset now — your labels are preserved!")

Output:

💾 Exporting all labeled records before any destructive operation...
   Found 187 labeled records to export

✅ Export complete!
   Records exported : 187
   Backup location  : oci://datalabeling-exports@ocitenancyexample/exports/dataset_xxxxxx_20260411_160322.jsonl

🔒 Safe to delete the dataset now — your labels are preserved!

📊 Complete Dataset Life Cycle — Full Summary Diagram

  ╔══════════════════════════════════════════════════════════════════════════╗
  ║                 OCI DATA LABELING — COMPLETE LIFECYCLE MAP               ║
  ╚══════════════════════════════════════════════════════════════════════════╝

  CREATE dataset
       │
       ▼
  ┌──────────┐   (OCI scans bucket, builds record list)
  │ CREATING │──────────────────────────────────────────────────► 2-5 min
  └────┬─────┘
       │
       ▼
  ┌──────────┐ ◄─────────────── This is where 95% of your time is spent!
  │  ACTIVE  │
  └────┬─────┘
       │
  ┌────┴──────────────────────────────────────────────────────────────┐
  │  OPERATIONS AVAILABLE IN ACTIVE STATE:                            │
  │                                                                   │
  │  ➕ add_dataset_labels()        ─── Add new label classes         │
  │  🗑️  remove_dataset_labels()    ─── Remove a label class          │
  │  📥 generate_dataset_records()  ─── Import more files (IMPORTING) │
  │  🏷️  create_annotation()        ─── Label a record                │
  │  ✏️  update_annotation()        ─── Fix a label                   │
  │  🗑️  delete_annotation()        ─── Remove a label                │
  │  📊 get_dataset()               ─── Check statistics              │
  │  🔍 list_records()              ─── Query labeled/unlabeled        │
  │  📤 list_annotations()          ─── Export labels                 │
  │  🗑️  delete_dataset()           ─── Begin permanent deletion       │
  └───────────────────────────────────────────────────────────────────┘
       │
       ├────────────────────────────────────► NEEDS_ATTENTION
       │                                      (bucket missing/
       │                                       permission issue)
       │
       └────────────────────────────────────► DELETING ──► (gone forever)

  INDIVIDUAL RECORD STATES (inside the dataset):
  ─────────────────────────────────────────────────
  ACTIVE + no annotation  =  Unlabeled (queue for labelers)
  ACTIVE + annotation     =  Labeled   (done ✅)
  FAILED                  =  File inaccessible (investigate!)

  ANNOTATION STATES (the actual label):
  ─────────────────────────────────────────────────
  ACTIVE   = Current valid label
  DELETED  = Old/corrected label (kept for audit trail)

📝 Quick Summary — What We Learned

  • CREATING → Dataset being provisioned — poll until ACTIVE before doing anything
  • ACTIVE → All operations available — label, query, export, add labels, import more files
  • NEEDS_ATTENTION → Source bucket issue — check IAM permissions and bucket existence
  • IMPORTING → Bulk file import running — existing records still labelable
  • DELETING → Permanent, irreversible — always export labels first!
  • Record states → ACTIVE (accessible) vs FAILED (file missing)
  • Annotation lifecycle → Create → Update → Delete — each change tracked for audit
  • Label schema → Can ADD labels anytime; REMOVE requires OCI to delete all matching annotations

Every AI project lives and dies by the quality of its labeled data. 🏆
Mastering the Dataset Life Cycle means you always know what state your data is in, what you can do with it, and how to protect it from accidental loss.🔄📦✨

Comments