Imagine you are baking a cake🎂
You go through stages: gather ingredients → mix them → bake → decorate → serve → finally clean up the kitchen.
A Dataset in OCI Data Labeling goes through a very similar journey —
from being created, to being actively labeled, to being completed, saved, and eventually retired.
This journey is called the Dataset Life Cycle.
Let's master every stage, step by step!
🔄 What is a Dataset Life Cycle?
Every OCI Data Labeling Dataset moves through a defined set of states during its lifetime — from the moment it is created to the moment it is deleted.
Understanding these states is critical because:
- 📋 Each state controls what actions you can perform on the dataset
- 🤖 Your code needs to check a dataset's state before calling APIs on it
- 💰 Inactive datasets still consume Object Storage — knowing when to delete saves cost
- 🔐 Security and access controls behave differently depending on the state
🗺️ The Complete Dataset Life Cycle — All States at a Glance
┌─────────────────────────────────────────────────────────────────────────┐
│ OCI DATA LABELING — DATASET LIFE CYCLE │
└─────────────────────────────────────────────────────────────────────────┘
CREATE
│
▼
┌───────────┐
│ CREATING │ ◄── OCI is setting up the dataset in the background
└─────┬─────┘
│ (auto-transitions in seconds to minutes)
▼
┌───────────┐ ┌────────────┐
│ ACTIVE │ ────► │ NEEDS │ ◄── Needs attention / data issue
│ (normal │ │ ATTENTION │
│ state) │ ◄──── └────────────┘
└─────┬─────┘
│
┌──────┼────────────────────┐
│ │ │
▼ ▼ ▼
┌──────┐ ┌──────────────┐ ┌──────────┐
│UPDATE│ │ DELETING │ │ IMPORTING│ ◄── Bulk import in progress
│ (add │ │ (permanent) │ └──────────┘
│labels│ └──────────────┘
│ or │
│records│
└──────┘
STATE SUMMARY:
─────────────────────────────────────────────────────────────────
CREATING → Dataset is being provisioned (read-only, wait)
ACTIVE → Fully operational — label, query, export, update
NEEDS ATTENTION → Something needs fixing (e.g. source bucket issue)
DELETING → Permanent deletion in progress
IMPORTING → Bulk record import is running (partial read-only)
─────────────────────────────────────────────────────────────────
📗 CREATING = The book is being printed — not available yet
📘 ACTIVE = Book is on the shelf — anyone can read or borrow it
📙 NEEDS ATTENTION = Book is damaged — needs repair before use
📕 IMPORTING = New pages being added — can read but not fully complete
🗑️ DELETING = Book is being shredded — gone forever
1️⃣ State: CREATING
When you first call the Create Dataset API,
the dataset enters the CREATING state immediately.
OCI is doing background work:
- 🔍 Scanning the source Object Storage bucket to find all files
- 📋 Building the initial list of records (one record per file)
- 🏷️ Setting up the label schema (the list of labels you defined)
- 🔐 Applying access controls and metadata
This code creates a new dataset and then watches it using a loop — checking every few seconds until the dataset finishes the CREATING stage and becomes ACTIVE. Think of it like watching the oven light — waiting for it to turn off so you know the cake is ready! 🎂⏳
import oci
import time
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
# Step 1: Submit the Create Dataset request
print("📦 Creating dataset...")
response = dl_client.create_dataset(
oci.data_labeling_service.models.CreateDatasetDetails(
display_name="Dog-Cat-Bird Image Classifier Dataset",
compartment_id=compartment_id,
annotation_format="SINGLE_LABEL",
dataset_format_details=oci.data_labeling_service.models.ImageDatasetFormatDetails(
format_type="IMAGE"
),
dataset_source_details=oci.data_labeling_service.models.ObjectStorageSourceDetails(
source_type="OBJECT_STORAGE",
namespace="your_namespace",
bucket="raw-training-images",
prefix=""
),
label_set=oci.data_labeling_service.models.LabelSet(items=[
oci.data_labeling_service.models.Label(name="Dog"),
oci.data_labeling_service.models.Label(name="Cat"),
oci.data_labeling_service.models.Label(name="Bird")
])
)
)
dataset_id = response.data.id
current_state = response.data.lifecycle_state
print(f" Dataset ID : {dataset_id}")
print(f" Initial State : {current_state}")
# Step 2: Poll until the dataset leaves CREATING state
print(f"\n⏳ Waiting for dataset to become ACTIVE...")
while current_state == "CREATING":
time.sleep(5)
dataset = dl_client.get_dataset(dataset_id)
current_state = dataset.data.lifecycle_state
record_count = dataset.data.dataset_statistics.record_count if dataset.data.dataset_statistics else "?"
print(f" State: {current_state:20} | Records found so far: {record_count}")
if current_state == "ACTIVE":
final = dl_client.get_dataset(dataset_id).data
print(f"\n✅ Dataset is now ACTIVE!")
print(f" Total records : {final.dataset_statistics.record_count}")
print(f" Labeled : {final.dataset_statistics.labeled_record_count}")
print(f" Unlabeled : {final.dataset_statistics.unlabeled_record_count}")
elif current_state == "NEEDS_ATTENTION":
print(f"\n⚠️ Dataset needs attention — check your source bucket!")
Output:
📦 Creating dataset... Dataset ID : ocid1.datalabelingdataset.oc1..examplexxxxxx Initial State : CREATING ⏳ Waiting for dataset to become ACTIVE... State: CREATING | Records found so far: 0 State: CREATING | Records found so far: 87 State: CREATING | Records found so far: 203 State: ACTIVE | Records found so far: 250 ✅ Dataset is now ACTIVE! Total records : 250 Labeled : 0 Unlabeled : 250
2️⃣ State: ACTIVE — The Working State
ACTIVE is the most important state —
it is where your dataset spends most of its life.
In this state, you can do everything:
- 🏷️ Create annotations — add labels to individual records
- 🔍 Query records — list labeled, unlabeled, or specific records
- 📊 Check statistics — see progress, label distribution, completion percentage
- ➕ Add new labels — extend the label schema with additional classes
- 📥 Import records — bulk-add new files from Object Storage
- 📤 Export — download the labeled dataset for model training
- ✏️ Update annotations — fix incorrect labels
- 🗑️ Delete records — remove specific files from the dataset
If a dataset is ACTIVE, you are free to work on it. All APIs are available. This is the green light state! 🟢
Reading Dataset Statistics (ACTIVE state)
This code reads the current statistics of your ACTIVE dataset — how many records are there, how many are labeled, and what percentage is complete. Think of it as checking a progress bar on a download — you want to know how much is done and how much is left! 📊
import oci
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
dl_dp = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
# Fetch the dataset details
dataset = dl_client.get_dataset(dataset_id).data
# Confirm it is ACTIVE before we proceed
print(f"📋 Dataset Status Report")
print(f"{'='*45}")
print(f" Name : {dataset.display_name}")
print(f" State : {dataset.lifecycle_state}")
print(f" Format : {dataset.dataset_format_details.format_type}")
print(f" Annotation : {dataset.annotation_format}")
stats = dataset.dataset_statistics
if stats:
total = stats.record_count
labeled = stats.labeled_record_count
unlabeled = stats.unlabeled_record_count
pct_complete = (labeled / total * 100) if total > 0 else 0
print(f"\n 📊 Progress:")
print(f" Total records : {total:,}")
print(f" Labeled : {labeled:,}")
print(f" Unlabeled : {unlabeled:,}")
print(f" Completion : {pct_complete:.1f}%")
# Draw a simple ASCII progress bar
filled = int(pct_complete / 5)
bar = "█" * filled + "░" * (20 - filled)
print(f"\n [{bar}] {pct_complete:.1f}% complete")
# Also get per-label breakdown
print(f"\n 🏷️ Label Distribution:")
annotations_summary = dl_dp.summarize_annotation(
dataset_id=dataset_id,
compartment_id=compartment_id,
request_type="COUNT"
)
# Count per label
label_counts = {}
for item in annotations_summary.data.items:
for entity in item.entities:
for label in entity.labels:
lname = label.label_name
label_counts[lname] = label_counts.get(lname, 0) + 1
for label, count in sorted(label_counts.items(), key=lambda x: -x[1]):
bar = "▓" * min(count, 30)
print(f" {label:<15 :="" count:="">5} {bar}")
15>
Output:
📋 Dataset Status Report ============================================= Name : Dog-Cat-Bird Image Classifier Dataset State : ACTIVE Format : IMAGE Annotation : SINGLE_LABEL 📊 Progress: Total records : 250 Labeled : 187 Unlabeled : 63 Completion : 74.8% [███████████████░░░░░] 74.8% complete 🏷️ Label Distribution: Dog : 82 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ Cat : 71 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ Bird : 34 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓
3️⃣ Life Cycle Operation: Adding New Labels to an ACTIVE Dataset
Imagine you started labeling images as "Dog" and "Cat" but then realised you also have rabbit photos in the dataset! In the ACTIVE state, you can add new labels to the existing schema without disturbing any records already labeled. 🐰
This code adds a new label ("Rabbit") to an existing dataset that already has Dog, Cat, Bird labels. It is like adding a new shelf section in a library without disturbing the existing books — all old labels stay exactly as they were! 📚➕
import oci
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
# Add a new label to an ACTIVE dataset
print("➕ Adding new label 'Rabbit' to dataset...")
response = dl_client.add_dataset_labels(
dataset_id=dataset_id,
add_dataset_labels_details=oci.data_labeling_service.models.AddDatasetLabelsDetails(
label_set=oci.data_labeling_service.models.LabelSet(
items=[
oci.data_labeling_service.models.Label(name="Rabbit")
]
)
)
)
print(f"✅ New label added!")
print(f" Work Request ID : {response.headers.get('opc-work-request-id', 'N/A')}")
print(f"\n Current label set now includes:")
updated_dataset = dl_client.get_dataset(dataset_id).data
for label in updated_dataset.label_set.items:
print(f" · {label.name}")
Output:
➕ Adding new label 'Rabbit' to dataset... ✅ New label added! Work Request ID : ocid1.workrequest.oc1..examplexxxxxx Current label set now includes: · Dog · Cat · Bird · Rabbit
OCI Data Labeling does not allow deleting a label that already exists in a dataset. If you made a typo (e.g. "Rabit" instead of "Rabbit"), you cannot fix it — you must either work around it or delete and recreate the dataset. Always double-check your label names before creating the dataset! 🔤
4️⃣ Life Cycle Operation: Importing Records (IMPORTING state)
When you need to add more files to an existing ACTIVE dataset
(say, your team found 500 more training images),
you trigger a bulk record import.
During this process, the dataset enters a temporary IMPORTING sub-state.
IMPORTING State Flow:
─────────────────────────────────────────────────────────────
ACTIVE ──► [Trigger import] ──► IMPORTING ──► ACTIVE
(background job runs,
new records appear
one by one as
they are added)
─────────────────────────────────────────────────────────────
⚠️ During IMPORTING: you CAN still label existing records
you CANNOT modify dataset metadata
This code triggers a bulk import of new image files from an Object Storage prefix into an existing dataset. It is like a delivery truck bringing new books to a library — the library stays open while the books are being unloaded and shelved, but the catalogue is not finalised until the delivery is complete! 🚚📚
import oci
import time
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
dl_dp = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
# Step 1: Trigger bulk record generation (import from Object Storage prefix)
print("📥 Triggering bulk record import from Object Storage...")
import_response = dl_client.generate_dataset_records(
dataset_id=dataset_id,
generate_dataset_records_details=oci.data_labeling_service.models.GenerateDatasetRecordsDetails(
limit=500.0 # import up to 500 new files
)
)
work_request_id = import_response.headers.get("opc-work-request-id")
print(f" Import job submitted!")
print(f" Work Request ID : {work_request_id}")
# Step 2: Poll the work request status
wrc = oci.work_requests.WorkRequestClient(config)
print(f"\n⏳ Monitoring import progress...")
while True:
wr = wrc.get_work_request(work_request_id)
state = wr.data.status
pct = wr.data.percent_complete
bar = "█" * int(pct / 5) + "░" * (20 - int(pct / 5))
print(f" [{bar}] {pct:.0f}% Status: {state}")
if state == "SUCCEEDED":
print(f"\n✅ Import complete!")
break
elif state in ["FAILED", "CANCELED"]:
print(f"\n❌ Import failed with state: {state}")
for err in wr.data.errors:
print(f" Error: {err.message}")
break
time.sleep(10)
# Step 3: Check updated record count
dataset = dl_client.get_dataset(dataset_id).data
print(f"\n Updated Total Records : {dataset.dataset_statistics.record_count:,}")
print(f" Newly Imported : {dataset.dataset_statistics.unlabeled_record_count:,} unlabeled")
Output:
📥 Triggering bulk record import from Object Storage... Import job submitted! Work Request ID : ocid1.workrequest.oc1..examplexxxxxx ⏳ Monitoring import progress... [░░░░░░░░░░░░░░░░░░░░] 0% Status: IN_PROGRESS [████░░░░░░░░░░░░░░░░] 20% Status: IN_PROGRESS [████████░░░░░░░░░░░░] 40% Status: IN_PROGRESS [████████████░░░░░░░░] 60% Status: IN_PROGRESS [████████████████░░░░] 80% Status: IN_PROGRESS [████████████████████]100% Status: SUCCEEDED ✅ Import complete! Updated Total Records : 750 Newly Imported : 500 unlabeled
5️⃣ State: NEEDS_ATTENTION — When Something Goes Wrong
The NEEDS_ATTENTION state is OCI's way of tapping you on the shoulder and saying:
"Hey, something is not right — please check this dataset!" 👋
Common reasons a dataset enters this state:
- 🪣 The source Object Storage bucket was deleted or renamed after the dataset was created
- 🔐 The IAM permissions changed and OCI can no longer access the source bucket
- 📁 The source prefix no longer exists or was moved
- 🔄 A background import job failed partway through
This code checks if any datasets in your compartment have entered the NEEDS_ATTENTION state and prints a helpful diagnosis for each one — like a nurse checking all patients in a ward to see who needs immediate care! 🏥
import oci
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
print("🔍 Scanning all datasets for issues...")
print("=" * 55)
# List ALL datasets in the compartment
all_datasets = dl_client.list_datasets(
compartment_id=compartment_id,
limit=100
).data.items
needs_attention = [ds for ds in all_datasets if ds.lifecycle_state == "NEEDS_ATTENTION"]
active_count = len([ds for ds in all_datasets if ds.lifecycle_state == "ACTIVE"])
print(f" Total datasets : {len(all_datasets)}")
print(f" ACTIVE : {active_count} ✅")
print(f" NEEDS_ATTENTION : {len(needs_attention)} ⚠️")
if needs_attention:
print(f"\n⚠️ Datasets requiring your attention:")
for ds in needs_attention:
print(f"\n Dataset : {ds.display_name}")
print(f" ID : {ds.id}")
print(f" State : {ds.lifecycle_state}")
# Fetch detailed info for diagnosis
detail = dl_client.get_dataset(ds.id).data
src = detail.dataset_source_details
print(f"\n 📁 Source details:")
print(f" Bucket : {src.bucket}")
print(f" Namespace : {src.namespace}")
print(f" Prefix : '{src.prefix or '(all files)'}'")
print(f"\n 🔧 Suggested fixes:")
print(f" 1. Confirm bucket '{src.bucket}' still exists in Object Storage")
print(f" 2. Confirm your IAM policies still allow access to this bucket")
print(f" 3. Check for any recent changes to the bucket permissions")
print(f" 4. Contact your OCI admin if the bucket was deleted")
else:
print(f"\n✅ All datasets are healthy! No attention needed.")
Output (when a dataset has an issue):
🔍 Scanning all datasets for issues...
=======================================================
Total datasets : 4
ACTIVE : 3 ✅
NEEDS_ATTENTION : 1 ⚠️
⚠️ Datasets requiring your attention:
Dataset : Product Defect Detection Dataset
ID : ocid1.datalabelingdataset.oc1..exampleyyyyyyy
State : NEEDS_ATTENTION
📁 Source details:
Bucket : factory-raw-images-old
Namespace : ocitenancyexample
Prefix : 'products/2025/'
🔧 Suggested fixes:
1. Confirm bucket 'factory-raw-images-old' still exists in Object Storage
2. Confirm your IAM policies still allow access to this bucket
3. Check for any recent changes to the bucket permissions
4. Contact your OCI admin if the bucket was deleted
6️⃣ Record Life Cycle — Inside the Dataset
Within an ACTIVE dataset, each individual Record also has its own lifecycle! A record is one file (one image or one text snippet) in the dataset.
RECORD LIFECYCLE (inside an ACTIVE dataset):
──────────────────────────────────────────────────────────────────
State │ Meaning
─────────────────┼────────────────────────────────────────────────
ACTIVE │ Record exists, ready to be labeled or queried
FAILED │ OCI could not access this file in Object Storage
│ (file was deleted or moved after dataset creation)
──────────────────────────────────────────────────────────────────
And each record can have one of these ANNOTATION states:
──────────────────────────────────────────────────────────────────
(no annotation) │ Not yet labeled — shows as Unlabeled
ACTIVE │ Has at least one label attached — Labeled
──────────────────────────────────────────────────────────────────
Querying Records by State
This code queries the dataset to find: all records that have been labeled (so you can review them), all records that still need labeling (so your team can focus on them), and all records that have FAILED (so you can investigate missing files). Like a school roll-call that shows who is present, absent, or has a problem! 📋
import oci
config = oci.config.from_file()
dl_dp = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
def fetch_all_records(is_labeled, lifecycle_state=None):
"""Fetch all records matching the given filters (handles pagination)."""
all_items = []
page_token = None
kwargs = dict(
compartment_id=compartment_id,
dataset_id=dataset_id,
is_labeled=is_labeled,
limit=100
)
if lifecycle_state:
kwargs["lifecycle_state"] = lifecycle_state
while True:
if page_token:
kwargs["page"] = page_token
resp = dl_dp.list_records(**kwargs)
all_items.extend(resp.data.items)
if resp.has_next_page:
page_token = resp.next_page
else:
break
return all_items
# Query each category of records
print("📋 Record Status Report for Dataset")
print("=" * 50)
labeled_records = fetch_all_records(is_labeled=True)
unlabeled_records = fetch_all_records(is_labeled=False)
failed_records = fetch_all_records(is_labeled=False, lifecycle_state="FAILED")
total = len(labeled_records) + len(unlabeled_records)
pct = len(labeled_records) / total * 100 if total > 0 else 0
print(f"\n ✅ Labeled records : {len(labeled_records):>5}")
print(f" ⏳ Unlabeled records : {len(unlabeled_records):>5}")
print(f" ❌ Failed records : {len(failed_records):>5}")
print(f" 📊 Completion : {pct:.1f}%")
if failed_records:
print(f"\n ❌ Failed records (files no longer accessible):")
for rec in failed_records[:5]: # show first 5
print(f" · {rec.name} (ID: {rec.id})")
if len(failed_records) > 5:
print(f" ... and {len(failed_records)-5} more")
# Show a sample of unlabeled records (next ones to label)
if unlabeled_records:
print(f"\n ⏳ Next records to label (sample of 5):")
for rec in unlabeled_records[:5]:
print(f" · {rec.name}")
Output:
📋 Record Status Report for Dataset
==================================================
✅ Labeled records : 187
⏳ Unlabeled records : 63
❌ Failed records : 2
📊 Completion : 74.8%
❌ Failed records (files no longer accessible):
· deleted_image_001.jpg (ID: ocid1.datalabelingrecord...001)
· moved_cat_photo.png (ID: ocid1.datalabelingrecord...002)
⏳ Next records to label (sample of 5):
· dog_batch2_001.jpg
· dog_batch2_002.jpg
· bird_outdoor_007.jpg
· cat_indoor_012.jpg
· unknown_animal_001.jpg
7️⃣ Annotation Life Cycle — The Label on Each Record
An Annotation is the actual label (or set of labels) attached to a record. Annotations also have their own lifecycle within a dataset!
ANNOTATION LIFECYCLE:
──────────────────────────────────────────────────────────────────
Action │ Result
─────────────────────┼────────────────────────────────────────────
Create annotation │ Record becomes "Labeled"
Update annotation │ Label changed (old annotation deleted,
│ new one created in ACTIVE state)
Delete annotation │ Record returns to "Unlabeled"
──────────────────────────────────────────────────────────────────
Annotation states:
ACTIVE = The currently valid label for this record
DELETED = A previous label that was corrected/removed
──────────────────────────────────────────────────────────────────
This code shows the complete annotation lifecycle operations: create a label, update it (fix a mistake), and delete it (remove a label from a record). Think of it like writing, erasing, and rewriting answers on an exam paper! ✏️📝
import oci
config = oci.config.from_file()
dl_dp = oci.data_labeling_service_dataplane.DataLabelingClient(config)
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
record_id = "ocid1.datalabelingrecord.oc1..examplerecord1"
# ── OPERATION 1: CREATE an annotation (label a record for the first time)
print("1️⃣ Creating annotation (labeling record as 'Dog')...")
create_resp = dl_dp.create_annotation(
oci.data_labeling_service_dataplane.models.CreateAnnotationDetails(
record_id=record_id,
compartment_id=compartment_id,
entities=[
oci.data_labeling_service_dataplane.models.ClassificationLabel(
entity_type="GENERIC",
labels=[
oci.data_labeling_service_dataplane.models.Label(label="Dog")
]
)
]
)
)
annotation_id = create_resp.data.id
print(f" ✅ Annotation created!")
print(f" Annotation ID : {annotation_id}")
print(f" Label : Dog")
print(f" State : {create_resp.data.lifecycle_state}")
# ── OPERATION 2: UPDATE the annotation (oops — it was actually a Cat!)
print(f"\n2️⃣ Updating annotation (correcting label to 'Cat')...")
update_resp = dl_dp.update_annotation(
annotation_id=annotation_id,
update_annotation_details=oci.data_labeling_service_dataplane.models.UpdateAnnotationDetails(
entities=[
oci.data_labeling_service_dataplane.models.ClassificationLabel(
entity_type="GENERIC",
labels=[
oci.data_labeling_service_dataplane.models.Label(label="Cat")
]
)
]
)
)
print(f" ✅ Annotation updated!")
print(f" New label : Cat (was: Dog)")
print(f" State : {update_resp.data.lifecycle_state}")
# ── OPERATION 3: DELETE the annotation (record is too blurry to label reliably)
print(f"\n3️⃣ Deleting annotation (image too blurry to label confidently)...")
dl_dp.delete_annotation(annotation_id=annotation_id)
print(f" ✅ Annotation deleted!")
print(f" Record is now back to: Unlabeled")
print(f" It will appear in the unlabeled queue again for a fresh labeler.")
Output:
1️⃣ Creating annotation (labeling record as 'Dog')... ✅ Annotation created! Annotation ID : ocid1.datalabelingannotation.oc1..exampleanno1 Label : Dog State : ACTIVE 2️⃣ Updating annotation (correcting label to 'Cat')... ✅ Annotation updated! New label : Cat (was: Dog) State : ACTIVE 3️⃣ Deleting annotation (image too blurry to label confidently)... ✅ Annotation deleted! Record is now back to: Unlabeled It will appear in the unlabeled queue again for a fresh labeler.
8️⃣ Life Cycle Operation: Removing a Label from the Schema
When a label is no longer needed in the dataset — for example, you merged "Puppy" and "Dog" into just "Dog" — you can remove a label from the label schema.
This code removes a label from the dataset's label schema. OCI will also automatically delete all existing annotations that used this label — those records become "Unlabeled" again and need to be re-labeled. Think of it like removing a subject from the school curriculum: all students who were enrolled in that subject need to re-enrol in a new one! 🏫
import oci
import time
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
label_to_remove = "Puppy" # we are merging this into "Dog"
print(f"🗑️ Removing label '{label_to_remove}' from dataset...")
print(f" ⚠️ All records labeled '{label_to_remove}' will become Unlabeled!")
response = dl_client.remove_dataset_labels(
dataset_id=dataset_id,
remove_dataset_labels_details=oci.data_labeling_service.models.RemoveDatasetLabelsDetails(
label_set=oci.data_labeling_service.models.LabelSet(
items=[oci.data_labeling_service.models.Label(name=label_to_remove)]
)
)
)
# Monitor the work request
work_request_id = response.headers.get("opc-work-request-id")
wrc = oci.work_requests.WorkRequestClient(config)
print(f"\n⏳ Waiting for label removal to complete...")
while True:
wr = wrc.get_work_request(work_request_id)
state = wr.data.status
pct = wr.data.percent_complete
print(f" {state} {pct:.0f}%")
if state == "SUCCEEDED":
break
elif state in ["FAILED", "CANCELED"]:
print(f"❌ Label removal failed!")
break
time.sleep(5)
print(f"\n✅ Label '{label_to_remove}' removed!")
print(f" Affected records are now Unlabeled.")
print(f" Go to the labeling UI and re-label them as 'Dog'.")
9️⃣ State: DELETING — The Final Stage
When you no longer need a dataset —
maybe the project is complete, or the data was moved elsewhere —
you delete it.
The dataset enters the DELETING state while OCI cleans up all its metadata.
This code deletes a dataset from OCI Data Labeling. It first checks the dataset is in ACTIVE state (safe to delete), then submits the delete request and monitors until it is fully gone. Think of it like returning a rented lab space — you clean everything out before handing back the keys! 🔑🏢
import oci
import time
config = oci.config.from_file()
dl_client = oci.data_labeling_service.DataLabelingManagementClient(config)
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
# Step 1: Confirm the dataset state before deleting
dataset = dl_client.get_dataset(dataset_id).data
print(f"📋 Dataset to delete:")
print(f" Name : {dataset.display_name}")
print(f" State : {dataset.lifecycle_state}")
print(f" Records: {dataset.dataset_statistics.record_count:,}")
if dataset.lifecycle_state != "ACTIVE":
print(f"\n⚠️ Cannot delete — dataset is in '{dataset.lifecycle_state}' state.")
print(f" Only ACTIVE datasets can be deleted.")
exit()
# Step 2: Safety confirmation (in a real script, ask for user input!)
print(f"\n⚠️ WARNING: This will permanently delete the dataset")
print(f" and ALL its annotations. This cannot be undone!")
print(f"\n (In production, add: input('Type CONFIRM to proceed: '))")
# Step 3: Delete the dataset
print(f"\n🗑️ Deleting dataset...")
response = dl_client.delete_dataset(dataset_id=dataset_id)
print(f" Delete request submitted.")
# Step 4: Monitor until deletion completes
print(f"\n⏳ Waiting for deletion to complete...")
while True:
try:
ds = dl_client.get_dataset(dataset_id)
state = ds.data.lifecycle_state
print(f" State: {state}")
if state == "DELETED":
break
time.sleep(5)
except oci.exceptions.ServiceError as e:
if e.status == 404:
# 404 means the dataset no longer exists — deletion complete!
print(f"\n✅ Dataset deleted successfully!")
print(f" The dataset and all its annotations are permanently removed.")
print(f"\n ℹ️ Note: The raw image files in Object Storage are NOT deleted.")
print(f" Only the labeling metadata was removed from OCI Data Labeling.")
break
else:
raise
Output:
📋 Dataset to delete:
Name : Old Product Test Dataset
State : ACTIVE
Records: 250
⚠️ WARNING: This will permanently delete the dataset
and ALL its annotations. This cannot be undone!
(In production, add: input('Type CONFIRM to proceed: '))
🗑️ Deleting dataset...
Delete request submitted.
⏳ Waiting for deletion to complete...
State: DELETING
State: DELETING
✅ Dataset deleted successfully!
The dataset and all its annotations are permanently removed.
ℹ️ Note: The raw image files in Object Storage are NOT deleted.
Only the labeling metadata was removed from OCI Data Labeling.
- All annotations (labels) created on this dataset are permanently lost
- Deletion cannot be undone — there is no "trash bin" or "soft delete"
- The raw files in Object Storage are NOT deleted — only the labeling metadata
- Always export your labeled dataset as JSON Lines first before deleting!
🏆 Best Practice: Always Export Before Deleting
This code exports all labeled records from the dataset to a JSON Lines file and uploads it to Object Storage as a permanent backup — before any deletion operation. Like making a photocopy of every page before shredding a document! 📷📄
import oci
import json
import datetime
config = oci.config.from_file()
dl_dp = oci.data_labeling_service_dataplane.DataLabelingClient(config)
os_client = oci.object_storage.ObjectStorageClient(config)
namespace = os_client.get_namespace().data
compartment_id = "ocid1.compartment.oc1..your_compartment_id"
dataset_id = "ocid1.datalabelingdataset.oc1..examplexxxxxx"
backup_bucket = "datalabeling-exports"
print("💾 Exporting all labeled records before any destructive operation...")
# Fetch all labeled records with pagination
all_records, page = [], None
while True:
kwargs = dict(compartment_id=compartment_id, dataset_id=dataset_id,
is_labeled=True, limit=100)
if page:
kwargs["page"] = page
resp = dl_dp.list_records(**kwargs)
all_records.extend(resp.data.items)
if resp.has_next_page:
page = resp.next_page
else:
break
print(f" Found {len(all_records)} labeled records to export")
# Build the JSONL export
export_lines = []
for record in all_records:
annotations = dl_dp.list_annotations(
compartment_id=compartment_id,
dataset_id=dataset_id,
record_id=record.id
).data.items
labels = []
for ann in annotations:
for entity in ann.entities:
for lbl in entity.labels:
labels.append(lbl.label_name)
export_lines.append(json.dumps({
"record_id" : record.id,
"source_file" : record.name,
"labels" : labels
}))
jsonl_content = "\n".join(export_lines).encode("utf-8")
# Upload to Object Storage as a backup
timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
export_key = f"exports/dataset_{dataset_id[-8:]}_{timestamp}.jsonl"
os_client.put_object(namespace, backup_bucket, export_key, jsonl_content)
print(f"\n✅ Export complete!")
print(f" Records exported : {len(all_records)}")
print(f" Backup location : oci://{backup_bucket}@{namespace}/{export_key}")
print(f"\n🔒 Safe to delete the dataset now — your labels are preserved!")
Output:
💾 Exporting all labeled records before any destructive operation... Found 187 labeled records to export ✅ Export complete! Records exported : 187 Backup location : oci://datalabeling-exports@ocitenancyexample/exports/dataset_xxxxxx_20260411_160322.jsonl 🔒 Safe to delete the dataset now — your labels are preserved!
📊 Complete Dataset Life Cycle — Full Summary Diagram
╔══════════════════════════════════════════════════════════════════════════╗
║ OCI DATA LABELING — COMPLETE LIFECYCLE MAP ║
╚══════════════════════════════════════════════════════════════════════════╝
CREATE dataset
│
▼
┌──────────┐ (OCI scans bucket, builds record list)
│ CREATING │──────────────────────────────────────────────────► 2-5 min
└────┬─────┘
│
▼
┌──────────┐ ◄─────────────── This is where 95% of your time is spent!
│ ACTIVE │
└────┬─────┘
│
┌────┴──────────────────────────────────────────────────────────────┐
│ OPERATIONS AVAILABLE IN ACTIVE STATE: │
│ │
│ ➕ add_dataset_labels() ─── Add new label classes │
│ 🗑️ remove_dataset_labels() ─── Remove a label class │
│ 📥 generate_dataset_records() ─── Import more files (IMPORTING) │
│ 🏷️ create_annotation() ─── Label a record │
│ ✏️ update_annotation() ─── Fix a label │
│ 🗑️ delete_annotation() ─── Remove a label │
│ 📊 get_dataset() ─── Check statistics │
│ 🔍 list_records() ─── Query labeled/unlabeled │
│ 📤 list_annotations() ─── Export labels │
│ 🗑️ delete_dataset() ─── Begin permanent deletion │
└───────────────────────────────────────────────────────────────────┘
│
├────────────────────────────────────► NEEDS_ATTENTION
│ (bucket missing/
│ permission issue)
│
└────────────────────────────────────► DELETING ──► (gone forever)
INDIVIDUAL RECORD STATES (inside the dataset):
─────────────────────────────────────────────────
ACTIVE + no annotation = Unlabeled (queue for labelers)
ACTIVE + annotation = Labeled (done ✅)
FAILED = File inaccessible (investigate!)
ANNOTATION STATES (the actual label):
─────────────────────────────────────────────────
ACTIVE = Current valid label
DELETED = Old/corrected label (kept for audit trail)
📝 Quick Summary — What We Learned
- CREATING → Dataset being provisioned — poll until ACTIVE before doing anything
- ACTIVE → All operations available — label, query, export, add labels, import more files
- NEEDS_ATTENTION → Source bucket issue — check IAM permissions and bucket existence
- IMPORTING → Bulk file import running — existing records still labelable
- DELETING → Permanent, irreversible — always export labels first!
- Record states → ACTIVE (accessible) vs FAILED (file missing)
- Annotation lifecycle → Create → Update → Delete — each change tracked for audit
- Label schema → Can ADD labels anytime; REMOVE requires OCI to delete all matching annotations
Every AI project lives and dies by the quality of its labeled data. 🏆
Mastering the Dataset Life Cycle means you always know what state your data is in,
what you can do with it, and how to protect it from accidental loss.🔄📦✨
Comments
Post a Comment