Skip to main content

Data Labelling — The Secret Ingredient That Makes AI Smart

Calculating read time…

Today we are going to learn about something that almost nobody talks about — but without it, no AI system in the world would work.

It is called Data Labelling. By the end of this post, you will understand it so completely that you could explain it to your parents over dinner! 🍽️

💡 Did You Know?
Data preparation and labelling takes up roughly 80% of the total time in any real AI project. The model training part — which everyone talks about — is only 20%! Labelling is the hidden giant of Machine Learning. 🐘


What is Data Labelling?

Imagine you are teaching your younger brother or sister the difference between a cat and a dog. 🐱🐶 You pick up a photo and say: "This is a CAT." Then another photo: "This is a DOG."

You are doing this again and again — hundreds, thousands of times — until your sibling can look at any new photo and say the correct animal name on their own.

Data Labelling is exactly this — but for computers!

You take a raw piece of data (a photo, an email, a sound clip) and you put a label (tag/answer) on it so the computer knows what it is. After seeing thousands of labelled examples, the computer learns on its own!

  BEFORE Labelling:                AFTER Labelling:
  ┌──────────────────┐             ┌──────────────────┐
  │  📷 Photo of cat │    ──────►  │  📷 Photo of cat │
  │  (just a photo,  │             │  Label: "CAT" 🏷️ │
  │   no meaning)    │             │  (now the AI     │
  └──────────────────┘             │   knows what it  │
                                   │   is looking at!)│
                                   └──────────────────┘
✅ Simple Definition:
Data Labelling = Adding meaningful tags, categories, or answers to raw data so that a Machine Learning model can learn from it.

🌍 Why Does Data Labelling Matter So Much?

Think of AI as a baby 👶. A baby learns by being shown things and being told what they are. The more correct examples the baby sees — the smarter it gets!

Without labels, the computer just sees random pixels, random words, random sounds. It has absolutely no idea what any of it means. Labels give it meaning.

  Raw Data (No Labels):               Labelled Data:
  ─────────────────────               ─────────────────────────────────────────
  📧 "FREE offer! Click now!"     →   📧 "FREE offer! Click now!"  [SPAM 🚫]
  📧 "Your invoice is attached"   →   📧 "Your invoice is attached" [NOT SPAM ✅]
  📧 "Team meeting at 3pm"        →   📧 "Team meeting at 3pm"      [NOT SPAM ✅]
  📧 "WIN A MILLION DOLLARS!"     →   📧 "WIN A MILLION DOLLARS!"   [SPAM 🚫]

  The AI cannot learn from left side. It CAN learn from right side!
❌ Without Labels:
Training an AI on unlabelled data is like asking someone to pass an exam without ever giving them the textbook or the answers. Impossible!

🗺️ End-to-End Flow — The Full Journey of Data Labelling

Let's trace the complete journey of data labelling — from raw data to a working AI model. Read it like a story! 📖

┌─────────────────────────────────────────────────────────────────────────────┐
│                  COMPLETE DATA LABELLING PIPELINE                           │
└─────────────────────────────────────────────────────────────────────────────┘

  STEP 1           STEP 2           STEP 3           STEP 4
  ────────         ────────         ────────         ────────
  Collect          Design           Label the        Quality
  Raw Data    →    Guidelines   →   Data        →    Check
  (photos,         (What counts     (Human or        (Is the
  text, audio,     as a cat?        AI or Both!)     label
  videos)          Edge cases?)                      correct?)
      │
      │            STEP 5           STEP 6           STEP 7
      │            ────────         ────────         ────────
      └──────►     Store &          Train the        Deploy
                   Version     →    Model       →    Model!
                   Labels           (Feed labels     (Working
                   (DVC, Git)       to algorithm)    AI App!)

Don't worry — we will cover each step in detail below with real examples! 🚀


🔹 STEP 1 — Collect Raw Data

Before you can label anything, you need the RAW DATA — the untagged, unprocessed stuff. Think of it as gathering blank answer sheets before a class exam. 📝

Types of Raw Data

  • 🖼️ Images — Photos of animals, X-rays, satellite images, product photos
  • 📝 Text — Emails, customer reviews, news articles, social media posts
  • 🔊 Audio — Voice recordings, customer calls, music clips
  • 🎬 Video — CCTV footage, dashcam recordings, movie clips
  • 📊 Tabular — Spreadsheets with rows and columns of numbers
  • 🗺️ 3D / LiDAR — Point clouds from self-driving car sensors (a trend!)
⚠️ Important Tip:
The most exciting new data types are multimodal data — combining images + text + audio together. For example, training an AI that watches a video AND listens to speech AND reads subtitles all at the same time!

🔹 STEP 2 — Design Labelling Guidelines (The Rule Book! 📖)

Before anyone starts labelling, you must write a clear rule book. Without rules, two different labellers will label the same image differently — and that makes your training data a mess!

Imagine if one teacher says a score of 35/100 is "Pass" and another teacher says it's "Fail". The students (AI model) would get totally confused! Same problem with inconsistent labels.

  EXAMPLE — Guidelines for labelling "Cats" in photos:
  ──────────────────────────────────────────────────────────────────
  ✅ Label as "CAT" if:
     - The photo contains a domestic cat (any breed, any colour)
     - The cat is the main subject of the photo
     - At least 50% of the cat's body is visible

  ❌ Do NOT label as "CAT" if:
     - It is a wild cat (lion, tiger, cheetah)
     - The cat is in a cartoon or drawing (label as "CARTOON")
     - The image is blurry and you cannot confirm it's a cat

  ⚠️ EDGE CASES (tricky ones!):
     - Cat + dog both visible → label as "MIXED" (define this!)
     - Only cat's paw visible → label as "UNCLEAR"
  ──────────────────────────────────────────────────────────────────
✅ Best Practice:
Always define edge cases in your guidelines BEFORE labelling starts. Edge cases are the tricky "what about this?" situations. Cover them early and your labels will be consistent across the whole team!

🔹 STEP 3 — Label the Data (The Core Task! 🏷️)

Now the actual labelling happens! , there are 6 different approaches to labelling — from fully manual (humans doing everything) to fully automated (AI doing everything).

The 6 Methods of Data Labelling

Method 1️⃣ — Manual / Human Labelling

A human expert looks at each piece of data and writes the label by hand. Like a teacher marking every exam paper one by one. 📝

  Human Labeller:
  ─────────────────────────────────────────────────────────
  Sees:  "This medication caused severe headaches."
  Types: Sentiment = NEGATIVE, Topic = SIDE_EFFECTS
  Done!  Next one...

  Sees:  "The drug completely cured my condition!"
  Types: Sentiment = POSITIVE, Topic = EFFECTIVENESS
  Done!  Next one...
  ─────────────────────────────────────────────────────────
  Speed: Slow (hundreds per day)
  Cost:  High ($10-$100 per hour per labeller)
  Quality: HIGH ✅ (if expert)
⚠️ When to use Manual Labelling:
Use it when data is complex, sensitive, or domain-specific — like medical images (X-rays), legal documents, or safety-critical AI systems. An expert's judgment cannot be replaced here!

Method 2️⃣ — Active Learning (Smart Labelling! 🧠)

Instead of labelling everything, the AI itself chooses WHICH samples are most confusing to it and asks a human to label only those! This saves enormous time and money.

Think of it like a student who instead of reading the entire textbook, asks the teacher: "I understand most of it — but can you explain just THESE 5 pages I'm confused about?"

  How Active Learning Works:
  ──────────────────────────────────────────────────────────────────
  Step 1: Train model on a SMALL set of labelled data
          (maybe just 100 examples)

  Step 2: Model looks at all UNLABELLED data and asks:
          "Which ones am I most UNSURE about?"
          Example: "I'm only 52% sure this email is spam"

  Step 3: Send ONLY the uncertain ones to human labellers
          (maybe 50 examples instead of 10,000!)

  Step 4: Human labels those 50 uncertain examples

  Step 5: Retrain model with new labels → it gets smarter!

  Step 6: Repeat from Step 2 → model gets confident faster!
  ──────────────────────────────────────────────────────────────────
  Result: 30-40% fewer labels needed! Same model quality! 🎯
✅ Trend:
Active learning is the fastest-growing labelling strategy in production. It cuts labelling costs by 30-40% while maintaining accuracy. Tools like modAL, Cleanlab, and Labelbox support this natively!

📋 What the code below does:
This shows a simple Active Learning loop using Python. We start with a tiny labelled dataset. The model trains on it, then figures out WHICH unlabelled samples it is most unsure about. Those uncertain samples get sent for human review first — not random ones! This way, you label less but learn more. Like studying your weak topics first instead of re-reading what you already know!

# ─────────────────────────────────────────────────────────────────
# Active Learning — Smart Labelling Strategy
# ─────────────────────────────────────────────────────────────────
# Instead of labelling 10,000 samples, we let the model CHOOSE
# which 100 samples it is most confused about.
# Then we ONLY label those 100 — saving 99% of labelling work!
# ─────────────────────────────────────────────────────────────────

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler

# ── Simulate our data ────────────────────────────────────────────
np.random.seed(42)

# 1000 total unlabelled samples (just random numbers for demo)
# In real life: these would be text features, image features, etc.
all_data = np.random.randn(1000, 10)

# We start with only 50 labelled samples (very small!)
labelled_indices = list(range(50))
X_labelled = all_data[:50]
y_labelled  = np.random.randint(0, 2, 50)   # 0 = Not Spam, 1 = Spam

# Remaining 950 samples are UNLABELLED
unlabelled_indices = list(range(50, 1000))
X_unlabelled = all_data[50:]

# ── Active Learning Loop ─────────────────────────────────────────
print("Starting Active Learning Loop...")
print(f"Initial labelled samples: {len(labelled_indices)}")

for round_num in range(1, 4):   # Do 3 rounds of active learning
    print(f"\n--- Round {round_num} ---")

    # STEP 1: Train the model on currently labelled data
    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X_labelled)
    model = LogisticRegression(max_iter=1000)
    model.fit(X_scaled, y_labelled)

    # STEP 2: Ask model: "How confident are you about each unlabelled sample?"
    # predict_proba returns probability for each class [P(Not Spam), P(Spam)]
    X_unlab_scaled = scaler.transform(X_unlabelled)
    probabilities  = model.predict_proba(X_unlab_scaled)

    # STEP 3: Find the MOST UNCERTAIN samples
    # Uncertainty = how close the probability is to 50% (= most confused!)
    # If model says 51% spam, 49% not spam → very uncertain → label this one!
    # If model says 99% spam, 1% not spam  → very confident → skip this one!
    max_probs   = np.max(probabilities, axis=1)   # Highest probability for each sample
    uncertainty = 1 - max_probs                   # Low confidence = high uncertainty

    # Pick TOP 20 most uncertain samples to send for human labelling
    top_uncertain_idx = np.argsort(uncertainty)[-20:][::-1]

    print(f"  Model trained on {len(X_labelled)} samples")
    print(f"  Top 20 uncertain samples selected for human review")
    print(f"  Most uncertain sample confidence: {max_probs[top_uncertain_idx[0]]:.2f}")

    # STEP 4: "Human" labels those 20 samples (simulated here with random labels)
    # In real life: these go to a labelling platform like Label Studio or Labelbox!
    new_labels = np.random.randint(0, 2, len(top_uncertain_idx))

    # STEP 5: Add newly labelled samples to our labelled pool
    new_X = X_unlabelled[top_uncertain_idx]
    X_labelled = np.vstack([X_labelled, new_X])
    y_labelled  = np.hstack([y_labelled, new_labels])

    # Remove labelled samples from unlabelled pool
    X_unlabelled = np.delete(X_unlabelled, top_uncertain_idx, axis=0)

    print(f"  Labelled pool now has: {len(X_labelled)} samples")
    print(f"  Unlabelled pool remaining: {len(X_unlabelled)} samples")

print("\n✅ Active learning complete!")
print(f"   Total labelled: {len(X_labelled)} (started with 50, added 60)")
print(f"   Saved labelling: {1000 - len(X_labelled)} samples! 🎉")

📋 Key lines explained:

  • predict_proba() → Returns the probability for each class. If Spam=0.51, Not Spam=0.49 → model is almost guessing! That sample needs a human label
  • uncertainty = 1 - max_probs → Samples where model is least confident get the highest uncertainty score
  • np.argsort(uncertainty)[-20:] → Gets indices of the 20 MOST uncertain samples to send for human review

Method 3️⃣ — Weak Supervision (Rules as Labels! 📐)

What if you have millions of data points and cannot possibly label them all by hand? Weak Supervision is the answer!

Instead of labelling one by one, you write simple rules (like "if the email contains the word FREE in caps → it's probably spam"). These rules are called Labelling Functions (LFs). They are not perfect — but thousands of imperfect rules combined are surprisingly accurate!

  Weak Supervision Example — Spam Detection:
  ─────────────────────────────────────────────────────────────────
  Rule 1: If email contains "FREE" or "WINNER" → SPAM (60% accurate)
  Rule 2: If email contains "@company.com" domain → NOT SPAM (80% accurate)
  Rule 3: If email has >3 exclamation marks → SPAM (70% accurate)
  Rule 4: If email is shorter than 20 words → might be SPAM (55% accurate)

  Combine all rules together using Snorkel →
  Final probabilistic label: 73% SPAM (confident enough to train on!)
  ─────────────────────────────────────────────────────────────────
  Result: Labelled 1,000,000 emails in minutes — not months!

📋 What the code below does:
We write three simple rules (called Labelling Functions). Each rule looks at an email and returns 1 (Spam), 0 (Not Spam), or -1 (I don't know — abstain!). Then Snorkel combines all three opinions and produces a final probability. Like taking a vote from three judges where each judge can say "Yes", "No", or "No opinion"!

# ─────────────────────────────────────────────────────────────────
# Weak Supervision — Label millions with rules, not human clicks!
# ─────────────────────────────────────────────────────────────────
# We write simple rules (Labelling Functions) that guess the label.
# No single rule is perfect — but combined, they're powerful!
# Tool: Snorkel (open-source weak supervision framework)
# ─────────────────────────────────────────────────────────────────

# Install with: pip install snorkel
from snorkel.labeling import labeling_function, LFApplier, LFAnalysis
from snorkel.labeling.model import LabelModel
import pandas as pd

# Define our class labels
SPAM     =  1    # This email IS spam
NOT_SPAM =  0    # This email is NOT spam
ABSTAIN  = -1    # We're not sure — skip this one

# ── LABELLING FUNCTIONS (our rules) ──────────────────────────────
# Each function receives one email text and returns SPAM, NOT_SPAM, or ABSTAIN

@labeling_function()
def lf_contains_free(email):
    """
    Rule 1: If email contains the word FREE in caps → probably spam!
    Like seeing "FREE MONEY" on a letter — usually junk mail!
    """
    return SPAM if "FREE" in email.text.upper() else ABSTAIN

@labeling_function()
def lf_exclamation_overload(email):
    """
    Rule 2: If email has 3 or more exclamation marks → probably spam!
    "WIN NOW!!! CLICK HERE!!! AMAZING OFFER!!!" ← Screams spam!
    """
    return SPAM if email.text.count("!") >= 3 else ABSTAIN

@labeling_function()
def lf_professional_domain(email):
    """
    Rule 3: If email is from a .edu or .gov address → probably NOT spam!
    Schools and government emails are usually legitimate.
    """
    if ".edu" in email.sender or ".gov" in email.sender:
        return NOT_SPAM
    return ABSTAIN

@labeling_function()
def lf_short_email(email):
    """
    Rule 4: Very short emails (< 20 words) with no formal structure → might be spam!
    """
    word_count = len(email.text.split())
    return SPAM if word_count < 20 else ABSTAIN

# ── Apply all rules to our unlabelled dataset ─────────────────────
lfs = [lf_contains_free, lf_exclamation_overload,
       lf_professional_domain, lf_short_email]

# LFApplier runs ALL rules on ALL data at once — super fast!
applier = LFApplier(lfs=lfs)
L_train = applier.apply(df=train_df)   # L_train = matrix of votes

# ── Analyze how well our rules agree ──────────────────────────────
# This tells us: How often does each rule vote? Do they agree?
lf_analysis = LFAnalysis(L=L_train, lfs=lfs).lf_summary()
print("Labelling Function Analysis:")
print(lf_analysis)

# ── Train the Label Model to combine all rule votes ────────────────
# Like a wise judge who listens to all 4 rules and gives final verdict!
label_model = LabelModel(cardinality=2, verbose=True)
label_model.fit(L_train=L_train, n_epochs=500, seed=42)

# Get final probabilistic labels for all data!
# Returns a probability: 0.73 = 73% chance of being spam
prob_labels = label_model.predict_proba(L=L_train)
print(f"\nGenerated {len(prob_labels)} labels automatically! 🎉")
print(f"Sample: Email 1 is {prob_labels[0][1]*100:.0f}% likely to be SPAM")

📋 Key lines explained:

  • @labeling_function() → A special decorator that tells Snorkel "this function is a labelling rule"
  • ABSTAIN = -1 → When a rule doesn't apply, it says "I don't know" — this is important! Forced guesses would hurt accuracy
  • LabelModel → The smart judge that weighs each rule's vote based on how reliable it has been historically

Method 4️⃣ — LLM-Assisted Labelling (AI Labels AI! )

One of the biggest trends! Instead of hiring humans to label data, we use a large language model (like GPT-4 or Claude) to generate the first draft of labels. Then a human just checks and corrects them — much faster!

Think of it like having a very fast student who does all the homework, and then the teacher just goes through quickly marking which ones are right and fixing the wrong ones. 📚

  LLM-Assisted Labelling Flow:
  ─────────────────────────────────────────────────────────────────
  1,000 customer reviews to label:

  OLD WAY (Manual only):
  Human reads each review → writes label → 50 reviews/hour
  Time needed: 1000 ÷ 50 = 20 hours 

  NEW WAY (LLM-Assisted):
  LLM reads each review → generates first-pass label → instant!
  Human checks LLM output → fixes mistakes → 200 reviews/hour ✅
  Time needed: 1000 ÷ 200 = 5 hours  (75% faster!)
  ─────────────────────────────────────────────────────────────────

📋 What the code below does:
We send each piece of text to an LLM with a carefully written instruction (prompt). The LLM reads the text and gives back its best guess for the label. This is called "zero-shot labelling" — the LLM was never trained specifically on your task, but it's smart enough to make a good first guess using its general knowledge!

# ─────────────────────────────────────────────────────────────────
# LLM-Assisted Labelling — Use AI to pre-label your data!
# ─────────────────────────────────────────────────────────────────
# We use an LLM to generate FIRST DRAFT labels for our data.
# This is WAY faster than a human reading each item one by one.
# A human expert then just REVIEWS and CORRECTS — not labels from scratch!
# ─────────────────────────────────────────────────────────────────

import json

# Install: pip install openai
# from openai import OpenAI
# client = OpenAI(api_key="your-key")

# (Using a mock function here for demonstration — replace with real API call)
def llm_label(text: str, categories: list) -> dict:
    """
    Ask an LLM to label a piece of text.

    This function sends the text + labelling instructions to an LLM.
    The LLM returns: the label + a confidence score + its reasoning.

    Why include reasoning? So human reviewers can quickly check
    IF the LLM understood the text correctly!
    """
    prompt = f"""You are a data labelling expert. 
    
Analyse this customer review and classify it into ONE of these categories:
{', '.join(categories)}

Customer review: "{text}"

Respond in JSON format only:
{{
  "label": "chosen category",
  "confidence": 0.0-1.0,
  "reasoning": "one sentence explanation"
}}"""

    # In production: send this prompt to GPT-4 or Claude API
    # response = client.chat.completions.create(
    #     model="gpt-4o",
    #     messages=[{"role": "user", "content": prompt}]
    # )
    # return json.loads(response.choices[0].message.content)

    # Mock response for demo:
    return {
        "label":      "POSITIVE",
        "confidence": 0.92,
        "reasoning":  "Review uses enthusiastic language and recommends the product"
    }


# ── Apply LLM labelling to a batch of reviews ─────────────────────
reviews = [
    "This product completely changed my life! I recommend it to everyone.",
    "Absolute garbage. Broke after 2 days. Waste of money.",
    "It is okay, nothing special. Works as described.",
    "AMAZING quality! Fast shipping! Will buy again!!! ⭐⭐⭐⭐⭐",
    "Disappointed. Expected much better for the price."
]

categories = ["POSITIVE", "NEGATIVE", "NEUTRAL"]
results    = []

for i, review in enumerate(reviews):
    result = llm_label(review, categories)
    results.append({
        "review":     review,
        "llm_label":  result["label"],
        "confidence": result["confidence"],
        "reasoning":  result["reasoning"],
        # Flag low-confidence items for human review!
        "needs_human_review": result["confidence"] < 0.80
    })

# Print results
for r in results:
    flag = "⚠️ Human Review Needed" if r["needs_human_review"] else "✅ Auto-Approved"
    print(f"{flag} | Label: {r['llm_label']} ({r['confidence']:.0%}) | {r['review'][:40]}...")

📋 Key lines explained:

  • Respond in JSON format only → We tell the LLM to respond in structured JSON so we can parse it programmatically — no messy free text!
  • "confidence": 0.0-1.0 → The LLM estimates how sure it is. Low confidence = human should double-check!
  • needs_human_review = confidence < 0.80 → Any label the LLM is less than 80% sure about → send to a human. Smart filtering!

Method 5️⃣ — RLHF (Teaching AI What Humans Prefer! ❤️)

This is the method that made ChatGPT, Claude, and Gemini so good at talking! RLHF stands for Reinforcement Learning from Human Feedback.

Instead of labelling data with categories (cat/dog/spam), humans compare two AI-generated answers and say which one is BETTER. The AI learns to generate responses that humans prefer!

  How RLHF Labelling Works:
  ─────────────────────────────────────────────────────────────────
  QUESTION: "How do I lose weight healthily?"

  AI Answer A: "Stop eating completely for a week."  ← Dangerous!

  AI Answer B: "Eat balanced meals, exercise regularly,
                sleep well, and consult a doctor."   ← Good advice!

  Human Labeller says: "B is better! ✅"

  The AI learns: "Answer B is preferred by humans → generate more like B"
  Over millions of such comparisons → AI learns human values!
  ─────────────────────────────────────────────────────────────────
  This is how ChatGPT learned to be helpful, harmless, and honest!
✅ RLHF Updates:

        RLHF has evolved into newer methods:
  • DPO (Direct Preference Optimization) → Simpler than RLHF, no separate reward model needed
  • GRPO → Even more efficient, used by DeepSeek-R1
  • RLAIF → AI gives feedback to AI (instead of humans) — scales better but needs careful oversight

Method 6️⃣ — Synthetic Data Generation (Creating Fake But Real Data! 🎭)

What if you don't have enough real data to label? Create synthetic data! Use AI to generate artificial but realistic data — complete with labels already attached.

  Real-World Examples of Synthetic Data:
  ──────────────────────────────────────────────────────────────────
  🚗 Self-driving cars → Simulate thousands of car crash scenarios
     in a game engine (like Unity) — no real crashes needed!

  🏥 Medical AI → Generate fake patient X-rays with known conditions
     — no privacy issues, no waiting for rare cases!

  🤖 Chatbots → Use an LLM to generate thousands of fake customer
     conversations with different intents already labelled!
  ──────────────────────────────────────────────────────────────────
  Market for synthetic data is expected to hit $3.7 BILLION by 2030!
⚠️ Important Warning:
Synthetic data is great for volume — but always validate on REAL data! A model trained only on synthetic data may fail on real-world edge cases. Use synthetic data to supplement real labelled data, not replace it!

🔹 STEP 4 — Quality Check (Is the Label Correct? ✅)

Garbage in = Garbage out. 🗑️ If your labels are wrong, your model learns the wrong things. Quality checking is NOT optional — it is critical!

Key Quality Metrics

Inter-Annotator Agreement (IAA)

If you have multiple people labelling the same data, do they AGREE with each other? Like having two doctors read the same X-ray — if they both say "normal", you're confident! If one says "tumour" and the other says "normal" — you have a problem! 🚨

  Inter-Annotator Agreement Example:
  ─────────────────────────────────────────────────────────────────
  Same 5 emails labelled by Labeller A and Labeller B:

  Email    Labeller A    Labeller B    Same?
  ──────   ──────────    ──────────    ─────
  Email 1  SPAM          SPAM          ✅ Agree
  Email 2  NOT SPAM      NOT SPAM      ✅ Agree
  Email 3  SPAM          NOT SPAM      ❌ Disagree!
  Email 4  NOT SPAM      NOT SPAM      ✅ Agree
  Email 5  SPAM          SPAM          ✅ Agree

  Agreement = 4/5 = 80%
  Cohen's Kappa score = 0.75 (considered "Substantial Agreement" ✅)
  ─────────────────────────────────────────────────────────────────
  Industry standard: Target IAA above 80% (Cohen's Kappa > 0.7)

📋 What the code below does:
We measure how much two labellers agree using Cohen's Kappa — a statistical formula that gives a score between -1 and 1. A score of 1 means perfect agreement. A score of 0 means they're just randomly guessing! A score above 0.7 is generally considered good enough for production-grade labelling.

# ─────────────────────────────────────────────────────────────────
# Quality Check — Measuring Inter-Annotator Agreement
# ─────────────────────────────────────────────────────────────────
# We measure HOW MUCH two human labellers agree with each other.
# If they agree a lot → labels are reliable!
# If they disagree a lot → guidelines are unclear or task is ambiguous!
# ─────────────────────────────────────────────────────────────────

from sklearn.metrics import cohen_kappa_score
import numpy as np

# Labels given by Labeller A (experienced domain expert)
labeller_A = np.array([1, 0, 1, 0, 1, 1, 0, 1, 0, 0])
#                      S  N  S  N  S  S  N  S  N  N
#                      (S=Spam, N=Not Spam)

# Labels given by Labeller B (another expert)
labeller_B = np.array([1, 0, 0, 0, 1, 1, 0, 1, 1, 0])
#                      S  N  N  N  S  S  N  S  S  N

# ── Calculate raw agreement ───────────────────────────────────────
# Simple percentage: how many labels are exactly the same?
raw_agreement = np.mean(labeller_A == labeller_B)
print(f"Raw Agreement: {raw_agreement * 100:.0f}%")

# ── Calculate Cohen's Kappa ───────────────────────────────────────
# More reliable than raw agreement because it adjusts for random chance!
# Even if labellers guessed randomly, they'd agree 50% of time by luck!
# Kappa removes that "luck agreement" and shows TRUE agreement.
kappa = cohen_kappa_score(labeller_A, labeller_B)
print(f"Cohen's Kappa: {kappa:.3f}")

# ── Interpret the Kappa score ─────────────────────────────────────
if kappa >= 0.8:
    print("✅ Almost Perfect Agreement — Labels are very reliable!")
elif kappa >= 0.6:
    print("✅ Substantial Agreement — Good enough for training!")
elif kappa >= 0.4:
    print("⚠️  Moderate Agreement — Review labelling guidelines!")
else:
    print("❌ Poor Agreement — Guidelines unclear, retrain labellers!")

# ── Find the disagreements ────────────────────────────────────────
# Show exactly WHICH samples the labellers disagreed on
disagreements = np.where(labeller_A != labeller_B)[0]
print(f"\nDisagreements on samples: {disagreements.tolist()}")
print(f"These {len(disagreements)} samples need a 3rd expert to decide!")

📋 Key lines explained:

  • cohen_kappa_score() → A statistical score that measures "true agreement" beyond what random chance would give
  • np.where(labeller_A != labeller_B) → Finds exactly which samples the two labellers disagreed on — send these to a third expert!
  • kappa >= 0.7 → Industry standard. Below this, your guidelines need improvement

🔹 STEP 5 — Store and Version Labels (Never Lose Your Work! 💾)

Imagine spending 3 months labelling 50,000 images and then accidentally overwriting the labels file. Nightmare! 😱 This is why we version control our labels — just like we version control code with Git.

  Label Versioning Example:
  ──────────────────────────────────────────────────────────────────
  labels_v1.0  ← First version: 10,000 images, basic categories
  labels_v1.1  ← Added 5,000 more images
  labels_v2.0  ← Redesigned categories (major change)
  labels_v2.1  ← Fixed 200 wrongly labelled images
  labels_v3.0  ← Added bounding box annotations
  ──────────────────────────────────────────────────────────────────
  Every model training run must record WHICH version of labels was used!
  This way, if a model behaves badly, you can trace it back to the labels.

📋 What the code below does:
We save labels as a JSON file and then use DVC (Data Version Control) to track changes. DVC is like Git — but for large data files and labels. It remembers every version of your label files so you can always go back in time!

# ─────────────────────────────────────────────────────────────────
# Label Storage and Versioning — Never Lose Your Labels!
# ─────────────────────────────────────────────────────────────────
# We save labels in a structured JSON format AND use DVC to
# version control them — just like Git controls your code!
# Install DVC: pip install dvc
# ─────────────────────────────────────────────────────────────────

import json
import hashlib
from datetime import datetime
from pathlib import Path

def save_labels_with_metadata(labels: dict, version: str, description: str):
    """
    Save a labelled dataset to disk with complete metadata.

    We store MORE than just the labels — we also store:
    - WHO labelled them (labeller IDs)
    - WHEN they were labelled
    - Version number (so we can track changes)
    - A checksum (to detect if the file is corrupted!)

    Think of it like a library book with a stamp showing
    who borrowed it, when, and its edition number! 📚
    """
    label_data = {
        # ── Metadata ────────────────────────────────────────────
        "version":     version,           # e.g. "v2.1"
        "created_at":  datetime.now().isoformat(),
        "description": description,
        "label_count": len(labels),

        # ── The actual labels ────────────────────────────────────
        "labels": labels,

        # ── Schema definition ────────────────────────────────────
        "schema": {
            "classes": ["POSITIVE", "NEGATIVE", "NEUTRAL"],
            "task":    "sentiment_classification",
        }
    }

    # Save to file
    filename = f"labels_{version}.json"
    filepath = Path("./label_store") / filename
    filepath.parent.mkdir(exist_ok=True)

    with open(filepath, "w") as f:
        json.dump(label_data, f, indent=2)

    # Compute checksum (a fingerprint of the file)
    # If the file gets corrupted or tampered, the checksum will change!
    file_content = json.dumps(label_data, sort_keys=True).encode()
    checksum = hashlib.md5(file_content).hexdigest()

    print(f"✅ Labels saved to: {filepath}")
    print(f"   Version: {version}")
    print(f"   Labels: {len(labels)} items")
    print(f"   Checksum: {checksum}")
    print(f"\n   Now run: dvc add {filepath}")
    print(f"            git commit -m 'Add labels {version}'")

    return str(filepath), checksum


# Example usage:
sample_labels = {
    "review_001": {"label": "POSITIVE", "labeller": "human_A", "confidence": 0.95},
    "review_002": {"label": "NEGATIVE", "labeller": "human_A", "confidence": 0.88},
    "review_003": {"label": "NEUTRAL",  "labeller": "llm_gpt4",  "confidence": 0.73},
}

filepath, checksum = save_labels_with_metadata(
    labels=sample_labels,
    version="v2.1",
    description="Added LLM-assisted labels for batch 3. Human reviewed all < 0.80 confidence."
)

🔹 STEP 6 — Common Types of Labelling Tasks

Different data types need different types of labels. Here's a complete overview:

Image Labelling Types

  1. IMAGE CLASSIFICATION
     ────────────────────
     Whole image → one label
     📷 Photo → "CAT" or "DOG" or "BIRD"

  2. OBJECT DETECTION (Bounding Boxes)
     ────────────────────────────────
     Draw boxes around objects in the image
     📷 Street photo → Box around each car, person, traffic light
     Each box has: [x, y, width, height, label]

  3. IMAGE SEGMENTATION
     ────────────────────────────────────────────
     Colour EVERY SINGLE PIXEL with its category!
     Much harder than bounding boxes!
     Used in: Self-driving cars, medical imaging, maps

  4. KEYPOINT LABELLING
     ─────────────────────────────────────────────
     Mark specific points on an object
     Example: 17 keypoints on a human body (nose, eyes, shoulders, etc.)
     Used for: Pose estimation, face recognition, sports analysis

Text Labelling Types

  1. TEXT CLASSIFICATION    → Whole text → one category (spam/not spam)

  2. NAMED ENTITY RECOGNITION (NER)
     "John Smith from London visited Apple headquarters."
      [PERSON]          [CITY]           [COMPANY]
     Tag each word with its entity type!

  3. SENTIMENT ANALYSIS    → Positive / Negative / Neutral

  4. INTENT LABELLING       → "Book me a flight" → INTENT: book_flight
     (Used for chatbots and voice assistants)

  5. QUESTION-ANSWER PAIRS  → Used for training LLMs!
     Q: "What is the capital of France?"  A: "Paris"

  6. PREFERENCE RANKING (RLHF) → Which answer is better? A or B?

🛠️ Popular Data Labelling Tools

You don't have to build a labelling system from scratch! Here are the best tools:

  ┌────────────────────┬─────────────────┬────────────────────────────────────┐
  │ Tool               │ Type            │ Best For                           │
  ├────────────────────┼─────────────────┼────────────────────────────────────┤
  │ Label Studio       │ Open Source 🆓  │ ALL data types, RLHF support       │
  │ Labelbox           │ Commercial 💰   │ Enterprise, active learning         │
  │ Encord             │ Commercial 💰   │ Video, 3D, multimodal, LLM          │
  │ CVAT               │ Open Source 🆓  │ Image and video annotation          │
  │ Snorkel            │ Open Source 🆓  │ Weak supervision, programmatic      │
  │ Scale AI           │ Managed 💰💰    │ Large-scale, autonomous vehicles    │
  │ Roboflow           │ Freemium        │ Computer vision, easy to use        │
  │ Prodigy            │ Commercial 💰   │ NLP, active learning, code-first    │
  └────────────────────┴─────────────────┴────────────────────────────────────┘
✅ Beginner Recommendation:
Start with Label Studio — it is completely FREE, runs on your own computer, supports images, text, audio, video, and even RLHF workflows. It's the best place to learn real data labelling without spending money!

Setting Up Label Studio Locally

📋 What these commands do:
These 3 commands install Label Studio on your computer, start the web server, and open a browser where you can upload data, draw bounding boxes, write labels, and export them. Like setting up your own mini labelling studio at home in 2 minutes!

# ─────────────────────────────────────────────────────────────────
# Install and Launch Label Studio — Your Free Labelling Studio!
# ─────────────────────────────────────────────────────────────────
# These 3 commands install Label Studio and open it in your browser.
# Once running, you can upload photos, draw bounding boxes,
# write text labels, and export everything to JSON/CSV — for FREE!
# ─────────────────────────────────────────────────────────────────

# Step 1: Install Label Studio (takes 1-2 minutes)
pip install label-studio

# Step 2: Start Label Studio server
# This launches a local website at http://localhost:8080
label-studio start

# Step 3: Open your browser and go to:
# http://localhost:8080
# Create an account, start a new project, and begin labelling! 🎉

# ── To export your labels (after labelling) ───────────────────────
# In the Label Studio UI: click Export → choose JSON or CSV
# OR use the command line:
label-studio export --project-id 1 --format JSON --output ./my_labels.json

📊 Data Labelling Landscape — Big Picture

  DATA LABELLING APPROACHES BY VOLUME AND ACCURACY:
  ──────────────────────────────────────────────────────────────────
  APPROACH              VOLUME    ACCURACY   COST     TREND 
  ──────────────────    ──────    ────────   ──────   ─────────────
  Manual Human          LOW       HIGHEST    $$$$     Still King for
                                                      complex tasks ✅

  Active Learning       MEDIUM    HIGH       $$$      Growing fast 📈
                                                      30-40% savings!

  Weak Supervision      HIGH      MEDIUM     $$       Widely adopted 📈

  LLM-Assisted          HIGH      HIGH       $$       Fastest growing🚀
  (AI + Human review)

  Synthetic Data        VERY HIGH MEDIUM     $        Booming market 💥
                                                      ($3.7B by 2030!)

  RLHF / Preference     LOW       HIGHEST    $$$$     Powers all LLMs ⭐
  Ranking
  ──────────────────────────────────────────────────────────────────
  Most production teams use HYBRID approaches:
  LLM labels first → Human reviews uncertain ones → Active learning
  selects the next batch → Repeat!

⚠️ Common Mistakes in Data Labelling (Don't Do These! 🚫)

❌ Mistake 1: No Labelling Guidelines
Starting labelling without a clear rulebook. Different labellers interpret the same data differently → inconsistent labels → bad model.

❌ Mistake 2: Ignoring Class Imbalance
If 98% of your labels are "Not Fraud" and only 2% are "Fraud" — the model will just always predict "Not Fraud" and be 98% "accurate" while missing all actual fraud!

❌ Mistake 3: Not Measuring IAA
Never checking if your labellers agree with each other. If they don't agree, your labels are noise — not signal.

❌ Mistake 4: No Versioning
Overwriting labels without keeping old versions. If the model gets worse after a labelling update, you can't go back.

❌ Mistake 5: Labelling Bias
Only labelling easy examples. Hard and edge-case examples are the most important ones for your model to learn from!
✅ Best Practices Summary:
  • Write clear labelling guidelines BEFORE you start
  • Always measure Inter-Annotator Agreement — target Kappa > 0.7
  • Use Active Learning to label smart, not label everything
  • Version control your labels with DVC
  • Do a small pilot (100-200 samples) before the full labelling run
  • Use LLM-assisted labelling for first pass — human for final review
  • Include edge cases and rare examples on purpose!

🎯 Quick Summary — What We Learned Today!

Data Labelling = Adding meaningful tags to raw data so ML models can learn

  • Manual Labelling → Human experts label each item. Slow but highest quality for complex tasks
  • Active Learning → Model chooses which uncertain samples to label first. 30-40% savings!
  • Weak Supervision → Write rules (Labelling Functions), combine them. Labels millions instantly!
  • LLM-Assisted → AI does first draft, human corrects. 75% faster than manual!
  • RLHF → Humans pick which AI answer is better. How ChatGPT was taught to be helpful!
  • Synthetic Data → Create fake-but-realistic labelled data when real data is scarce
  • Quality = IAA → Measure inter-annotator agreement. Below 70%? Fix your guidelines!
  • Version Everything → Use DVC to version labels just like Git versions code

🌟 You did it! You now understand Data Labelling from zero to hero level — one of the most critical (and most underrated) skills in real-world Machine Learning. Every AI system that exists today — from your phone's face recognition to ChatGPT — runs on labelled data created by people like you. Keep learning! 💪

Happy labelling! 🏷️✨

Comments