Skip to main content

Labeling Functions (LFs) in Machine Learning: A Practical Guide to Weak Supervision

Calculating read time…

Imagine you are the teacher of a class of 1,000,000 students. You need to give every single student a grade: Pass or Fail. Reading every student's paper by hand would take your entire lifetime!

Now imagine you could write a few clever rules instead — "If the paper mentions the formula correctly, probably Pass. If it's under 50 words, probably Fail." You write 10 such rules, and they automatically grade all 1,000,000 papers for you in minutes!

That's exactly what Labeling Functions do in Machine Learning. They are small, smart rules that automatically label enormous amounts of data — saving months of human effort, and costing almost nothing.



One-line summary: A Labeling Function is a mini-rule you write that says: "If you see THIS pattern, give THAT label."


📋 What You Will Learn 

  • The biggest problem in Machine Learning — the labeling bottleneck
  • What Weak Supervision is and why it's a game-changer
  • What a Labeling Function is (with simple analogies)
  • The 7 types of Labeling Functions with real examples
  • What ABSTAIN means and why it's a superpower
  • How multiple LFs vote together (the Label Model)
  • The 4 key metrics to measure LF quality
  • Full hands-on code: writing LFs with Snorkel (the most popular LF library)
  • The trend: LLM-powered Labeling Functions
  • The full pipeline: LFs → Label Model → Train your ML model
  • Best practices, DOs and DON'Ts

1. 😩 The Biggest Problem in ML — "We Need Labels!"

Machine Learning models learn by example. You show them thousands of examples with correct answers, and they learn the pattern. But here's the painful truth:

Getting those "correct answers" (called labels) is incredibly hard, slow, and expensive.

🧒 Think of it like this:

Imagine you want to teach a robot to recognise spam emails. You need to show it 100,000 emails with each one marked: SPAM or NOT SPAM. A human expert has to read and mark every single one — that's months of work! If your definition of "spam" changes tomorrow, you start all over again. 

This problem — called the labeling bottleneck — is the #1 reason most ML projects fail or take forever.

 THE LABELING BOTTLENECK

 UNLABELED DATA         MANUAL LABELING         LABELED DATA
 ─────────────          ───────────────          ────────────
 Email 1: ???    →→  Human reads it  →→    Email 1: SPAM
 Email 2: ???    →→  Human reads it  →→    Email 2: NOT SPAM
 Email 3: ???    →→  Human reads it  →→    Email 3: SPAM
 ...
 Email 100,000:  →→  [3 months later] →→   Email 100,000: ???

 ⚠️ Cost: $$$$$   Time: Months   Scalability: 😭

Researchers at Stanford found that labeling training data has become the single largest bottleneck in deploying production ML systems. Companies like Google have hired entire teams of hundreds of people just to label data!

2. 💡 Weak Supervision — The Smart Shortcut

Weak Supervision is a completely different approach. Instead of labeling data one by one, you write rules, heuristics, and shortcuts that automatically assign labels to thousands of data points at once.

These rules won't be perfect. Some will make mistakes. But here's the clever part: you use many imperfect rules together, and a mathematical model figures out the best combined answer — just like a jury reaching a verdict!

🧒 The jury analogy:

Imagine 10 jurors in a courtroom, each with partial evidence. Juror 1 is very reliable. Juror 5 is sometimes wrong. The judge (Label Model) doesn't take a simple vote — it weighs each juror's reliability automatically and reaches the best possible verdict.

 WEAK SUPERVISION vs MANUAL LABELING

                  MANUAL LABELING         WEAK SUPERVISION
                  ───────────────         ────────────────
 Speed:           Months                  Hours / Days
 Cost:            Very High ($$$$$)       Very Low ($)
 Scalability:     Limited                 Millions of examples
 Flexibility:     Rigid                   Update rules in minutes
 Accuracy:        High (if careful)       Good (improves with more LFs)
 Used at:         Small startups          Google, Intel, Stanford,
                                          BNY Mellon, Chubb Insurance...

Research from Stanford's AI Lab showed that teams using Weak Supervision built models 2.8x faster and increased predictive performance by an average of 45.5% compared to traditional 7-hour manual labeling sessions!

3. 🏷️ What Exactly is a Labeling Function?

A Labeling Function (LF) is a Python function you write that:

  • Takes ONE data point as input (e.g., an email text)
  • Looks at it using some rule or heuristic
  • Returns ONE of three things: a positive label, a negative label, or ABSTAIN (I don't know)

That's it. It's really that simple! The LF doesn't need to be right 100% of the time. It just needs to be right more often than random guessing.

The Three Possible Outputs of a Labeling Function

 LABELING FUNCTION OUTPUTS
 ──────────────────────────

 Input: An email text

 LF Output 1 → SPAM (1)       "I think this IS spam"
 LF Output 2 → NOT SPAM (0)   "I think this is NOT spam"
 LF Output 3 → ABSTAIN (-1)   "I have no idea, skip me"

 ABSTAIN is the LF saying: "This data point is outside my area of expertise.
 Let other LFs handle it. Don't count my vote here."
✅ ABSTAIN is a SUPERPOWER, not a weakness!
A LF that abstains when it's uncertain is much better than one that guesses randomly. Abstaining prevents a bad LF from polluting good labels on examples it knows nothing about. Think of a doctor saying "I'm not sure, let a specialist check this" — that's wisdom, not failure!

A Simple Mental Model

🧒 Think of LFs as detectives, each specialising in one clue:

  • Detective 1 checks: "Does the email mention FREE MONEY?" → If yes: SPAM. If not: ABSTAIN.
  • Detective 2 checks: "Is it from a known colleague's email address?" → If yes: NOT SPAM. If not: ABSTAIN.
  • Detective 3 checks: "Does it have 50+ exclamation marks?" → If yes: SPAM. If not: ABSTAIN.
  • Detective 4 checks: "Does it contain a meeting invitation?" → If yes: NOT SPAM. If not: ABSTAIN.

No single detective catches every spam. But together, when they all vote, the combined verdict is surprisingly accurate!

4. 🔧 The 7 Types of Labeling Functions (with Examples)

Labeling Functions come in many flavours. You choose the type based on what kind of "signal" you have available. Let's explore all of them with the spam email detection example throughout:

Type 1 — Keyword-Based LF

The simplest type. You give the LF a list of words or phrases. If the data point contains any of them, return a label.

🧒 Analogy: A security guard with a list of forbidden words. If you say one of them at the door, you're flagged!

 KEYWORD LF — MENTAL MODEL

 Spam keywords list: ["FREE MONEY", "CLICK HERE", "WINNER", "PRIZE", "$$$$"]

 Email A: "You WON a PRIZE! CLICK HERE to claim."
    → Contains "PRIZE" + "CLICK HERE" → Label: SPAM ✅

 Email B: "Let's schedule a meeting for Tuesday."
    → No spam keywords → Label: ABSTAIN 🔇

Type 2 — Regex (Pattern) LF

Uses regular expressions — pattern-matching rules for text. Perfect for structured patterns like phone numbers, URLs, email formats, dollar amounts.

🧒 Analogy: Like using a template. "Does this sentence fit the shape of a phone number pattern?"

 REGEX LF — MENTAL MODEL

 Pattern: "Emails with suspicious URLs (e.g., bit.ly/xxx)"

 Email A: "Claim your prize here: bit.ly/abc123"
    → Matches URL shortener pattern → Label: SPAM ✅

 Email B: "Project report attached. See slide 3."
    → No URL pattern → Label: ABSTAIN 🔇

Type 3 — Heuristic / Rule-Based LF

Uses domain expertise in the form of if-then logic. A human expert writes: "If X AND Y, then label Z." These are the most flexible type!

🧒 Analogy: Your grandma's wisdom. "If someone calls at midnight AND asks for money, it's suspicious."

 HEURISTIC LF — MENTAL MODEL

 Rule: "If email is ALL CAPS AND has more than 3 exclamation marks → SPAM"

 Email A: "YOU WON!!! CLAIM NOW!!! LIMITED TIME!!!"
    → ALL CAPS: Yes. Exclamations: 6 → Label: SPAM ✅

 Email B: "Quick reminder about tomorrow's meeting!"
    → Not ALL CAPS. Exclamations: 1 → Label: ABSTAIN 🔇

Type 4 — Distant Supervision LF

Uses an external knowledge base (a list, database, or lookup table) to generate labels. If a data point matches an entry in the knowledge base, apply the known label.

🧒 Analogy: A bouncer with a VIP guest list. If your name is on the list, you're automatically in (or out).

 DISTANT SUPERVISION LF — MENTAL MODEL

 Known spam sender list: ["spammer@phishing.com", "noreply@scam.ru"]

 Email A: From "spammer@phishing.com"
    → Matches spam sender list → Label: SPAM ✅

 Email B: From "alice@mycompany.com"
    → Not in any list → Label: ABSTAIN 🔇

Type 5 — Model-Based LF

Uses a pre-trained ML model as a labeling function! Maybe you have an old, slightly inaccurate model from a previous project. Instead of throwing it away, you use it as one voter among many.

🧒 Analogy: An experienced but slightly rusty expert. Their judgement isn't perfect, but it's still useful when combined with others.

 MODEL-BASED LF — MENTAL MODEL

 Old spam classifier (trained 2 years ago, 78% accuracy):
    → Predicts probability of spam for each email
    → If probability > 0.85 → Label: SPAM
    → If probability < 0.15 → Label: NOT SPAM
    → Otherwise → ABSTAIN (it's not confident enough)

Type 6 — Crowdsourced / Human LF

Uses input from non-expert humans (e.g., via surveys, clicks, ratings, or feedback data) as a source of weak labels. Each individual human label might be noisy, but aggregated, they're powerful!

🧒 Analogy: Asking 1,000 people "Is this funny?" — even if some people have bad taste in humour, the majority vote is usually right!

Type 7 — LLM-Powered LF (The Trend! 🔥)

The newest and most exciting type. Instead of writing rules in code, you write a natural language prompt and ask a Large Language Model (like GPT or Llama) to label the data. The LLM's answer becomes the LF's vote.

🧒 Analogy: Instead of training your robot guard, you just tell it in plain English: "If the email sounds suspicious or pushy, it's spam."

 LLM-POWERED LF — MENTAL MODEL

 Prompt to LLM:
 "The following email was received. Is it spam?
  Email: {email_text}
  Answer with ONLY: 'SPAM' or 'NOT SPAM' or 'UNSURE'"

 Email A: "CONGRATULATIONS! You've been selected for a free gift!"
    → LLM: "SPAM" → Label: SPAM ✅

 Email B: "Please review the attached quarterly report."
    → LLM: "NOT SPAM" → Label: NOT SPAM ✅

 Email C: "Hi, following up on our last chat."
    → LLM: "UNSURE" → Label: ABSTAIN 🔇
⚠️ Trend Alert: LLM-powered LFs are now being used by research teams at Stanford, Google, and enterprise AI companies. A 2024 paper in the ACM/IMS Journal of Data Science showed that treating LLMs as labeling functions within a weak supervision pipeline outperformed zero-shot and few-shot LLM prompting alone — because the label model learns how much to trust each LLM prompt variant!

5. 📊 The Label Matrix — How All LFs Work Together

When you have multiple LFs, you apply all of them to all your data points. The result is called the Label Matrix — a big table of votes.

 THE LABEL MATRIX

 Each row = one data point (email)
 Each column = one Labeling Function
 Values: 1=SPAM, 0=NOT SPAM, -1=ABSTAIN

            LF1    LF2    LF3    LF4    LF5
 Email 1:    1      1     -1      1     -1    → Probably SPAM (3 votes for SPAM)
 Email 2:    0     -1      0     -1      0    → Probably NOT SPAM (3 votes against)
 Email 3:   -1     -1     -1     -1     -1    → No coverage! (all abstain)
 Email 4:    1     -1      0      1      0    → Conflicted (2 SPAM, 2 NOT SPAM)
 Email 5:    1      1      1      0     -1    → Probably SPAM (3 SPAM, 1 against)

 ↓
 Label Model reads this matrix and computes the BEST GUESS label for each row.

Notice Email 4 has a conflict — some LFs say SPAM, some say NOT SPAM. The Label Model resolves this by figuring out which LFs are more trustworthy and weighting their votes accordingly. It's like the judge deciding which witnesses in a trial are more credible!

6. 📏 The 4 Key Metrics to Judge Your LFs

After applying your LFs to your data, Snorkel gives you 4 important statistics about each LF. Think of these as a report card for each LF.

Metric 1 — Coverage

What fraction of your data does this LF label? (i.e., how often does it NOT abstain?)

🧒 Like asking: "How many students did this teacher actually grade?" A teacher who only graded 3 out of 1000 papers isn't very useful!

 Coverage = (Number of non-ABSTAIN labels) / (Total data points)

 LF1 labels 800 out of 1000 emails → Coverage = 0.80 (80%)  ✅ Good!
 LF2 labels  12 out of 1000 emails → Coverage = 0.01 (1%)   ⚠️ Too low!

Metric 2 — Polarity

What unique labels does this LF output? (e.g., does it only say SPAM, or does it say both SPAM and NOT SPAM?)

🧒 Like asking: "Does this judge only ever convict people, or does she also acquit?" A judge who always says "Guilty!" regardless of evidence is suspicious!

Metric 3 — Overlaps

How often does this LF label the same data point as at least one other LF? High overlap = this LF's territory is well-covered by other LFs too.

Metric 4 — Conflicts

How often does this LF disagree with another LF on the same data point? High conflicts = the LF contradicts others a lot. Some conflict is normal and healthy — it means your LFs are covering different signals. But if one LF conflicts with ALL others, it might be wrong!

 LF REPORT CARD (from Snorkel)

          Coverage  Polarity  Overlaps  Conflicts
 LF1:     0.82      [0,1]     0.65      0.08     ← Healthy LF ✅
 LF2:     0.14      [1]       0.09      0.03     ← Low coverage ⚠️
 LF3:     0.91      [0,1]     0.78      0.22     ← High conflict, check it ⚠️
 LF4:     0.55      [0]       0.40      0.00     ← Only says NOT SPAM, ok ✅
 LF5:     0.03      [1]       0.02      0.01     ← Almost useless ❌

7. 💻 Hands-On Code — Building Labeling Functions with Snorkel

Snorkel is the most popular and powerful Python library for Labeling Functions. It was built at Stanford AI Lab, published in a landmark research paper, and is now used by companies like Google, Intel, BNY Mellon, and Stanford Medicine.

🧒 Think of Snorkel as a smart referee. You bring your team of imperfect judges (LFs). Snorkel watches how they vote, figures out who's most trustworthy, and gives you the final score!

First — Install Snorkel

📝 What the command below does:
This installs the Snorkel library onto your computer using pip (Python's package manager). Snorkel contains all the tools we need: decorators for writing LFs, the LF applier (runs LFs on data), and the Label Model (combines LF votes).
pip install snorkel pandas scikit-learn

Step 1 — Set Up Our Data

📝 What the code below does:
We create a tiny dataset of 8 emails. Each email is stored in a Pandas DataFrame (think of it as a spreadsheet in Python). We also define our label constants: SPAM=1, NOT_SPAM=0, ABSTAIN=-1. In real projects, this dataset would have thousands or millions of rows — all without labels! That's the whole point — we'll use LFs to label them automatically.
# Step 1: Import libraries and create our email dataset
# Pandas lets us work with data in a table format (like Excel in Python)
# Snorkel's ABSTAIN constant (-1) means "I don't know" — our LF skips this example

import pandas as pd
from snorkel.labeling import labeling_function

# ── Define what our labels mean ─────────────────────────────────────────────
ABSTAIN   = -1   # LF says: "I have no opinion on this one"
NOT_SPAM  =  0   # LF says: "This is definitely NOT spam"
SPAM      =  1   # LF says: "This IS spam"

# ── Create a small example email dataset ────────────────────────────────────
# In a real project you'd load millions of unlabeled emails from a file or database.
# For learning, we use 8 simple examples with known answers (we'll pretend we don't know them!).

data = {
    "email_id": [1, 2, 3, 4, 5, 6, 7, 8],
    "text": [
        "Congratulations! You've WON a FREE prize! CLICK HERE NOW!!!",   # SPAM
        "Hi, please review the attached quarterly report by Friday.",       # NOT SPAM
        "URGENT: Your account will be closed! Verify NOW for FREE access", # SPAM
        "Can we reschedule Tuesday's meeting to 3pm?",                     # NOT SPAM
        "Earn $5000 from home! Limited time offer! Click the link below!", # SPAM
        "Following up on our discussion about the Q3 roadmap.",            # NOT SPAM
        "You are our lucky winner! Claim your FREE iPad today!",           # SPAM
        "The server logs from yesterday look fine. No anomalies found.",    # NOT SPAM
    ]
}

# Create a DataFrame (like a spreadsheet) from our data
df_train = pd.DataFrame(data)
print(df_train.head())
print(f"\n📊 Dataset size: {len(df_train)} emails")

Key lines explained:

  • ABSTAIN = -1 → The magic number in Snorkel. Any LF returning -1 is saying "I have no opinion here."
  • pd.DataFrame(data) → Creates a spreadsheet-like table from our Python dictionary.
  • In production, replace this toy data with: df = pd.read_csv("my_million_emails.csv")

Step 2 — Write Your Labeling Functions

📝 What the code below does:
We define 6 different Labeling Functions — each one is a Python function decorated with @labeling_function(). That @ decorator is Snorkel's way of saying "treat this Python function as an LF." Each function receives one row from our DataFrame (one email) and returns SPAM, NOT_SPAM, or ABSTAIN. Each LF is a different type: keyword-based, regex-based, heuristic-based, and length-based.
# Step 2: Write our Labeling Functions
# The @labeling_function() decorator registers each function as an official Snorkel LF.
# Each LF receives "x" which is ONE row from our DataFrame — one email.
# We access the text with: x.text

import re
from snorkel.labeling import labeling_function

# ────────────────────────────────────────────────────────────────────────────
# LF 1: KEYWORD-BASED — checks for known spam trigger words
# If any spam keyword is found → SPAM
# If no keywords found → ABSTAIN (not "NOT SPAM"! We just don't know.)
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_spam_keywords(x):
    """
    Checks if the email contains classic spam trigger words.
    This is the most basic type of LF — just a keyword list.
    """
    spam_keywords = ["free prize", "click here", "winner", "limited time",
                     "claim your", "earn $", "lucky winner", "free access"]

    text_lower = x.text.lower()   # Convert to lowercase so matching is case-insensitive

    for keyword in spam_keywords:
        if keyword in text_lower:
            return SPAM     # Found a spam keyword → vote SPAM

    return ABSTAIN          # No spam keywords found → I don't have an opinion


# ────────────────────────────────────────────────────────────────────────────
# LF 2: ALL CAPS DETECTION — spammers love using ALL CAPS to create urgency
# If more than 30% of letter characters are uppercase → SPAM
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_excessive_caps(x):
    """
    Spammers use excessive CAPS to create a sense of urgency.
    If over 30% of letters are uppercase, it's suspicious.
    """
    letters = [c for c in x.text if c.isalpha()]   # Get all letters only

    if len(letters) == 0:
        return ABSTAIN       # No letters at all? Skip it.

    caps_ratio = sum(1 for c in letters if c.isupper()) / len(letters)

    if caps_ratio > 0.30:    # More than 30% caps → suspicious!
        return SPAM
    else:
        return ABSTAIN       # Normal caps ratio → I don't have strong opinion


# ────────────────────────────────────────────────────────────────────────────
# LF 3: EXCLAMATION MARKS — spam emails LOVE exclamation marks!!!
# If 3 or more exclamation marks → SPAM
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_many_exclamations(x):
    """
    Counts exclamation marks. Spam loves urgency = many exclamation marks!
    If 3 or more → vote SPAM. Otherwise → abstain.
    """
    exclamation_count = x.text.count("!")

    if exclamation_count >= 3:
        return SPAM
    return ABSTAIN


# ────────────────────────────────────────────────────────────────────────────
# LF 4: PROFESSIONAL LANGUAGE — work emails use professional business words
# If professional keywords found → NOT SPAM
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_professional_language(x):
    """
    Legitimate business emails contain professional language.
    If we find professional terms → vote NOT SPAM.
    """
    professional_words = ["report", "meeting", "roadmap", "schedule",
                          "reschedule", "attached", "logs", "anomalies",
                          "quarterly", "discussion", "following up"]

    text_lower = x.text.lower()

    for word in professional_words:
        if word in text_lower:
            return NOT_SPAM    # Professional word found → vote NOT SPAM

    return ABSTAIN             # No professional words → I don't know


# ────────────────────────────────────────────────────────────────────────────
# LF 5: REGEX-BASED — detects suspicious URL patterns using regular expressions
# Short URL services (bit.ly, tinyurl, etc.) are common in spam
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_suspicious_url(x):
    """
    Uses regex to detect short/suspicious URLs — very common in phishing spam.
    regex pattern: looks for bit.ly, tinyurl, or similar URL shorteners.
    """
    # This pattern matches common URL shortener patterns
    url_pattern = re.compile(r"(bit\.ly|tinyurl|t\.co|goo\.gl|ow\.ly)/\S+",
                             re.IGNORECASE)

    if url_pattern.search(x.text):
        return SPAM        # Found a suspicious short URL → SPAM
    return ABSTAIN         # No suspicious URL → I don't know


# ────────────────────────────────────────────────────────────────────────────
# LF 6: HEURISTIC — combines two signals: urgency words + money mentions
# If BOTH are present → very likely SPAM
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_urgency_plus_money(x):
    """
    Spam often combines URGENCY ("URGENT", "NOW", "TODAY") with MONEY mentions.
    This heuristic LF checks for both signals at once.
    If both are present → strong vote for SPAM.
    """
    text_lower = x.text.lower()

    urgency_words = ["urgent", "now!", "today!", "immediately", "act now"]
    money_pattern = re.compile(r"\$\d+|\bfree\b|\bprize\b", re.IGNORECASE)

    has_urgency = any(word in text_lower for word in urgency_words)
    has_money   = bool(money_pattern.search(x.text))

    if has_urgency and has_money:
        return SPAM      # Both signals present → strong SPAM vote
    return ABSTAIN       # Not both present → I'm not sure

Key lines explained:

  • @labeling_function() → The Snorkel decorator that registers your function as an LF. Without this, it's just a regular Python function.
  • x.text → Accesses the "text" column of the current row (one email).
  • x.text.lower() → Converts to lowercase so "FREE" matches "free".
  • re.compile(pattern) → Creates a compiled regex pattern for efficient repeated matching.
  • Every LF returns ABSTAIN as the default when it doesn't see its specific signal. This is critical!

Step 3 — Apply All LFs to Your Data

📝 What the code below does:
We use Snorkel's PandasLFApplier to run ALL our LFs on ALL our data at once. This produces the Label Matrix — a table where each row is an email and each column is one LF's vote. Then we use LFAnalysis to print the report card for all our LFs: coverage, polarity, overlaps, and conflicts.
# Step 3: Apply all LFs to our dataset and see the Label Matrix

from snorkel.labeling import PandasLFApplier, LFAnalysis

# Put all LFs into a list — the order matters!
lfs = [
    lf_spam_keywords,
    lf_excessive_caps,
    lf_many_exclamations,
    lf_professional_language,
    lf_suspicious_url,
    lf_urgency_plus_money
]

# PandasLFApplier runs every LF from our list on every row of our DataFrame
applier = PandasLFApplier(lfs=lfs)

# Apply! This returns the Label Matrix — shape is (n_emails, n_lfs)
# Each cell is 1 (SPAM), 0 (NOT SPAM), or -1 (ABSTAIN)
L_train = applier.apply(df=df_train)

print("📊 Label Matrix (rows=emails, cols=LFs):")
print("   LF1  LF2  LF3  LF4  LF5  LF6")
print(L_train)

print("\n📋 LF Report Card:")
# LFAnalysis gives you the quality metrics for each LF
# Pass the label matrix and the list of LF names
lf_analysis = LFAnalysis(L=L_train, lfs=lfs).lf_summary()
print(lf_analysis)

Key lines explained:

  • PandasLFApplier(lfs=lfs) → Creates the "runner" that applies all LFs. It's smart — it can run LFs in parallel for speed.
  • applier.apply(df=df_train) → The magic line. Runs every LF on every row and returns the full Label Matrix as a NumPy array.
  • LFAnalysis(L=L_train, lfs=lfs).lf_summary() → Produces the report card table showing Coverage, Polarity, Overlaps, and Conflicts for each LF.

Step 4 — Train the Label Model (Combine LF Votes)

📝 What the code below does:
This is the most important step — the Label Model. It reads the Label Matrix (all the LF votes) and learns which LFs are more trustworthy, which ones tend to conflict, and which ones make systematic errors. It then produces a probabilistic label for each data point — a confidence score like "87% probability this is SPAM." We can also convert these to hard labels (SPAM or NOT SPAM) using predict().
# Step 4: Train the Label Model to intelligently combine all LF votes

from snorkel.labeling.model import LabelModel

# LabelModel is the "smart referee" that learns which LFs to trust more.
# cardinality=2 means binary classification (SPAM vs NOT SPAM)
# verbose=True prints progress during training
label_model = LabelModel(cardinality=2, verbose=True)

# Fit the Label Model on the Label Matrix
# It learns the reliability and correlations of each LF without seeing true labels!
# n_epochs=200 = number of training rounds for the label model itself
label_model.fit(L_train=L_train, n_epochs=200, seed=42)

print("✅ Label Model trained!")

# ── Get probabilistic labels ──────────────────────────────────────────────────
# predict_proba gives [P(NOT_SPAM), P(SPAM)] for each email
# E.g., [0.08, 0.92] means 92% probability of being SPAM
probas = label_model.predict_proba(L=L_train)
print("\n📊 Probabilistic labels (first 5 emails):")
print("   [P(NOT_SPAM), P(SPAM)]")
for i, prob in enumerate(probas[:5]):
    print(f"   Email {i+1}: NOT_SPAM={prob[0]:.2f}, SPAM={prob[1]:.2f}")

# ── Get hard (definitive) labels ──────────────────────────────────────────────
# predict() returns the final label: 0 (NOT_SPAM) or 1 (SPAM) for each email
# tie_break_policy="abstain" means: if truly uncertain, output -1 (don't guess)
labels = label_model.predict(L=L_train, tie_break_policy="abstain")

df_train["weak_label"] = labels
print("\n📧 Emails with automatically generated labels:")
for _, row in df_train.iterrows():
    label_name = {1: "SPAM", 0: "NOT_SPAM", -1: "ABSTAIN"}[row["weak_label"]]
    print(f"   Email {row['email_id']}: {label_name:10} | {row['text'][:55]}...")

Key lines explained:

  • LabelModel(cardinality=2) → Creates the label model. cardinality=2 means two classes (binary). For a 5-class problem, use cardinality=5.
  • label_model.fit(L_train=L_train) → Trains the label model on the Label Matrix. It learns LF reliability without needing any ground truth labels!
  • predict_proba() → Returns soft probabilities — more nuanced than just a hard label.
  • predict(tie_break_policy="abstain") → Returns a definitive label. When tied, returns ABSTAIN instead of guessing randomly.

Step 5 — Train Your Final ML Model on the Weak Labels

📝 What the code below does:
Now we use the automatically generated labels to train a real Machine Learning model. Here we use a simple Logistic Regression classifier (from scikit-learn). In production, you would use a more powerful model (BERT, XGBoost, neural network). The key insight: the final model generalises beyond the LFs — it learns patterns from the data itself, not just the LF rules. A classifier trained on 10,000 weakly labelled examples can outperform the LFs that created those labels!
# Step 5: Train a real ML model using the weak labels from Step 4
# In this example, we use a simple TF-IDF text vectoriser + Logistic Regression.
# In production: replace with BERT, RoBERTa, XGBoost, or any powerful model!

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report

# ── Filter out ABSTAIN labels before training ──────────────────────────────
# We only want to train on data points where the label model gave a confident label
# Data points where tie_break_policy produced ABSTAIN (-1) are excluded
df_labeled = df_train[df_train["weak_label"] != ABSTAIN].copy()
print(f"✅ Training on {len(df_labeled)} examples with confident labels")
print(f"   (Excluded {len(df_train) - len(df_labeled)} ABSTAIN examples)")

# ── Create text features using TF-IDF ──────────────────────────────────────
# TF-IDF converts raw text into numerical features the ML model can understand
# Think of it as: each word gets a score based on how important it is in this email
vectorizer = TfidfVectorizer(ngram_range=(1, 2), max_features=5000)
X_train = vectorizer.fit_transform(df_labeled["text"])
y_train = df_labeled["weak_label"]

# ── Train the Logistic Regression classifier ───────────────────────────────
# LogisticRegression is a simple but powerful text classification algorithm
clf = LogisticRegression(random_state=42, max_iter=1000)
clf.fit(X_train, y_train)

print("\n✅ Classifier trained on weakly labeled data!")

# ── Test on new, unseen emails ──────────────────────────────────────────────
new_emails = [
    "Win a free iPhone! Click now to claim your prize!",
    "Attached is the slides deck for tomorrow's all-hands meeting.",
    "FREE MONEY - limited offer - Act now or lose your chance!",
]

X_new = vectorizer.transform(new_emails)
predictions = clf.predict(X_new)

print("\n🔮 Predictions on new emails:")
for email, pred in zip(new_emails, predictions):
    label = "SPAM 🚨" if pred == SPAM else "NOT SPAM ✅"
    print(f"   {label}: {email[:60]}...")

8. Trend — LLM-Powered Labeling Functions

One of the most exciting developments is using Large Language Models (LLMs) like GPT-4, Llama 3, or Mistral as labeling functions. Instead of writing code, you write a natural language prompt — and the LLM's answer becomes the LF's vote.

This technique, sometimes called Prompted Weak Supervision, was pioneered in a 2022 Stanford paper and has become mainstream . Companies are now using it to label medical records, legal contracts, financial reports — all domains where expert rules are hard to write in code but easy to describe in English!

📝 What the code below does:
This shows how to create an LLM-powered Labeling Function. We write a function that takes an email, builds a structured prompt, sends it to an LLM (here using OpenAI's GPT-4o-mini as an example), and maps the LLM's answer back to SPAM / NOT_SPAM / ABSTAIN. This LF can then be added to your LF list alongside your rule-based LFs — the Label Model will automatically learn how much to trust it!
# LLM-Powered Labeling Function 
# This LF asks GPT-4o-mini to classify each email and maps the answer to a label.
# You need: pip install openai

import openai
from snorkel.labeling import labeling_function

# Set your API key (in production, use environment variables — never hardcode!)
# openai.api_key = os.environ.get("OPENAI_API_KEY")
client = openai.OpenAI()   # reads OPENAI_API_KEY from environment automatically

@labeling_function()
def lf_llm_classifier(x):
    """
    Uses GPT-4o-mini as a labeling function.
    Sends the email to the LLM with a clear prompt and maps the answer to a label.
    This is the way to get powerful labeling without writing complex rules!
    """
    prompt = f"""You are a spam email classifier.

Classify the following email as SPAM or NOT_SPAM or UNSURE.
Rules:
- SPAM: unsolicited commercial email, phishing, prize scams, too-good-to-be-true offers
- NOT_SPAM: professional business communication, meeting requests, work discussions
- UNSURE: you are not confident about the classification

Email text:
"{x.text}"

Reply with ONLY one of these three words: SPAM, NOT_SPAM, or UNSURE.
No explanation needed."""

    try:
        response = client.chat.completions.create(
            model="gpt-4o-mini",          # Use a fast, cheap model for LF labeling
            messages=[{"role": "user", "content": prompt}],
            max_tokens=10,                 # We only need one word back
            temperature=0.0               # temperature=0 = deterministic, no randomness
        )
        # Extract the LLM's answer and strip whitespace
        answer = response.choices[0].message.content.strip().upper()

        # Map LLM answer → Snorkel label
        if answer == "SPAM":
            return SPAM
        elif answer == "NOT_SPAM":
            return NOT_SPAM
        else:
            return ABSTAIN    # "UNSURE" or unexpected answer → abstain

    except Exception as e:
        # If the API call fails for any reason → abstain (don't crash the whole pipeline)
        print(f"LLM LF error: {e}")
        return ABSTAIN


# ── Use this LLM LF alongside your rule-based LFs ────────────────────────────
# The Label Model will automatically learn to weight the LLM LF's votes
# against your other rule-based LFs!
# lfs_with_llm = [
#     lf_spam_keywords,
#     lf_excessive_caps,
#     lf_many_exclamations,
#     lf_professional_language,
#     lf_suspicious_url,
#     lf_urgency_plus_money,
#     lf_llm_classifier        # ← Add the LLM LF just like any other LF!
# ]

Key lines explained:

  • temperature=0.0 → Makes the LLM deterministic — same input always gives same output. Critical for reproducible labels!
  • max_tokens=10 → We only need one word ("SPAM" or "NOT_SPAM"). Limiting tokens saves time and API cost.
  • try/except → ABSTAIN → If the LLM API call fails (network error, rate limit), we safely abstain instead of crashing.
  • The LLM LF is added to the list exactly like any other LF — the Label Model doesn't care whether it's rule-based or LLM-based!
💡 Cost Tip for LLM LFs:
Running an LLM API call for each of 1,000,000 emails could get expensive. The trick: run the LLM LF on a representative sample (e.g., 10,000 emails), then train the Label Model on both the sample LLM votes + the full rule-based LF votes. Or use a small, fast local model (Llama 3.1 8B) — free and runs on your own machine!

9. 🗺️ The Complete LF Pipeline — From Raw Data to Trained Model

Let's put everything together and see the full picture. This is the exact workflow used by ML teams at Google, Stanford Medicine, and Fortune 500 companies:

 ═══════════════════════════════════════════════════════════════════════
              COMPLETE LABELING FUNCTION PIPELINE
 ═══════════════════════════════════════════════════════════════════════

 PHASE 1: DATA COLLECTION
 ─────────────────────────
 Collect large amounts of UNLABELED raw data
 (emails, documents, images, medical records...)
 No labels needed yet! → df_train (no "y" column)

         ↓

 PHASE 2: EXPLORE & UNDERSTAND DATA (EDA)
 ─────────────────────────────────────────
 Read samples. Find patterns. Talk to domain experts.
 "What signals separate class A from class B?"

         ↓

 PHASE 3: WRITE LABELING FUNCTIONS
 ────────────────────────────────────────────────────────────────
 LF Type 1: Keyword LFs        (simple word matches)
 LF Type 2: Regex LFs          (pattern matching)
 LF Type 3: Heuristic LFs      (if-then rules from experts)
 LF Type 4: Distant supervision (external knowledge bases)
 LF Type 5: Model-based LFs    (old/pretrained models)
 LF Type 6: LLM-powered LFs    (GPT/Llama prompts)  ← New in 2025!

         ↓

 PHASE 4: APPLY LFs → GET LABEL MATRIX
 ──────────────────────────────────────
 PandasLFApplier.apply(df_train)
 Result: Label Matrix L_train (n_examples × n_lfs)
 Each cell: 1, 0, or -1 (ABSTAIN)

         ↓

 PHASE 5: ANALYSE LF QUALITY
 ────────────────────────────
 LFAnalysis(L_train, lfs).lf_summary()
 Check: Coverage, Polarity, Overlaps, Conflicts
 Remove or fix LFs with low coverage or high conflict

         ↓

 PHASE 6: TRAIN LABEL MODEL
 ───────────────────────────
 LabelModel.fit(L_train)
 Learns: which LFs are trustworthy, which conflict
 Outputs: probabilistic labels  [P(class0), P(class1)]

         ↓

 PHASE 7: FILTER + GET TRAINING LABELS
 ──────────────────────────────────────
 label_model.predict(L_train, tie_break_policy="abstain")
 Filter out ABSTAIN examples
 Result: df_labeled with "weak_label" column

         ↓

 PHASE 8: TRAIN FINAL END MODEL
 ───────────────────────────────
 Use df_labeled to train any ML model:
 → BERT, XGBoost, Logistic Regression, Neural Network
 The end model GENERALISES beyond the LF rules!

         ↓

 PHASE 9: EVALUATE + ITERATE
 ─────────────────────────────
 Test on a small gold-standard validation set
 If accuracy too low → add more LFs, fix existing ones
 Repeat from Phase 3

 ═══════════════════════════════════════════════════════════════════════

10. ✅❌ Best Practices — DOs and DON'Ts

✅ DO: Write many diverse LFs, not just many similar ones
10 LFs that each catch a different pattern are far more powerful than 10 LFs that all check for the same keyword. Diversity reduces the chance of systematic errors — when LFs disagree, the Label Model learns from the disagreement!
❌ DON'T: Worry about making each LF perfect
This is the most common beginner mistake. An LF with 60% accuracy that has good coverage is much more valuable than an LF with 95% accuracy that only covers 1% of your data. Imperfect is fine — that's the whole point of combining many LFs!
✅ DO: Use ABSTAIN generously
When your LF is not confident, return ABSTAIN. A LF that abstains 70% of the time but is 95% accurate on the 30% it does label is genuinely useful. A LF that never abstains but is only 55% accurate is almost useless (barely better than a coin flip).
❌ DON'T: Forget to hold out a validation set
Always keep a small set of manually labeled examples (100–500 samples) that no LF ever sees. Use this validation set to measure the true quality of your Label Model and final ML model. Without this, you're flying blind!
✅ DO: Involve domain experts in writing LFs
A doctor can write a medical LF in 5 minutes using domain knowledge it would take a data scientist months to learn. LFs are the perfect interface between subject matter experts and ML systems. The expert describes the rule in plain English; the developer translates it to 3 lines of Python.
❌ DON'T: Use LFs as a replacement for your ML model
LFs generate training labels. They are NOT the final classifier. A key insight from Snorkel's research: the final ML model trained on weak labels consistently outperforms using the LFs directly as a classifier — because the model generalises beyond the rules!
✅ DO: Monitor and update LFs over time
The real world changes. New types of spam appear. New business terms emerge. LFs need to be updated as data distributions shift. The great news: updating a rule takes minutes, while re-labeling a dataset takes months!

11. 📝 Quick Reference — Everything in One Place

LF Types Cheat Sheet

  • Keyword LF → Match a word from a list → fastest to write, good first LF to build
  • Regex LF → Match a pattern (URL, phone, date format) → powerful for structured text
  • Heuristic LF → if-then rule from domain expert → most flexible, requires expertise
  • Distant Supervision LF → lookup in external database/knowledge base → great for named entities
  • Model-Based LF → use an old pretrained model as a voter → reuse existing work
  • Crowdsourced LF → use click/rating data as weak signal → good for user-facing products
  • LLM-Powered LF → prompt GPT/Llama to label → most powerful for fuzzy concepts 

Key Numbers to Know

  • 🏅 Research result: Teams with Snorkel built models 2.8x faster with 45.5% better performance vs manual labeling
  • 📊 Minimum LFs: Start with at least 5–10 LFs. The Label Model needs enough signal to learn from.
  • 🎯 Coverage goal: Aim for each LF to cover at least 5–10% of your dataset. Otherwise it's too sparse to be useful.
  • ⚖️ Accuracy floor: An LF should be right at least 55–60% of the time when it doesn't abstain. Below that, it might hurt more than help.

Core Snorkel Functions Cheat Sheet

  • @labeling_function() → Decorator to register a Python function as a Snorkel LF
  • PandasLFApplier(lfs=lfs).apply(df) → Runs all LFs on a DataFrame, returns Label Matrix
  • LFAnalysis(L, lfs).lf_summary() → Prints the quality report card for all LFs
  • LabelModel(cardinality=n) → Creates the Label Model (n = number of classes)
  • label_model.fit(L_train) → Trains the Label Model on the Label Matrix
  • label_model.predict_proba(L) → Returns soft probability labels for each data point
  • label_model.predict(L, tie_break_policy="abstain") → Returns hard labels (0, 1, or -1)

Happy labeling! 🏷️🤖✨

Comments