Labeling Functions (LFs) in Machine Learning: A Practical Guide to Weak Supervision
Imagine you are the teacher of a class of 1,000,000 students. You need to give every single student a grade: Pass or Fail. Reading every student's paper by hand would take your entire lifetime!
Now imagine you could write a few clever rules instead — "If the paper mentions the formula correctly, probably Pass. If it's under 50 words, probably Fail." You write 10 such rules, and they automatically grade all 1,000,000 papers for you in minutes!
That's exactly what Labeling Functions do in Machine Learning. They are small, smart rules that automatically label enormous amounts of data — saving months of human effort, and costing almost nothing.
One-line summary: A Labeling Function is a mini-rule you write that says: "If you see THIS pattern, give THAT label."
📋 What You Will Learn
- The biggest problem in Machine Learning — the labeling bottleneck
- What Weak Supervision is and why it's a game-changer
- What a Labeling Function is (with simple analogies)
- The 7 types of Labeling Functions with real examples
- What ABSTAIN means and why it's a superpower
- How multiple LFs vote together (the Label Model)
- The 4 key metrics to measure LF quality
- Full hands-on code: writing LFs with Snorkel (the most popular LF library)
- The trend: LLM-powered Labeling Functions
- The full pipeline: LFs → Label Model → Train your ML model
- Best practices, DOs and DON'Ts
1. 😩 The Biggest Problem in ML — "We Need Labels!"
Machine Learning models learn by example. You show them thousands of examples with correct answers, and they learn the pattern. But here's the painful truth:
Getting those "correct answers" (called labels) is incredibly hard, slow, and expensive.
🧒 Think of it like this:
Imagine you want to teach a robot to recognise spam emails.
You need to show it 100,000 emails with each one marked: SPAM or NOT SPAM.
A human expert has to read and mark every single one — that's months of work!
If your definition of "spam" changes tomorrow, you start all over again.
This problem — called the labeling bottleneck — is the #1 reason most ML projects fail or take forever.
THE LABELING BOTTLENECK UNLABELED DATA MANUAL LABELING LABELED DATA ───────────── ─────────────── ──────────── Email 1: ??? →→ Human reads it →→ Email 1: SPAM Email 2: ??? →→ Human reads it →→ Email 2: NOT SPAM Email 3: ??? →→ Human reads it →→ Email 3: SPAM ... Email 100,000: →→ [3 months later] →→ Email 100,000: ??? ⚠️ Cost: $$$$$ Time: Months Scalability: 😭
Researchers at Stanford found that labeling training data has become the single largest bottleneck in deploying production ML systems. Companies like Google have hired entire teams of hundreds of people just to label data!
2. 💡 Weak Supervision — The Smart Shortcut
Weak Supervision is a completely different approach. Instead of labeling data one by one, you write rules, heuristics, and shortcuts that automatically assign labels to thousands of data points at once.
These rules won't be perfect. Some will make mistakes. But here's the clever part: you use many imperfect rules together, and a mathematical model figures out the best combined answer — just like a jury reaching a verdict!
🧒 The jury analogy:
Imagine 10 jurors in a courtroom, each with partial evidence. Juror 1 is very reliable. Juror 5 is sometimes wrong. The judge (Label Model) doesn't take a simple vote — it weighs each juror's reliability automatically and reaches the best possible verdict.
WEAK SUPERVISION vs MANUAL LABELING
MANUAL LABELING WEAK SUPERVISION
─────────────── ────────────────
Speed: Months Hours / Days
Cost: Very High ($$$$$) Very Low ($)
Scalability: Limited Millions of examples
Flexibility: Rigid Update rules in minutes
Accuracy: High (if careful) Good (improves with more LFs)
Used at: Small startups Google, Intel, Stanford,
BNY Mellon, Chubb Insurance...
Research from Stanford's AI Lab showed that teams using Weak Supervision built models 2.8x faster and increased predictive performance by an average of 45.5% compared to traditional 7-hour manual labeling sessions!
3. 🏷️ What Exactly is a Labeling Function?
A Labeling Function (LF) is a Python function you write that:
- Takes ONE data point as input (e.g., an email text)
- Looks at it using some rule or heuristic
- Returns ONE of three things: a positive label, a negative label, or ABSTAIN (I don't know)
That's it. It's really that simple! The LF doesn't need to be right 100% of the time. It just needs to be right more often than random guessing.
The Three Possible Outputs of a Labeling Function
LABELING FUNCTION OUTPUTS ────────────────────────── Input: An email text LF Output 1 → SPAM (1) "I think this IS spam" LF Output 2 → NOT SPAM (0) "I think this is NOT spam" LF Output 3 → ABSTAIN (-1) "I have no idea, skip me" ABSTAIN is the LF saying: "This data point is outside my area of expertise. Let other LFs handle it. Don't count my vote here."
A LF that abstains when it's uncertain is much better than one that guesses randomly. Abstaining prevents a bad LF from polluting good labels on examples it knows nothing about. Think of a doctor saying "I'm not sure, let a specialist check this" — that's wisdom, not failure!
A Simple Mental Model
🧒 Think of LFs as detectives, each specialising in one clue:
- Detective 1 checks: "Does the email mention FREE MONEY?" → If yes: SPAM. If not: ABSTAIN.
- Detective 2 checks: "Is it from a known colleague's email address?" → If yes: NOT SPAM. If not: ABSTAIN.
- Detective 3 checks: "Does it have 50+ exclamation marks?" → If yes: SPAM. If not: ABSTAIN.
- Detective 4 checks: "Does it contain a meeting invitation?" → If yes: NOT SPAM. If not: ABSTAIN.
No single detective catches every spam. But together, when they all vote, the combined verdict is surprisingly accurate!
4. 🔧 The 7 Types of Labeling Functions (with Examples)
Labeling Functions come in many flavours. You choose the type based on what kind of "signal" you have available. Let's explore all of them with the spam email detection example throughout:
Type 1 — Keyword-Based LF
The simplest type. You give the LF a list of words or phrases. If the data point contains any of them, return a label.
🧒 Analogy: A security guard with a list of forbidden words. If you say one of them at the door, you're flagged!
KEYWORD LF — MENTAL MODEL
Spam keywords list: ["FREE MONEY", "CLICK HERE", "WINNER", "PRIZE", "$$$$"]
Email A: "You WON a PRIZE! CLICK HERE to claim."
→ Contains "PRIZE" + "CLICK HERE" → Label: SPAM ✅
Email B: "Let's schedule a meeting for Tuesday."
→ No spam keywords → Label: ABSTAIN 🔇
Type 2 — Regex (Pattern) LF
Uses regular expressions — pattern-matching rules for text. Perfect for structured patterns like phone numbers, URLs, email formats, dollar amounts.
🧒 Analogy: Like using a template. "Does this sentence fit the shape of a phone number pattern?"
REGEX LF — MENTAL MODEL
Pattern: "Emails with suspicious URLs (e.g., bit.ly/xxx)"
Email A: "Claim your prize here: bit.ly/abc123"
→ Matches URL shortener pattern → Label: SPAM ✅
Email B: "Project report attached. See slide 3."
→ No URL pattern → Label: ABSTAIN 🔇
Type 3 — Heuristic / Rule-Based LF
Uses domain expertise in the form of if-then logic. A human expert writes: "If X AND Y, then label Z." These are the most flexible type!
🧒 Analogy: Your grandma's wisdom. "If someone calls at midnight AND asks for money, it's suspicious."
HEURISTIC LF — MENTAL MODEL
Rule: "If email is ALL CAPS AND has more than 3 exclamation marks → SPAM"
Email A: "YOU WON!!! CLAIM NOW!!! LIMITED TIME!!!"
→ ALL CAPS: Yes. Exclamations: 6 → Label: SPAM ✅
Email B: "Quick reminder about tomorrow's meeting!"
→ Not ALL CAPS. Exclamations: 1 → Label: ABSTAIN 🔇
Type 4 — Distant Supervision LF
Uses an external knowledge base (a list, database, or lookup table) to generate labels. If a data point matches an entry in the knowledge base, apply the known label.
🧒 Analogy: A bouncer with a VIP guest list. If your name is on the list, you're automatically in (or out).
DISTANT SUPERVISION LF — MENTAL MODEL
Known spam sender list: ["spammer@phishing.com", "noreply@scam.ru"]
Email A: From "spammer@phishing.com"
→ Matches spam sender list → Label: SPAM ✅
Email B: From "alice@mycompany.com"
→ Not in any list → Label: ABSTAIN 🔇
Type 5 — Model-Based LF
Uses a pre-trained ML model as a labeling function! Maybe you have an old, slightly inaccurate model from a previous project. Instead of throwing it away, you use it as one voter among many.
🧒 Analogy: An experienced but slightly rusty expert. Their judgement isn't perfect, but it's still useful when combined with others.
MODEL-BASED LF — MENTAL MODEL
Old spam classifier (trained 2 years ago, 78% accuracy):
→ Predicts probability of spam for each email
→ If probability > 0.85 → Label: SPAM
→ If probability < 0.15 → Label: NOT SPAM
→ Otherwise → ABSTAIN (it's not confident enough)
Type 6 — Crowdsourced / Human LF
Uses input from non-expert humans (e.g., via surveys, clicks, ratings, or feedback data) as a source of weak labels. Each individual human label might be noisy, but aggregated, they're powerful!
🧒 Analogy: Asking 1,000 people "Is this funny?" — even if some people have bad taste in humour, the majority vote is usually right!
Type 7 — LLM-Powered LF (The Trend! 🔥)
The newest and most exciting type. Instead of writing rules in code, you write a natural language prompt and ask a Large Language Model (like GPT or Llama) to label the data. The LLM's answer becomes the LF's vote.
🧒 Analogy: Instead of training your robot guard, you just tell it in plain English: "If the email sounds suspicious or pushy, it's spam."
LLM-POWERED LF — MENTAL MODEL
Prompt to LLM:
"The following email was received. Is it spam?
Email: {email_text}
Answer with ONLY: 'SPAM' or 'NOT SPAM' or 'UNSURE'"
Email A: "CONGRATULATIONS! You've been selected for a free gift!"
→ LLM: "SPAM" → Label: SPAM ✅
Email B: "Please review the attached quarterly report."
→ LLM: "NOT SPAM" → Label: NOT SPAM ✅
Email C: "Hi, following up on our last chat."
→ LLM: "UNSURE" → Label: ABSTAIN 🔇
5. 📊 The Label Matrix — How All LFs Work Together
When you have multiple LFs, you apply all of them to all your data points. The result is called the Label Matrix — a big table of votes.
THE LABEL MATRIX
Each row = one data point (email)
Each column = one Labeling Function
Values: 1=SPAM, 0=NOT SPAM, -1=ABSTAIN
LF1 LF2 LF3 LF4 LF5
Email 1: 1 1 -1 1 -1 → Probably SPAM (3 votes for SPAM)
Email 2: 0 -1 0 -1 0 → Probably NOT SPAM (3 votes against)
Email 3: -1 -1 -1 -1 -1 → No coverage! (all abstain)
Email 4: 1 -1 0 1 0 → Conflicted (2 SPAM, 2 NOT SPAM)
Email 5: 1 1 1 0 -1 → Probably SPAM (3 SPAM, 1 against)
↓
Label Model reads this matrix and computes the BEST GUESS label for each row.
Notice Email 4 has a conflict — some LFs say SPAM, some say NOT SPAM. The Label Model resolves this by figuring out which LFs are more trustworthy and weighting their votes accordingly. It's like the judge deciding which witnesses in a trial are more credible!
6. 📏 The 4 Key Metrics to Judge Your LFs
After applying your LFs to your data, Snorkel gives you 4 important statistics about each LF. Think of these as a report card for each LF.
Metric 1 — Coverage
What fraction of your data does this LF label? (i.e., how often does it NOT abstain?)
🧒 Like asking: "How many students did this teacher actually grade?" A teacher who only graded 3 out of 1000 papers isn't very useful!
Coverage = (Number of non-ABSTAIN labels) / (Total data points) LF1 labels 800 out of 1000 emails → Coverage = 0.80 (80%) ✅ Good! LF2 labels 12 out of 1000 emails → Coverage = 0.01 (1%) ⚠️ Too low!
Metric 2 — Polarity
What unique labels does this LF output? (e.g., does it only say SPAM, or does it say both SPAM and NOT SPAM?)
🧒 Like asking: "Does this judge only ever convict people, or does she also acquit?" A judge who always says "Guilty!" regardless of evidence is suspicious!
Metric 3 — Overlaps
How often does this LF label the same data point as at least one other LF? High overlap = this LF's territory is well-covered by other LFs too.
Metric 4 — Conflicts
How often does this LF disagree with another LF on the same data point? High conflicts = the LF contradicts others a lot. Some conflict is normal and healthy — it means your LFs are covering different signals. But if one LF conflicts with ALL others, it might be wrong!
LF REPORT CARD (from Snorkel)
Coverage Polarity Overlaps Conflicts
LF1: 0.82 [0,1] 0.65 0.08 ← Healthy LF ✅
LF2: 0.14 [1] 0.09 0.03 ← Low coverage ⚠️
LF3: 0.91 [0,1] 0.78 0.22 ← High conflict, check it ⚠️
LF4: 0.55 [0] 0.40 0.00 ← Only says NOT SPAM, ok ✅
LF5: 0.03 [1] 0.02 0.01 ← Almost useless ❌
7. 💻 Hands-On Code — Building Labeling Functions with Snorkel
Snorkel is the most popular and powerful Python library for Labeling Functions. It was built at Stanford AI Lab, published in a landmark research paper, and is now used by companies like Google, Intel, BNY Mellon, and Stanford Medicine.
🧒 Think of Snorkel as a smart referee. You bring your team of imperfect judges (LFs). Snorkel watches how they vote, figures out who's most trustworthy, and gives you the final score!
First — Install Snorkel
This installs the Snorkel library onto your computer using pip (Python's package manager). Snorkel contains all the tools we need: decorators for writing LFs, the LF applier (runs LFs on data), and the Label Model (combines LF votes).
pip install snorkel pandas scikit-learn
Step 1 — Set Up Our Data
We create a tiny dataset of 8 emails. Each email is stored in a Pandas DataFrame (think of it as a spreadsheet in Python). We also define our label constants: SPAM=1, NOT_SPAM=0, ABSTAIN=-1. In real projects, this dataset would have thousands or millions of rows — all without labels! That's the whole point — we'll use LFs to label them automatically.
# Step 1: Import libraries and create our email dataset
# Pandas lets us work with data in a table format (like Excel in Python)
# Snorkel's ABSTAIN constant (-1) means "I don't know" — our LF skips this example
import pandas as pd
from snorkel.labeling import labeling_function
# ── Define what our labels mean ─────────────────────────────────────────────
ABSTAIN = -1 # LF says: "I have no opinion on this one"
NOT_SPAM = 0 # LF says: "This is definitely NOT spam"
SPAM = 1 # LF says: "This IS spam"
# ── Create a small example email dataset ────────────────────────────────────
# In a real project you'd load millions of unlabeled emails from a file or database.
# For learning, we use 8 simple examples with known answers (we'll pretend we don't know them!).
data = {
"email_id": [1, 2, 3, 4, 5, 6, 7, 8],
"text": [
"Congratulations! You've WON a FREE prize! CLICK HERE NOW!!!", # SPAM
"Hi, please review the attached quarterly report by Friday.", # NOT SPAM
"URGENT: Your account will be closed! Verify NOW for FREE access", # SPAM
"Can we reschedule Tuesday's meeting to 3pm?", # NOT SPAM
"Earn $5000 from home! Limited time offer! Click the link below!", # SPAM
"Following up on our discussion about the Q3 roadmap.", # NOT SPAM
"You are our lucky winner! Claim your FREE iPad today!", # SPAM
"The server logs from yesterday look fine. No anomalies found.", # NOT SPAM
]
}
# Create a DataFrame (like a spreadsheet) from our data
df_train = pd.DataFrame(data)
print(df_train.head())
print(f"\n📊 Dataset size: {len(df_train)} emails")
Key lines explained:
ABSTAIN = -1→ The magic number in Snorkel. Any LF returning -1 is saying "I have no opinion here."pd.DataFrame(data)→ Creates a spreadsheet-like table from our Python dictionary.- In production, replace this toy data with:
df = pd.read_csv("my_million_emails.csv")
Step 2 — Write Your Labeling Functions
We define 6 different Labeling Functions — each one is a Python function decorated with
@labeling_function().
That @ decorator is Snorkel's way of saying "treat this Python function as an LF."
Each function receives one row from our DataFrame (one email) and returns SPAM, NOT_SPAM, or ABSTAIN.
Each LF is a different type: keyword-based, regex-based, heuristic-based, and length-based.
# Step 2: Write our Labeling Functions
# The @labeling_function() decorator registers each function as an official Snorkel LF.
# Each LF receives "x" which is ONE row from our DataFrame — one email.
# We access the text with: x.text
import re
from snorkel.labeling import labeling_function
# ────────────────────────────────────────────────────────────────────────────
# LF 1: KEYWORD-BASED — checks for known spam trigger words
# If any spam keyword is found → SPAM
# If no keywords found → ABSTAIN (not "NOT SPAM"! We just don't know.)
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_spam_keywords(x):
"""
Checks if the email contains classic spam trigger words.
This is the most basic type of LF — just a keyword list.
"""
spam_keywords = ["free prize", "click here", "winner", "limited time",
"claim your", "earn $", "lucky winner", "free access"]
text_lower = x.text.lower() # Convert to lowercase so matching is case-insensitive
for keyword in spam_keywords:
if keyword in text_lower:
return SPAM # Found a spam keyword → vote SPAM
return ABSTAIN # No spam keywords found → I don't have an opinion
# ────────────────────────────────────────────────────────────────────────────
# LF 2: ALL CAPS DETECTION — spammers love using ALL CAPS to create urgency
# If more than 30% of letter characters are uppercase → SPAM
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_excessive_caps(x):
"""
Spammers use excessive CAPS to create a sense of urgency.
If over 30% of letters are uppercase, it's suspicious.
"""
letters = [c for c in x.text if c.isalpha()] # Get all letters only
if len(letters) == 0:
return ABSTAIN # No letters at all? Skip it.
caps_ratio = sum(1 for c in letters if c.isupper()) / len(letters)
if caps_ratio > 0.30: # More than 30% caps → suspicious!
return SPAM
else:
return ABSTAIN # Normal caps ratio → I don't have strong opinion
# ────────────────────────────────────────────────────────────────────────────
# LF 3: EXCLAMATION MARKS — spam emails LOVE exclamation marks!!!
# If 3 or more exclamation marks → SPAM
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_many_exclamations(x):
"""
Counts exclamation marks. Spam loves urgency = many exclamation marks!
If 3 or more → vote SPAM. Otherwise → abstain.
"""
exclamation_count = x.text.count("!")
if exclamation_count >= 3:
return SPAM
return ABSTAIN
# ────────────────────────────────────────────────────────────────────────────
# LF 4: PROFESSIONAL LANGUAGE — work emails use professional business words
# If professional keywords found → NOT SPAM
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_professional_language(x):
"""
Legitimate business emails contain professional language.
If we find professional terms → vote NOT SPAM.
"""
professional_words = ["report", "meeting", "roadmap", "schedule",
"reschedule", "attached", "logs", "anomalies",
"quarterly", "discussion", "following up"]
text_lower = x.text.lower()
for word in professional_words:
if word in text_lower:
return NOT_SPAM # Professional word found → vote NOT SPAM
return ABSTAIN # No professional words → I don't know
# ────────────────────────────────────────────────────────────────────────────
# LF 5: REGEX-BASED — detects suspicious URL patterns using regular expressions
# Short URL services (bit.ly, tinyurl, etc.) are common in spam
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_suspicious_url(x):
"""
Uses regex to detect short/suspicious URLs — very common in phishing spam.
regex pattern: looks for bit.ly, tinyurl, or similar URL shorteners.
"""
# This pattern matches common URL shortener patterns
url_pattern = re.compile(r"(bit\.ly|tinyurl|t\.co|goo\.gl|ow\.ly)/\S+",
re.IGNORECASE)
if url_pattern.search(x.text):
return SPAM # Found a suspicious short URL → SPAM
return ABSTAIN # No suspicious URL → I don't know
# ────────────────────────────────────────────────────────────────────────────
# LF 6: HEURISTIC — combines two signals: urgency words + money mentions
# If BOTH are present → very likely SPAM
# ────────────────────────────────────────────────────────────────────────────
@labeling_function()
def lf_urgency_plus_money(x):
"""
Spam often combines URGENCY ("URGENT", "NOW", "TODAY") with MONEY mentions.
This heuristic LF checks for both signals at once.
If both are present → strong vote for SPAM.
"""
text_lower = x.text.lower()
urgency_words = ["urgent", "now!", "today!", "immediately", "act now"]
money_pattern = re.compile(r"\$\d+|\bfree\b|\bprize\b", re.IGNORECASE)
has_urgency = any(word in text_lower for word in urgency_words)
has_money = bool(money_pattern.search(x.text))
if has_urgency and has_money:
return SPAM # Both signals present → strong SPAM vote
return ABSTAIN # Not both present → I'm not sure
Key lines explained:
@labeling_function()→ The Snorkel decorator that registers your function as an LF. Without this, it's just a regular Python function.x.text→ Accesses the "text" column of the current row (one email).x.text.lower()→ Converts to lowercase so "FREE" matches "free".re.compile(pattern)→ Creates a compiled regex pattern for efficient repeated matching.- Every LF returns
ABSTAINas the default when it doesn't see its specific signal. This is critical!
Step 3 — Apply All LFs to Your Data
We use Snorkel's
PandasLFApplier to run ALL our LFs on ALL our data at once.
This produces the Label Matrix — a table where each row is an email and each column is one LF's vote.
Then we use LFAnalysis to print the report card for all our LFs:
coverage, polarity, overlaps, and conflicts.
# Step 3: Apply all LFs to our dataset and see the Label Matrix
from snorkel.labeling import PandasLFApplier, LFAnalysis
# Put all LFs into a list — the order matters!
lfs = [
lf_spam_keywords,
lf_excessive_caps,
lf_many_exclamations,
lf_professional_language,
lf_suspicious_url,
lf_urgency_plus_money
]
# PandasLFApplier runs every LF from our list on every row of our DataFrame
applier = PandasLFApplier(lfs=lfs)
# Apply! This returns the Label Matrix — shape is (n_emails, n_lfs)
# Each cell is 1 (SPAM), 0 (NOT SPAM), or -1 (ABSTAIN)
L_train = applier.apply(df=df_train)
print("📊 Label Matrix (rows=emails, cols=LFs):")
print(" LF1 LF2 LF3 LF4 LF5 LF6")
print(L_train)
print("\n📋 LF Report Card:")
# LFAnalysis gives you the quality metrics for each LF
# Pass the label matrix and the list of LF names
lf_analysis = LFAnalysis(L=L_train, lfs=lfs).lf_summary()
print(lf_analysis)
Key lines explained:
PandasLFApplier(lfs=lfs)→ Creates the "runner" that applies all LFs. It's smart — it can run LFs in parallel for speed.applier.apply(df=df_train)→ The magic line. Runs every LF on every row and returns the full Label Matrix as a NumPy array.LFAnalysis(L=L_train, lfs=lfs).lf_summary()→ Produces the report card table showing Coverage, Polarity, Overlaps, and Conflicts for each LF.
Step 4 — Train the Label Model (Combine LF Votes)
This is the most important step — the Label Model. It reads the Label Matrix (all the LF votes) and learns which LFs are more trustworthy, which ones tend to conflict, and which ones make systematic errors. It then produces a probabilistic label for each data point — a confidence score like "87% probability this is SPAM." We can also convert these to hard labels (SPAM or NOT SPAM) using
predict().
# Step 4: Train the Label Model to intelligently combine all LF votes
from snorkel.labeling.model import LabelModel
# LabelModel is the "smart referee" that learns which LFs to trust more.
# cardinality=2 means binary classification (SPAM vs NOT SPAM)
# verbose=True prints progress during training
label_model = LabelModel(cardinality=2, verbose=True)
# Fit the Label Model on the Label Matrix
# It learns the reliability and correlations of each LF without seeing true labels!
# n_epochs=200 = number of training rounds for the label model itself
label_model.fit(L_train=L_train, n_epochs=200, seed=42)
print("✅ Label Model trained!")
# ── Get probabilistic labels ──────────────────────────────────────────────────
# predict_proba gives [P(NOT_SPAM), P(SPAM)] for each email
# E.g., [0.08, 0.92] means 92% probability of being SPAM
probas = label_model.predict_proba(L=L_train)
print("\n📊 Probabilistic labels (first 5 emails):")
print(" [P(NOT_SPAM), P(SPAM)]")
for i, prob in enumerate(probas[:5]):
print(f" Email {i+1}: NOT_SPAM={prob[0]:.2f}, SPAM={prob[1]:.2f}")
# ── Get hard (definitive) labels ──────────────────────────────────────────────
# predict() returns the final label: 0 (NOT_SPAM) or 1 (SPAM) for each email
# tie_break_policy="abstain" means: if truly uncertain, output -1 (don't guess)
labels = label_model.predict(L=L_train, tie_break_policy="abstain")
df_train["weak_label"] = labels
print("\n📧 Emails with automatically generated labels:")
for _, row in df_train.iterrows():
label_name = {1: "SPAM", 0: "NOT_SPAM", -1: "ABSTAIN"}[row["weak_label"]]
print(f" Email {row['email_id']}: {label_name:10} | {row['text'][:55]}...")
Key lines explained:
LabelModel(cardinality=2)→ Creates the label model.cardinality=2means two classes (binary). For a 5-class problem, usecardinality=5.label_model.fit(L_train=L_train)→ Trains the label model on the Label Matrix. It learns LF reliability without needing any ground truth labels!predict_proba()→ Returns soft probabilities — more nuanced than just a hard label.predict(tie_break_policy="abstain")→ Returns a definitive label. When tied, returns ABSTAIN instead of guessing randomly.
Step 5 — Train Your Final ML Model on the Weak Labels
Now we use the automatically generated labels to train a real Machine Learning model. Here we use a simple Logistic Regression classifier (from scikit-learn). In production, you would use a more powerful model (BERT, XGBoost, neural network). The key insight: the final model generalises beyond the LFs — it learns patterns from the data itself, not just the LF rules. A classifier trained on 10,000 weakly labelled examples can outperform the LFs that created those labels!
# Step 5: Train a real ML model using the weak labels from Step 4
# In this example, we use a simple TF-IDF text vectoriser + Logistic Regression.
# In production: replace with BERT, RoBERTa, XGBoost, or any powerful model!
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
# ── Filter out ABSTAIN labels before training ──────────────────────────────
# We only want to train on data points where the label model gave a confident label
# Data points where tie_break_policy produced ABSTAIN (-1) are excluded
df_labeled = df_train[df_train["weak_label"] != ABSTAIN].copy()
print(f"✅ Training on {len(df_labeled)} examples with confident labels")
print(f" (Excluded {len(df_train) - len(df_labeled)} ABSTAIN examples)")
# ── Create text features using TF-IDF ──────────────────────────────────────
# TF-IDF converts raw text into numerical features the ML model can understand
# Think of it as: each word gets a score based on how important it is in this email
vectorizer = TfidfVectorizer(ngram_range=(1, 2), max_features=5000)
X_train = vectorizer.fit_transform(df_labeled["text"])
y_train = df_labeled["weak_label"]
# ── Train the Logistic Regression classifier ───────────────────────────────
# LogisticRegression is a simple but powerful text classification algorithm
clf = LogisticRegression(random_state=42, max_iter=1000)
clf.fit(X_train, y_train)
print("\n✅ Classifier trained on weakly labeled data!")
# ── Test on new, unseen emails ──────────────────────────────────────────────
new_emails = [
"Win a free iPhone! Click now to claim your prize!",
"Attached is the slides deck for tomorrow's all-hands meeting.",
"FREE MONEY - limited offer - Act now or lose your chance!",
]
X_new = vectorizer.transform(new_emails)
predictions = clf.predict(X_new)
print("\n🔮 Predictions on new emails:")
for email, pred in zip(new_emails, predictions):
label = "SPAM 🚨" if pred == SPAM else "NOT SPAM ✅"
print(f" {label}: {email[:60]}...")
8. Trend — LLM-Powered Labeling Functions
One of the most exciting developments is using Large Language Models (LLMs) like GPT-4, Llama 3, or Mistral as labeling functions. Instead of writing code, you write a natural language prompt — and the LLM's answer becomes the LF's vote.
This technique, sometimes called Prompted Weak Supervision, was pioneered in a 2022 Stanford paper and has become mainstream . Companies are now using it to label medical records, legal contracts, financial reports — all domains where expert rules are hard to write in code but easy to describe in English!
This shows how to create an LLM-powered Labeling Function. We write a function that takes an email, builds a structured prompt, sends it to an LLM (here using OpenAI's GPT-4o-mini as an example), and maps the LLM's answer back to SPAM / NOT_SPAM / ABSTAIN. This LF can then be added to your LF list alongside your rule-based LFs — the Label Model will automatically learn how much to trust it!
# LLM-Powered Labeling Function
# This LF asks GPT-4o-mini to classify each email and maps the answer to a label.
# You need: pip install openai
import openai
from snorkel.labeling import labeling_function
# Set your API key (in production, use environment variables — never hardcode!)
# openai.api_key = os.environ.get("OPENAI_API_KEY")
client = openai.OpenAI() # reads OPENAI_API_KEY from environment automatically
@labeling_function()
def lf_llm_classifier(x):
"""
Uses GPT-4o-mini as a labeling function.
Sends the email to the LLM with a clear prompt and maps the answer to a label.
This is the way to get powerful labeling without writing complex rules!
"""
prompt = f"""You are a spam email classifier.
Classify the following email as SPAM or NOT_SPAM or UNSURE.
Rules:
- SPAM: unsolicited commercial email, phishing, prize scams, too-good-to-be-true offers
- NOT_SPAM: professional business communication, meeting requests, work discussions
- UNSURE: you are not confident about the classification
Email text:
"{x.text}"
Reply with ONLY one of these three words: SPAM, NOT_SPAM, or UNSURE.
No explanation needed."""
try:
response = client.chat.completions.create(
model="gpt-4o-mini", # Use a fast, cheap model for LF labeling
messages=[{"role": "user", "content": prompt}],
max_tokens=10, # We only need one word back
temperature=0.0 # temperature=0 = deterministic, no randomness
)
# Extract the LLM's answer and strip whitespace
answer = response.choices[0].message.content.strip().upper()
# Map LLM answer → Snorkel label
if answer == "SPAM":
return SPAM
elif answer == "NOT_SPAM":
return NOT_SPAM
else:
return ABSTAIN # "UNSURE" or unexpected answer → abstain
except Exception as e:
# If the API call fails for any reason → abstain (don't crash the whole pipeline)
print(f"LLM LF error: {e}")
return ABSTAIN
# ── Use this LLM LF alongside your rule-based LFs ────────────────────────────
# The Label Model will automatically learn to weight the LLM LF's votes
# against your other rule-based LFs!
# lfs_with_llm = [
# lf_spam_keywords,
# lf_excessive_caps,
# lf_many_exclamations,
# lf_professional_language,
# lf_suspicious_url,
# lf_urgency_plus_money,
# lf_llm_classifier # ← Add the LLM LF just like any other LF!
# ]
Key lines explained:
temperature=0.0→ Makes the LLM deterministic — same input always gives same output. Critical for reproducible labels!max_tokens=10→ We only need one word ("SPAM" or "NOT_SPAM"). Limiting tokens saves time and API cost.try/except → ABSTAIN→ If the LLM API call fails (network error, rate limit), we safely abstain instead of crashing.- The LLM LF is added to the list exactly like any other LF — the Label Model doesn't care whether it's rule-based or LLM-based!
Running an LLM API call for each of 1,000,000 emails could get expensive. The trick: run the LLM LF on a representative sample (e.g., 10,000 emails), then train the Label Model on both the sample LLM votes + the full rule-based LF votes. Or use a small, fast local model (Llama 3.1 8B) — free and runs on your own machine!
9. 🗺️ The Complete LF Pipeline — From Raw Data to Trained Model
Let's put everything together and see the full picture. This is the exact workflow used by ML teams at Google, Stanford Medicine, and Fortune 500 companies:
═══════════════════════════════════════════════════════════════════════
COMPLETE LABELING FUNCTION PIPELINE
═══════════════════════════════════════════════════════════════════════
PHASE 1: DATA COLLECTION
─────────────────────────
Collect large amounts of UNLABELED raw data
(emails, documents, images, medical records...)
No labels needed yet! → df_train (no "y" column)
↓
PHASE 2: EXPLORE & UNDERSTAND DATA (EDA)
─────────────────────────────────────────
Read samples. Find patterns. Talk to domain experts.
"What signals separate class A from class B?"
↓
PHASE 3: WRITE LABELING FUNCTIONS
────────────────────────────────────────────────────────────────
LF Type 1: Keyword LFs (simple word matches)
LF Type 2: Regex LFs (pattern matching)
LF Type 3: Heuristic LFs (if-then rules from experts)
LF Type 4: Distant supervision (external knowledge bases)
LF Type 5: Model-based LFs (old/pretrained models)
LF Type 6: LLM-powered LFs (GPT/Llama prompts) ← New in 2025!
↓
PHASE 4: APPLY LFs → GET LABEL MATRIX
──────────────────────────────────────
PandasLFApplier.apply(df_train)
Result: Label Matrix L_train (n_examples × n_lfs)
Each cell: 1, 0, or -1 (ABSTAIN)
↓
PHASE 5: ANALYSE LF QUALITY
────────────────────────────
LFAnalysis(L_train, lfs).lf_summary()
Check: Coverage, Polarity, Overlaps, Conflicts
Remove or fix LFs with low coverage or high conflict
↓
PHASE 6: TRAIN LABEL MODEL
───────────────────────────
LabelModel.fit(L_train)
Learns: which LFs are trustworthy, which conflict
Outputs: probabilistic labels [P(class0), P(class1)]
↓
PHASE 7: FILTER + GET TRAINING LABELS
──────────────────────────────────────
label_model.predict(L_train, tie_break_policy="abstain")
Filter out ABSTAIN examples
Result: df_labeled with "weak_label" column
↓
PHASE 8: TRAIN FINAL END MODEL
───────────────────────────────
Use df_labeled to train any ML model:
→ BERT, XGBoost, Logistic Regression, Neural Network
The end model GENERALISES beyond the LF rules!
↓
PHASE 9: EVALUATE + ITERATE
─────────────────────────────
Test on a small gold-standard validation set
If accuracy too low → add more LFs, fix existing ones
Repeat from Phase 3
═══════════════════════════════════════════════════════════════════════
10. ✅❌ Best Practices — DOs and DON'Ts
10 LFs that each catch a different pattern are far more powerful than 10 LFs that all check for the same keyword. Diversity reduces the chance of systematic errors — when LFs disagree, the Label Model learns from the disagreement!
This is the most common beginner mistake. An LF with 60% accuracy that has good coverage is much more valuable than an LF with 95% accuracy that only covers 1% of your data. Imperfect is fine — that's the whole point of combining many LFs!
When your LF is not confident, return ABSTAIN. A LF that abstains 70% of the time but is 95% accurate on the 30% it does label is genuinely useful. A LF that never abstains but is only 55% accurate is almost useless (barely better than a coin flip).
Always keep a small set of manually labeled examples (100–500 samples) that no LF ever sees. Use this validation set to measure the true quality of your Label Model and final ML model. Without this, you're flying blind!
A doctor can write a medical LF in 5 minutes using domain knowledge it would take a data scientist months to learn. LFs are the perfect interface between subject matter experts and ML systems. The expert describes the rule in plain English; the developer translates it to 3 lines of Python.
LFs generate training labels. They are NOT the final classifier. A key insight from Snorkel's research: the final ML model trained on weak labels consistently outperforms using the LFs directly as a classifier — because the model generalises beyond the rules!
The real world changes. New types of spam appear. New business terms emerge. LFs need to be updated as data distributions shift. The great news: updating a rule takes minutes, while re-labeling a dataset takes months!
11. 📝 Quick Reference — Everything in One Place
LF Types Cheat Sheet
- Keyword LF → Match a word from a list → fastest to write, good first LF to build
- Regex LF → Match a pattern (URL, phone, date format) → powerful for structured text
- Heuristic LF → if-then rule from domain expert → most flexible, requires expertise
- Distant Supervision LF → lookup in external database/knowledge base → great for named entities
- Model-Based LF → use an old pretrained model as a voter → reuse existing work
- Crowdsourced LF → use click/rating data as weak signal → good for user-facing products
- LLM-Powered LF → prompt GPT/Llama to label → most powerful for fuzzy concepts
Key Numbers to Know
- 🏅 Research result: Teams with Snorkel built models 2.8x faster with 45.5% better performance vs manual labeling
- 📊 Minimum LFs: Start with at least 5–10 LFs. The Label Model needs enough signal to learn from.
- 🎯 Coverage goal: Aim for each LF to cover at least 5–10% of your dataset. Otherwise it's too sparse to be useful.
- ⚖️ Accuracy floor: An LF should be right at least 55–60% of the time when it doesn't abstain. Below that, it might hurt more than help.
Core Snorkel Functions Cheat Sheet
@labeling_function()→ Decorator to register a Python function as a Snorkel LFPandasLFApplier(lfs=lfs).apply(df)→ Runs all LFs on a DataFrame, returns Label MatrixLFAnalysis(L, lfs).lf_summary()→ Prints the quality report card for all LFsLabelModel(cardinality=n)→ Creates the Label Model (n = number of classes)label_model.fit(L_train)→ Trains the Label Model on the Label Matrixlabel_model.predict_proba(L)→ Returns soft probability labels for each data pointlabel_model.predict(L, tie_break_policy="abstain")→ Returns hard labels (0, 1, or -1)
Happy labeling! 🏷️🤖✨
Comments
Post a Comment