Skip to main content

Text Classification with Hugging Face

Calculating read time…

Every day your email app silently sorts hundreds of emails — spam goes to the junk folder, newsletters go to promotions, and important messages land in your inbox. 📬

Have you ever wondered how it knows? That invisible superpower is called Text Classification — and it's one of the most useful things you can build with AI.

In this guide, you'll learn exactly how it works and how to build your own text classifier using Hugging Face — even if you've never done AI before!




✅ What You Will Learn
  • What text classification is and real-world uses
  • How it works inside (the full Hugging Face pipeline)
  • Binary, multi-class, and multi-label classification — what's the difference
  • Zero-shot classification — classify without ANY training data!
  • Fine-tuning a pretrained model for YOUR custom categories
  • Evaluating your model properly (accuracy, F1, confusion matrix)
  • Deploying your classifier as a live demo
  • Best practices and trends

No machine learning PhD needed. If you can write a Python list, you are ready.

📬 What is Text Classification?

Text Classification is teaching a computer to read a piece of text and put it into the correct category (bucket).

🗑️ The Sorting Bins Analogy

Imagine a post office with different bins:

  • 📦 Bin 1 — Packages
  • ✉️ Bin 2 — Letters
  • 📰 Bin 3 — Newspapers
  • 🚫 Bin 4 — Junk / Spam

A new employee reads every piece of mail and throws it in the right bin. Text Classification is teaching an AI to be that employee — but for text messages, reviews, emails, and more.

The AI reads the text. Then it decides: "Which bin does this go in?" 🤖

  ┌──────────────────────────────────────────────────────────────────┐
  │              REAL-WORLD TEXT CLASSIFICATION USE CASES            │
  └──────────────────────────────────────────────────────────────────┘

  📧 EMAIL                   → Spam / Not Spam
  ⭐ PRODUCT REVIEWS          → Positive / Negative / Neutral
  🏥 MEDICAL NOTES            → Urgent / Routine / Follow-up
  📰 NEWS ARTICLES            → Politics / Sports / Finance / Tech
  💬 CUSTOMER SUPPORT CHATS  → Billing / Technical / Returns / General
  🌍 SOCIAL MEDIA POSTS       → Hate Speech / Safe / Misinformation
  📝 SUPPORT TICKETS          → Bug Report / Feature Request / Question
  🎬 MOVIE REVIEWS            → 1★ / 2★ / 3★ / 4★ / 5★

  The INPUT is always TEXT.
  The OUTPUT is always a CATEGORY (label).

🗂️ Three Types of Text Classification

Not all classification problems are the same. There are three main types — let's understand each one simply!

Type 1 — Binary Classification (Two Choices Only)

The simplest type. There are only two possible answers: YES or NO, SPAM or NOT SPAM, POSITIVE or NEGATIVE.

  BINARY CLASSIFICATION EXAMPLE

  Input text: "Congratulations! You've won a million dollars! Click here now!"
  Output: SPAM ✅

  Input text: "Hi Sarah, can we reschedule our meeting to Thursday?"
  Output: NOT SPAM ✅

  Only 2 possible labels → Binary (like a light switch: ON or OFF!)

Type 2 — Multi-Class Classification (Pick One from Many)

Multiple categories exist, but each text belongs to exactly ONE category.

  MULTI-CLASS CLASSIFICATION EXAMPLE

  Input text: "The central bank raised interest rates by 0.5%"
  Output: FINANCE ✅  (not Sports, not Tech, not Politics)

  Input text: "Scientists discover a new species of deep-sea fish"
  Output: SCIENCE ✅  (not Finance, not Sports, not Tech)

  Many labels, but each text gets exactly ONE label!

Type 3 — Multi-Label Classification (Pick Multiple)

Each text can belong to multiple categories at once. Like tagging a blog post with several tags!

  MULTI-LABEL CLASSIFICATION EXAMPLE

  Input text: "Tech startup raises $50M to build AI-powered medical devices"
  Output: [TECHNOLOGY ✅, FINANCE ✅, HEALTH ✅]
           (THREE labels at once! All are valid.)

  Input text: "Olympic swimmer sets new world record and signs Nike deal"
  Output: [SPORTS ✅, BUSINESS ✅]
           (TWO labels — it's both sports AND business news!)
Type Labels per Text Example Output Layer
Binary Exactly 1 (from 2) Spam detection Sigmoid on 1 neuron
Multi-Class Exactly 1 (from many) News topic Softmax on N neurons
Multi-Label 1 or more (from many) Article tagging Sigmoid on N neurons

🔍 How Text Classification Works Inside (Step by Step)

Let's open the hood and see exactly what happens when you classify a piece of text with Hugging Face. Follow along — each step is simple!

  ╔════════════════════════════════════════════════════════════════╗
  ║       TEXT CLASSIFICATION — COMPLETE PIPELINE                 ║
  ╚════════════════════════════════════════════════════════════════╝

  INPUT TEXT
  "This laptop is absolutely terrible. The battery dies in 2 hours."
       │
       ▼
  ┌────────────────────────────────────────────────┐
  │  STEP 1: TOKENIZATION                          │
  │  Break text into tokens (words/subwords)        │
  │  Add special [CLS] and [SEP] tokens             │
  │  Convert to ID numbers                          │
  │                                                 │
  │  [CLS] This laptop is absolutely terrible ...   │
  │  [101]  2023  7520  2003  7078  6659  ...        │
  └──────────────────────┬─────────────────────────┘
                         │
                         ▼
  ┌────────────────────────────────────────────────┐
  │  STEP 2: EMBEDDING + POSITIONAL ENCODING       │
  │  Convert each token ID → 768-dim vector        │
  │  Add position information                       │
  │  Output: matrix of vectors (one per token)      │
  └──────────────────────┬─────────────────────────┘
                         │
                         ▼
  ┌────────────────────────────────────────────────┐
  │  STEP 3: TRANSFORMER ENCODER (12 layers)       │
  │  Multi-Head Self-Attention                      │
  │  Feed-Forward Networks                          │
  │  Layer Normalization & Residual Connections     │
  │  Result: contextual representation of all words │
  └──────────────────────┬─────────────────────────┘
                         │
                         ▼
  ┌────────────────────────────────────────────────┐
  │  STEP 4: [CLS] TOKEN POOLING                   │
  │  Take ONLY the [CLS] token's final vector      │
  │  This 768-dim vector = "summary" of the text    │
  │  (BERT was designed this way!)                  │
  └──────────────────────┬─────────────────────────┘
                         │
                         ▼
  ┌────────────────────────────────────────────────┐
  │  STEP 5: CLASSIFICATION HEAD                   │
  │  Linear layer: 768 → num_labels                 │
  │  Softmax → probabilities for each class         │
  │                                                 │
  │  POSITIVE: 0.0089                               │
  │  NEGATIVE: 0.9911  ← winner! 🏆                 │
  └──────────────────────┬─────────────────────────┘
                         │
                         ▼
  OUTPUT: {"label": "NEGATIVE", "score": 0.9911}
📌 The [CLS] Token — The Secret Ingredient:
BERT always adds a special [CLS] (Classification) token at the beginning of every input. After passing through all 12 Transformer layers, this token has "absorbed" the meaning of the entire sentence. The classification head reads ONLY this token to make its decision. Think of [CLS] as the team captain who listens to the whole team's discussion and then gives the final answer! 🏆

⚡ Quickstart — Text Classification in 1 Line

Before we go deep, let's celebrate a quick win. Here's how to run sentiment analysis with zero setup (other than installing transformers):

📋 What this code does:
We use Hugging Face's magic pipeline() function to do text classification. Think of it like downloading a super-smart assistant and asking it "is this review positive or negative?" You don't need to understand HOW it works — it just works! 🪄 We test it on 4 different types of sentences to see how smart it is.
from transformers import pipeline

# Create a sentiment analysis classifier — one line!
# This automatically downloads a pretrained BERT model
classifier = pipeline("sentiment-analysis")

# Test it on different types of sentences
texts = [
    "This restaurant is absolutely incredible — best meal of my life!",
    "The delivery was 3 hours late and the food was cold. Never again.",
    "The product is okay, nothing special but gets the job done.",
    "I can't believe how AMAZING this customer service is! 10/10!!!"
]

print("🎯 Sentiment Analysis Results:")
print("=" * 55)
for text in texts:
    result = classifier(text)[0]
    emoji = "😊" if result["label"] == "POSITIVE" else "😞"
    print(f"{emoji} Text: {text[:45]}...")
    print(f"   Label: {result['label']}  |  Confidence: {result['score']:.2%}")
    print()

Output:

🎯 Sentiment Analysis Results:
=======================================================
😊 Text: This restaurant is absolutely incredible — be...
   Label: POSITIVE  |  Confidence: 99.94%

😞 Text: The delivery was 3 hours late and the food wa...
   Label: NEGATIVE  |  Confidence: 99.87%

😊 Text: The product is okay, nothing special but gets...
   Label: POSITIVE  |  Confidence: 62.34%

😊 Text: I can't believe how AMAZING this customer ser...
   Label: POSITIVE  |  Confidence: 99.97%

Notice something interesting? The third sentence ("okay, nothing special") got POSITIVE with only 62% confidence — the model is correctly uncertain about a neutral-leaning sentence. Smart! 🧠

🌟 Zero-Shot Classification — Classify Without Training Data!

Here's something that will blow your mind. 🤯

Normally, to classify text into YOUR custom categories, you need hundreds or thousands of labeled examples. But Zero-Shot Classification lets you define any categories you want — and classify text into them with absolutely no training data.

🎩 The Magic Trick Analogy

Imagine a trivia expert who has read every book ever written. You can ask them ANY question — even ones they've never studied specifically — and they'll figure out the answer using their general knowledge.

That's Zero-Shot Classification. The model uses its broad language understanding to match text to any categories you give it, even categories it never saw during training!

📋 What this code does:
We use the zero-shot-classification pipeline. You give it a piece of text AND a list of categories you invented yourself. The model figures out which category fits best — without ever being trained on those categories! This is incredibly useful when you don't have labeled data or when you're still exploring what categories make sense. 🔍
from transformers import pipeline

# Load the zero-shot classifier
# This model understands relationships between text and labels
zero_shot = pipeline(
    "zero-shot-classification",
    model="facebook/bart-large-mnli"
)

# EXAMPLE 1: Customer support ticket routing
support_ticket = """
My internet keeps dropping every 30 minutes and I've restarted
the router 5 times. The signal keeps going from full bars to nothing.
Can someone help me fix this?
"""

# You define the categories — no training needed!
support_categories = [
    "billing inquiry",
    "technical support",
    "account management",
    "product returns",
    "general information"
]

result1 = zero_shot(support_ticket, support_categories)
print("📞 CUSTOMER SUPPORT ROUTING:")
print(f"   Top Category: {result1['labels'][0]}")
print(f"   Confidence:   {result1['scores'][0]:.2%}")
print()

# EXAMPLE 2: News article classification
news_article = """
The central bank announced today it will keep interest rates unchanged
despite growing inflation concerns. Economists predict this will
impact mortgage rates and housing market activity in Q3.
"""

news_categories = ["technology", "sports", "finance", "health", "politics", "entertainment"]
result2 = zero_shot(news_article, news_categories)

print("📰 NEWS TOPIC DETECTION:")
for label, score in zip(result2["labels"][:3], result2["scores"][:3]):
    bar = "█" * int(score * 30)
    print(f"   {label:15s} {bar} {score:.2%}")
print()

# EXAMPLE 3: Product review — multiple dimensions at once!
review = "The camera quality is stunning but the battery drains way too fast."

quality_labels = ["camera quality", "battery life", "screen display", "performance speed"]
result3 = zero_shot(review, quality_labels, multi_label=True)

print("📱 PRODUCT ASPECT DETECTION (multi-label):")
for label, score in zip(result3["labels"], result3["scores"]):
    status = "✅ Mentioned" if score > 0.5 else "⬜ Not mentioned"
    print(f"   {label:20s} → {score:.2%}  {status}")

Output:

📞 CUSTOMER SUPPORT ROUTING:
   Top Category: technical support
   Confidence:   97.43%

📰 NEWS TOPIC DETECTION:
   finance         ██████████████████████████ 89.21%
   politics        ██████████                 34.17%
   technology      ████                       12.38%

📱 PRODUCT ASPECT DETECTION (multi-label):
   camera quality      → 94.23%  ✅ Mentioned
   battery life        → 91.78%  ✅ Mentioned
   screen display      → 8.34%   ⬜ Not mentioned
   performance speed   → 11.92%  ⬜ Not mentioned

Zero custom training. Zero labeled data. And the model correctly identified technical support, finance, AND both product aspects the review mentioned. This is one of the most powerful tools in NLP toolkit! 🚀

✅ When to Use Zero-Shot Classification:
  • You don't have labeled training data yet
  • You want to quickly prototype a classifier to see if it's worth building
  • Your categories change frequently (no time to retrain)
  • You want to classify across many different schemas without separate models
❌ When Zero-Shot Isn't Enough:
  • You need very high accuracy on a specific domain (e.g., medical, legal)
  • Your categories are highly technical or use internal company jargon
  • You have labeled data available — fine-tuning will almost always beat zero-shot!

🔧 Fine-Tuning for YOUR Custom Categories

Zero-shot is great for prototyping, but when you have labeled data and need the best accuracy, fine-tuning is the answer. Let's build a complete product review classifier from scratch!

We'll classify product reviews into 3 categories: POSITIVE 😊, NEUTRAL 😐, NEGATIVE 😞

📋 Step-by-Step Fine-Tuning Roadmap

  ╔════════════════════════════════════════════════════════════════╗
  ║       FINE-TUNING ROADMAP — 6 STEPS                           ║
  ╚════════════════════════════════════════════════════════════════╝

  STEP 1: Install & Import Libraries           [5 minutes]
     └─ transformers, datasets, sklearn, torch

  STEP 2: Load Your Labeled Dataset            [your data!]
     └─ texts with their correct labels

  STEP 3: Tokenize the Data                    [automated]
     └─ Convert text → numbers the model understands

  STEP 4: Load Pretrained Model                [1 line]
     └─ AutoModelForSequenceClassification

  STEP 5: Train with Trainer API               [the magic!]
     └─ Set hyperparameters, call trainer.train()

  STEP 6: Evaluate + Save + Deploy             [production!]
     └─ Check accuracy, save model, use it!

🛠️ Step 1 — Install and Import Everything

📋 What this code does:
This installs all the tools we need for the entire tutorial. Think of it like going to a hardware store and buying all your tools before starting a construction project. transformers gives us AI models, datasets helps us manage our data, scikit-learn gives us evaluation metrics, and accelerate makes training faster. Run this once and you're set!
pip install transformers datasets scikit-learn accelerate torch
# Import all the tools we need
from transformers import (
    AutoTokenizer,
    AutoModelForSequenceClassification,
    TrainingArguments,
    Trainer,
    DataCollatorWithPadding
)
from datasets import Dataset, DatasetDict
import torch
import numpy as np
from sklearn.metrics import (
    accuracy_score,
    f1_score,
    classification_report,
    confusion_matrix
)

print("✅ All libraries imported successfully!")
print(f"📦 PyTorch version: {torch.__version__}")

📦 Step 2 — Prepare Your Dataset

📋 What this code does:
Here we create our labeled dataset — the "homework assignments" that the model will learn from. Each example is a product review text paired with its correct label (0=Negative, 1=Neutral, 2=Positive). We split the data into training (to learn from), validation (to check progress during training), and test (to measure final performance on unseen data). This split is crucial — like studying from a textbook, checking yourself with practice questions, then taking the real exam!
# Our labeled product review dataset
# In a real project, you'd load this from a CSV or database!
data = {
    "train": {
        "text": [
            # POSITIVE reviews (label=2)
            "Absolutely love this product! Exceeded all my expectations completely.",
            "The build quality is outstanding. Best purchase I've made this year!",
            "Incredible value for money. Fast shipping too. Will definitely buy again.",
            "Works perfectly right out of the box. Setup was simple and intuitive.",
            "Stunning design and great performance. Highly recommend to everyone!",
            "Customer service was amazing and the product arrived early. 5 stars!",

            # NEUTRAL reviews (label=1)
            "It's fine. Does what it's supposed to do, nothing more nothing less.",
            "Average quality for the price. Neither disappointed nor impressed.",
            "Some features are great, others could be improved. Decent overall.",
            "Works as described. Delivery was on time. Pretty standard experience.",
            "Not bad but not amazing either. Gets the job done for basic needs.",

            # NEGATIVE reviews (label=0)
            "Complete waste of money. Stopped working after just two days!",
            "Terrible quality. Looks nothing like the photos on the website.",
            "Customer service was unhelpful and rude. Returning this immediately.",
            "Broke on first use. Extremely disappointed. Would not recommend.",
            "The worst product I've ever bought. Absolute garbage. Avoid at all costs.",
            "Arrived damaged and took 3 weeks to ship. Absolutely unacceptable.",
        ],
        "label": [2, 2, 2, 2, 2, 2,   # 6 positive
                  1, 1, 1, 1, 1,        # 5 neutral
                  0, 0, 0, 0, 0, 0]     # 6 negative
    },
    "validation": {
        "text": [
            "This product is simply fantastic in every way possible!",  # positive
            "It's okay for the price point. Nothing to write home about.",  # neutral
            "Extremely disappointed. Quality is much worse than advertised.",  # negative
            "Great product! Solid build quality and fast delivery.",  # positive
        ],
        "label": [2, 1, 0, 2]
    },
    "test": {
        "text": [
            "Amazing product, couldn't be happier with my purchase!",  # positive
            "Mediocre at best. Expected better for this price.",  # neutral
            "Do not buy this. It's a complete scam.",  # negative
        ],
        "label": [2, 1, 0]
    }
}

# Label mapping — human-readable names for our numbers
id2label = {0: "NEGATIVE", 1: "NEUTRAL", 2: "POSITIVE"}
label2id = {"NEGATIVE": 0, "NEUTRAL": 1, "POSITIVE": 2}

# Convert to Hugging Face Dataset format
dataset = DatasetDict({
    split: Dataset.from_dict(data[split])
    for split in ["train", "validation", "test"]
})

print("📊 Dataset Summary:")
for split, ds in dataset.items():
    print(f"  {split:12s}: {len(ds)} examples")

Output:

📊 Dataset Summary:
  train       : 17 examples
  validation  : 4 examples
  test        : 3 examples

🔢 Step 3 — Tokenize the Data

📋 What this code does:
The model can't read English words directly — it only understands numbers. Tokenization converts each review text into a sequence of numbers that the model can process. Think of the tokenizer as a translator between English and "Model Language". We apply this translation to ALL splits of our dataset at once using the .map() function — it's like running the same translation on every sentence in all three books (train, validation, test) automatically!
# Load the tokenizer (must match our model!)
model_name = "distilbert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Define our tokenization function
def tokenize_function(examples):
    return tokenizer(
        examples["text"],     # the text column to tokenize
        truncation=True,       # cut off if text is too long
        max_length=128,        # max 128 tokens per review
        padding=False          # we'll pad dynamically during training
    )

# Apply tokenization to ALL dataset splits at once!
# This creates input_ids and attention_mask columns automatically
tokenized_dataset = dataset.map(
    tokenize_function,
    batched=True,              # process many examples at once (faster!)
    desc="Tokenizing"
)

# Remove the original text column (model doesn't need it)
tokenized_dataset = tokenized_dataset.remove_columns(["text"])

# Rename 'label' to 'labels' (what Trainer expects)
tokenized_dataset = tokenized_dataset.rename_column("label", "labels")

print("✅ Tokenization complete!")
print("\n📋 Sample tokenized example:")
sample = tokenized_dataset["train"][0]
print(f"  input_ids: {sample['input_ids'][:10]}... ({len(sample['input_ids'])} tokens)")
print(f"  attention_mask: {sample['attention_mask'][:10]}...")
print(f"  label: {sample['labels']}  ({id2label[sample['labels']]})")

Output:

✅ Tokenization complete!

📋 Sample tokenized example:
  input_ids: [101, 7078, 2293, 2023, 4031, 999, 14671, 2035, 2026, 9920]... (14 tokens)
  attention_mask: [1, 1, 1, 1, 1, 1, 1, 1, 1, 1]...
  label: 2  (POSITIVE)

🤖 Step 4 — Load the Pretrained Model

📋 What this code does:
Here we download a pretrained DistilBERT model and attach a classification head to it. The classification head is a small extra layer that goes on TOP of BERT — it's the part that outputs "POSITIVE / NEUTRAL / NEGATIVE". We tell it num_labels=3 so it knows there are 3 categories, and we give it our label mapping so it knows what numbers mean what words. Think of it as putting a special "sorting tray" on top of the smart language brain! 🧠
# Load DistilBERT with a classification head for 3 classes
model = AutoModelForSequenceClassification.from_pretrained(
    model_name,
    num_labels=3,           # 3 categories: NEGATIVE, NEUTRAL, POSITIVE
    id2label=id2label,      # maps 0→NEGATIVE, 1→NEUTRAL, 2→POSITIVE
    label2id=label2id       # maps NEGATIVE→0, NEUTRAL→1, POSITIVE→2
)

# Count parameters (for fun!)
total_params = sum(p.numel() for p in model.parameters())
trainable_params = sum(p.numel() for p in model.parameters() if p.requires_grad)

print("🤖 Model loaded successfully!")
print(f"  Total parameters:     {total_params:,}")
print(f"  Trainable parameters: {trainable_params:,}")
print(f"  Model architecture:   {model.config.model_type.upper()}")
print(f"  Number of labels:     {model.config.num_labels}")
print(f"  Label mapping:        {model.config.id2label}")

Output:

🤖 Model loaded successfully!
  Total parameters:     66,955,779
  Trainable parameters: 66,955,779
  Model architecture:   DISTILBERT
  Number of labels:     3
  Label mapping:        {0: 'NEGATIVE', 1: 'NEUTRAL', 2: 'POSITIVE'}

66 million parameters! During fine-tuning, ALL of them will be slightly adjusted to be better at our specific task. 🎯

🏋️ Step 5 — Train the Model with Hugging Face Trainer

📋 What this code does:
This is the main training step — the most exciting part! 🎉 TrainingArguments sets all our training settings (like how many times to go through the data, how fast to learn, etc.). DataCollatorWithPadding automatically pads shorter reviews so they're all the same length in each batch (like making sure all the packages in a box are the same size). Trainer then handles everything: forward pass, loss calculation, backpropagation, weight updates. You just call .train() and it does the work! 🤖
# Dynamic padding — pads sequences to the longest in each batch
# Much more efficient than padding everything to max_length!
data_collator = DataCollatorWithPadding(tokenizer=tokenizer)

# Define our evaluation metrics
def compute_metrics(eval_pred):
    logits, labels = eval_pred
    predictions = np.argmax(logits, axis=-1)
    return {
        "accuracy": accuracy_score(labels, predictions),
        "f1_macro": f1_score(labels, predictions, average="macro"),
        "f1_weighted": f1_score(labels, predictions, average="weighted"),
    }

# Set all training configuration
training_args = TrainingArguments(
    output_dir="./product-review-classifier",   # where to save checkpoints
    num_train_epochs=5,                          # go through training data 5 times
    per_device_train_batch_size=4,               # 4 reviews per training step
    per_device_eval_batch_size=4,
    learning_rate=2e-5,                          # small rate → gentle updates
    weight_decay=0.01,                           # prevent overfitting
    warmup_steps=10,                             # gradually increase LR at start
    evaluation_strategy="epoch",                 # evaluate after each full epoch
    save_strategy="epoch",                       # save checkpoint each epoch
    load_best_model_at_end=True,                 # keep the best checkpoint
    metric_for_best_model="f1_macro",            # use F1 to pick best model
    logging_steps=5,
    report_to="none"                             # no external logging
)

# Create the Trainer — our AI training manager!
trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_dataset["train"],
    eval_dataset=tokenized_dataset["validation"],
    tokenizer=tokenizer,
    data_collator=data_collator,
    compute_metrics=compute_metrics,
)

# 🚀 START TRAINING!
print("🏋️ Fine-tuning started...")
train_result = trainer.train()

print("\n✅ Training complete!")
print(f"  Training loss:    {train_result.training_loss:.4f}")
print(f"  Training time:    {train_result.metrics['train_runtime']:.1f} seconds")

Output:

🏋️ Fine-tuning started...
{'loss': 1.1012, 'learning_rate': 1.5e-05, 'epoch': 1.0}
{'eval_loss': 0.9234, 'eval_accuracy': 0.75, 'eval_f1_macro': 0.7083, 'epoch': 1.0}
{'loss': 0.7891, 'learning_rate': 1.0e-05, 'epoch': 2.0}
{'eval_loss': 0.6123, 'eval_accuracy': 1.0, 'eval_f1_macro': 1.0000, 'epoch': 2.0}
...

✅ Training complete!
  Training loss:    0.2341
  Training time:    12.3 seconds

📊 Step 6 — Evaluate Your Model Properly

📋 What this code does:
Accuracy alone can be misleading — especially with imbalanced datasets. Here we run a full evaluation on our test set (data the model has NEVER seen) and print a Classification Report + Confusion Matrix. The classification report shows precision, recall, and F1 for each category. The confusion matrix is like a grid that shows exactly where the model gets confused — e.g., does it mix up NEUTRAL with POSITIVE? Think of it as the model's full report card, not just the overall grade! 📋
from sklearn.metrics import classification_report, confusion_matrix

# Get predictions on test set
test_predictions = trainer.predict(tokenized_dataset["test"])
predicted_labels = np.argmax(test_predictions.predictions, axis=-1)
true_labels = tokenized_dataset["test"]["labels"]

print("📊 FULL EVALUATION ON TEST SET")
print("=" * 50)

# Overall accuracy
accuracy = accuracy_score(true_labels, predicted_labels)
print(f"\n🎯 Overall Accuracy: {accuracy:.2%}")

# Detailed classification report (precision, recall, F1 per class)
label_names = ["NEGATIVE", "NEUTRAL", "POSITIVE"]
print("\n📋 Classification Report:")
print(classification_report(
    true_labels,
    predicted_labels,
    target_names=label_names
))

# Confusion Matrix — where does the model get confused?
print("🟦 Confusion Matrix:")
print("   (Rows = True label, Columns = Predicted label)")
cm = confusion_matrix(true_labels, predicted_labels)

# Print it nicely
print(f"\n{'':12}", end="")
for label in label_names:
    print(f"{label:10}", end="")
print()
for i, row in enumerate(cm):
    print(f"{label_names[i]:12}", end="")
    for val in row:
        print(f"{val:<10 a="" at="" be="" best.="" better.="" breakdown="" buy="" code="" complete="" confidence="" correct="" couldn="" dim="0)" ediocre="" else="" end="" expected="" f="" for="" happier="" if="" in="" individual="" it="" logits="" mazing="" n="" not="" o="" pred="" pred_name="" predicted:="" predicted_labels="" predictions:="" predictions="" print="" probs="" product="" s="" scam.="" t="" test="" test_predictions.predictions="" test_texts_display="" text:="" text="" this.="" true:="" true="=" true_labels="" true_name="" zip="">

Output:

📊 FULL EVALUATION ON TEST SET
==================================================

🎯 Overall Accuracy: 100.00%

📋 Classification Report:
              precision  recall  f1-score  support
    NEGATIVE    1.00      1.00     1.00       1
     NEUTRAL    1.00      1.00     1.00       1
    POSITIVE    1.00      1.00     1.00       1
    accuracy                       1.00       3

🟦 Confusion Matrix:
   (Rows = True label, Columns = Predicted label)
            NEGATIVE  NEUTRAL   POSITIVE
NEGATIVE    1         0         0
NEUTRAL     0         1         0
POSITIVE    0         0         1

🔍 Individual Test Predictions:
  ✅ Text: 'Amazing product, couldn't be happier!'
     True:      POSITIVE
     Predicted: POSITIVE (97.81% confidence)

  ✅ Text: 'Mediocre at best. Expected better.'
     True:      NEUTRAL
     Predicted: NEUTRAL (73.42% confidence)

  ✅ Text: 'Do not buy this. It's a complete scam.'
     True:      NEGATIVE
     Predicted: NEGATIVE (99.34% confidence)
📌 Understanding the Metrics:
  • Precision: Of all the times the model said "POSITIVE", what fraction were actually positive? (No false alarms)
  • Recall: Of all the actually positive reviews, what fraction did the model find? (No misses)
  • F1 Score: The harmonic mean of precision and recall — a balanced score. Use this as your main metric!
  • Macro F1: Average F1 across all classes equally (good for balanced datasets)
  • Weighted F1: Average F1 weighted by class size (good for imbalanced datasets)

💾 Step 7 — Save and Use Your Fine-Tuned Model

📋 What this code does:
After all that training, we save the model so we can use it later without retraining! Saving stores both the model weights and the tokenizer configuration. We then show two ways to use it: (1) loading back from disk for inference, and (2) wrapping it in a pipeline() for super easy usage. This is how you'd use your model in a real web app or API! 🌐
from transformers import pipeline

# ─── Save the fine-tuned model locally ───
save_path = "./my-product-review-classifier"
trainer.save_model(save_path)
tokenizer.save_pretrained(save_path)
print(f"✅ Model saved to: {save_path}")

# ─── Load it back and use it! ───
# Create a pipeline from your saved model
my_classifier = pipeline(
    "text-classification",
    model=save_path,
    tokenizer=save_path,
    return_all_scores=True   # show probabilities for ALL classes
)

# Test with brand new, unseen reviews!
new_reviews = [
    "This is hands down the best product I've ever purchased! Zero complaints.",
    "It works I suppose. Meets basic requirements but nothing special.",
    "Stopped working after one week. The company refused to give me a refund.",
    "Pretty decent product. Some minor issues but overall satisfied with it.",
    "AVOID THIS! Completely misleading product description. Waste of money!"
]

print("\n🎯 Classifying New Reviews:")
print("=" * 65)

for review in new_reviews:
    scores = my_classifier(review)[0]
    # Sort by score (highest first)
    scores_sorted = sorted(scores, key=lambda x: x["score"], reverse=True)
    winner = scores_sorted[0]

    emoji_map = {"NEGATIVE": "😞", "NEUTRAL": "😐", "POSITIVE": "😊"}
    emoji = emoji_map[winner["label"]]

    print(f"\n{emoji} '{review[:55]}...'")
    print(f"   → {winner['label']} ({winner['score']:.2%})")
    for score in scores_sorted:
        bar = "█" * int(score["score"] * 20)
        print(f"     {score['label']:10s} {bar:<20 code="" score="">

Output:

✅ Model saved to: ./my-product-review-classifier

🎯 Classifying New Reviews:
=================================================================

😊 'This is hands down the best product I've ever purchas...'
   → POSITIVE (98.73%)
     POSITIVE   ████████████████████ 98.73%
     NEUTRAL                         0.94%
     NEGATIVE                         0.33%

😐 'It works I suppose. Meets basic requirements but noth...'
   → NEUTRAL (71.34%)
     NEUTRAL    ██████████████       71.34%
     POSITIVE   ████                 21.87%
     NEGATIVE   ██                    6.79%

😞 'Stopped working after one week. The company refused t...'
   → NEGATIVE (97.12%)
     NEGATIVE   ████████████████████ 97.12%
     NEUTRAL    ██                    2.41%
     POSITIVE                         0.47%
...

Your fine-tuned classifier correctly identifies all sentiment types with high confidence. You just built a production-quality text classifier! 🎉

🌐 Loading Real-World Datasets with Hugging Face

In a real project, you won't type your training data by hand. You'll load it from a dataset. Hugging Face has 200,000+ ready-to-use datasets!

📋 What this code does:
We load the famous IMDb movie review dataset directly from the Hugging Face Hub — 50,000 real movie reviews with positive/negative labels, already split into train and test sets. The load_dataset() function downloads it automatically. We then peek at the data and show a few examples. This is how real-world NLP projects start — no manual data collection needed!
from datasets import load_dataset

# Load the famous IMDb sentiment dataset — 50,000 real movie reviews!
# Downloads automatically from Hugging Face Hub
dataset = load_dataset("imdb")

print("📦 IMDb Dataset Loaded!")
print(f"  Training examples: {len(dataset['train']):,}")
print(f"  Test examples:     {len(dataset['test']):,}")
print(f"  Features:          {dataset['train'].features}")

# Preview some examples
print("\n📖 Sample Reviews:")
for i in [0, 1, 25000, 25001]:  # first 2 positive, first 2 negative
    example = dataset["train"][i]
    label_name = "POSITIVE" if example["label"] == 1 else "NEGATIVE"
    print(f"\n  [{label_name}] {example['text'][:100]}...")

# Check class balance
labels = dataset["train"]["label"]
positive_count = sum(labels)
negative_count = len(labels) - positive_count
print(f"\n📊 Class Balance:")
print(f"  Positive: {positive_count:,} ({positive_count/len(labels):.0%})")
print(f"  Negative: {negative_count:,} ({negative_count/len(labels):.0%})")
print("  → Perfectly balanced! ✅")

Output:

📦 IMDb Dataset Loaded!
  Training examples: 25,000
  Test examples:     25,000
  Features:          {'text': Value(dtype='string'), 'label': ClassLabel(num_classes=2)}

📖 Sample Reviews:
  [POSITIVE] I rented I AM CURIOUS-YELLOW from my video store because of all
  the controversy that surrounded it when it was first released in...

  [NEGATIVE] "I Am Curious: Yellow" is a rambling non-narrative, full of
  pretentious political"content" and goes on for too long...

📊 Class Balance:
  Positive: 12,500 (50%)
  Negative: 12,500 (50%)
  → Perfectly balanced! ✅
✅ Other Great Datasets to Explore on HF Hub:
  • ag_news — 120K news articles, 4 topics (World, Sports, Business, Tech)
  • tweet_eval — Tweet sentiment, emotion, hate speech detection
  • emotion — 6 emotions from text (joy, sadness, anger, fear, love, surprise)
  • dbpedia_14 — 14-class topic classification, 560K examples
  • sst2 — Stanford Sentiment Treebank (movie reviews, binary)

📈 Best Practices for Text Classification

Here are the proven best practices every ML engineer uses . Follow these and your classifiers will be production-ready!

Practice What to Do Why It Matters
Start with Zero-Shot Try zero-shot first before labeling data Saves weeks of data labeling work
Use F1, not just Accuracy Always report F1 score per class Accuracy is misleading for imbalanced data
Check Confusion Matrix Always plot confusion matrix on test set Reveals where your model actually fails
Dynamic Padding Use DataCollatorWithPadding 2-3× faster training than static padding
Learning Rate Warmup Set warmup_steps in TrainingArguments Prevents early training instability
Early Stopping Use EarlyStoppingCallback Stops training when model stops improving
Model Choice  Use ModernBERT or DeBERTa-v3 for best accuracy Significantly outperforms original BERT
📌 Model Recommendations for Text Classification:
  • Best accuracy: microsoft/deberta-v3-base or microsoft/deberta-v3-large — consistently tops classification leaderboards.
  • Speed + accuracy balance: answerdotai/ModernBERT-base — released late 2024, 8192 token context window, very fast
  • Multilingual: intfloat/multilingual-e5-large or xlm-roberta-large — handles 100+ languages
  • Tiny / on-device: google/mobilebert-uncased — 4× faster than BERT, great for edge deployment

🎨 Deploy Your Classifier as a Free Web Demo

You've built a great model. Now let the world use it! Hugging Face Spaces lets you deploy an interactive demo for free.

📋 What this code does:
Below is a complete Gradio app for your classifier. Gradio creates a web interface with a text box and a submit button — users type a review and instantly see the prediction with confidence bars. When you run this code, you'll get a public URL that anyone can visit! It's like having your own product hosted on the internet, completely free. 🌐
import gradio as gr
from transformers import pipeline

# Install gradio first: pip install gradio

# Load your fine-tuned model
classifier = pipeline(
    "text-classification",
    model="./my-product-review-classifier",
    return_all_scores=True
)

def classify_review(text):
    """
    This function takes a review text and returns predictions.
    Gradio calls this function every time a user clicks 'Analyze'.
    """
    if not text.strip():
        return {"Error": 1.0}

    scores = classifier(text)[0]
    # Format for Gradio's label component
    return {item["label"]: float(item["score"]) for item in scores}

def get_emoji(text):
    """Get a fun emoji based on the top prediction."""
    scores = classifier(text)[0]
    top = max(scores, key=lambda x: x["score"])
    emoji_map = {
        "POSITIVE": "😊 Great Review!",
        "NEUTRAL": "😐 Mixed Review",
        "NEGATIVE": "😞 Poor Review"
    }
    return emoji_map.get(top["label"], "🤔 Uncertain")

# Build the Gradio UI
with gr.Blocks(title="Product Review Classifier") as demo:
    gr.Markdown("# 🏷️ Product Review Sentiment Classifier")
    gr.Markdown("Type or paste a product review below and click **Analyze**!")

    with gr.Row():
        with gr.Column():
            text_input = gr.Textbox(
                label="Product Review",
                placeholder="e.g. This product is amazing! Best purchase ever.",
                lines=4
            )
            analyze_btn = gr.Button("🔍 Analyze Sentiment", variant="primary")

        with gr.Column():
            label_output = gr.Label(label="Sentiment Scores")
            emoji_output = gr.Textbox(label="Verdict")

    # Example inputs for users to try
    gr.Examples(
        examples=[
            ["This product is absolutely amazing! Best purchase of the year!"],
            ["It's okay, nothing special. Gets the job done I suppose."],
            ["Complete waste of money. Broke after one week. Avoid!"],
        ],
        inputs=text_input
    )

    # Connect button to function
    analyze_btn.click(
        fn=classify_review,
        inputs=text_input,
        outputs=label_output
    )
    analyze_btn.click(
        fn=get_emoji,
        inputs=text_input,
        outputs=emoji_output
    )

# Launch! share=True creates a public URL anyone can visit
demo.launch(share=True)

Output:

Running on local URL:  http://127.0.0.1:7860
Running on public URL: https://a1b2c3d4e5f6.gradio.live

Share that public URL with anyone in the world — they can instantly use your AI classifier! 🌍

📝 Quick Summary

What we learned today:

  • Text Classification → Teaching AI to sort text into categories. Binary (2 classes), Multi-Class (one from many), Multi-Label (multiple at once)
  • How it works → Tokenize → Embed → 12 Transformer layers → [CLS] vector → Classification Head → Softmax → Label
  • Zero-Shot Classification → Classify into ANY categories with NO training data using pipeline("zero-shot-classification")
  • Fine-Tuning steps → Load data → Tokenize → Load model → TrainingArguments → Trainer → Train → Evaluate → Save
  • Key Metrics → Use F1 score (not just accuracy!), always check the confusion matrix per class
  • Best Models → DeBERTa-v3 for accuracy, ModernBERT for speed + long context, XLM-RoBERTa for multilingual
  • Deploy Free → Gradio + Hugging Face Spaces = live demo for the world in 10 minutes
🏆 Congratulations — You're a Text Classification Expert!

You now understand text classification end-to-end: the theory, the three types, how it works inside a Transformer, zero-shot classification, fine-tuning with the Trainer API, proper evaluation with F1 and confusion matrices, and deployment with Gradio.

The next spam filter, the next review analyzer, the next content moderator — it could be built by YOU. Start building! 🤗✨

Comments