Skip to main content

Hugging Face Transformers Pipeline

Calculating read time…

What if you could add AI superpowers to your app — without a PhD in machine learning? No training. No math. No GPU setup. Just three lines of code.

That's exactly what the Transformers pipeline() function does. It's the single most powerful shortcut in all of modern AI development.




✅ What You'll Learn in This Blog:
  • What a pipeline is and why it exists
  • Every major pipeline task with working code examples
  • How to customize pipelines with any model from the Hub
  • Batch processing, device selection, and performance tips
  • Advanced patterns: chaining pipelines, custom pipelines, and async inference
  • Best practices for using pipelines in real production apps

🍕 Section 1 — What Exactly IS a Pipeline?

Let's start with an analogy you'll never forget.

🏭 Think of a Pizza Factory

When you order a pizza, you don't knead the dough, slice the vegetables, operate the oven, and cut it yourself. You just say "I want a Margherita" — and the factory handles the rest.

A Hugging Face pipeline is exactly that factory for AI. You say what task you want (e.g., "analyze sentiment"), and the pipeline handles ALL the boring steps for you:

⚙️ What Happens Inside a Pipeline

📝
Your Text
"I love pizza!"
→
🔪
Tokenizer
Chop text to numbers
→
🧠
Model
AI thinks & processes
→
🔄
Post-process
Numbers → human answer
→
✅
Result
POSITIVE 😊

↑ The pipeline runs ALL these steps automatically when you call it with one line of code.

💡 Before Pipelines Existed...
Developers had to write 50–100 lines of code just to run a single model: load tokenizer, tokenize text, convert to tensors, run model, apply softmax, decode output, handle errors...
Now? 3 lines. That's the pipeline revolution.

⚙️ Section 2 — Setup: Get Ready in 2 Minutes

📋 What This Does:
Installs the Hugging Face transformers library and torch (the AI engine that powers the models). Run this once — you're set for everything in this blog.
# Install Hugging Face libraries
pip install transformers torch

# Optional but recommended: faster downloads + audio support
pip install accelerate datasets torchaudio

Verify it works:

📋 What This Does:
Loads the pipeline function and runs a quick hello-world test. If you see a result (not an error), everything is installed correctly!
from transformers import pipeline

# The world's simplest AI program
classifier = pipeline("sentiment-analysis")
print(classifier("I absolutely love learning AI!"))

Expected Output:

[{'label': 'POSITIVE', 'score': 0.9998689889907837}]

🎉 That's it. You just ran an AI model. Let's go deeper.

✅ Tip: Use Google Colab for FREE GPU!
Go to colab.research.google.com, click Runtime → Change runtime type → GPU. Transformers pipelines run 10–50× faster on GPU. It's completely free!

🔬 Section 3 — Anatomy of the Pipeline Function

The pipeline() function has several important parameters. Let's break down exactly what each one does.

pipeline(
    task,            ← WHAT to do (e.g., "sentiment-analysis")
    model=None,      ← WHICH model to use (optional — auto-picked if omitted)
    tokenizer=None,  ← HOW to chop text (usually auto-matched to model)
    device=None,     ← WHERE to run: -1=CPU, 0=GPU, "mps"=Apple Silicon
    batch_size=1,    ← HOW MANY examples to process at once
    **kwargs         ← Extra task-specific options
)
📋 What This Code Does:
Shows four different ways to create the same pipeline — from the simplest (auto everything) to the most explicit (specify every detail). Think of it like ordering coffee: "coffee" vs "double-shot oat milk latte at 65°C".
from transformers import pipeline

# ── Level 1: Simplest — auto picks everything ────────────────
pipe = pipeline("sentiment-analysis")

# ── Level 2: Specify a model ─────────────────────────────────
pipe = pipeline(
    "sentiment-analysis",
    model="distilbert-base-uncased-finetuned-sst-2-english"
)

# ── Level 3: Specify model + device ──────────────────────────
pipe = pipeline(
    "sentiment-analysis",
    model="distilbert-base-uncased-finetuned-sst-2-english",
    device=0  # Use GPU (device=0 = first GPU). Use -1 for CPU.
)

# ── Level 4: Maximum control ─────────────────────────────────
from transformers import AutoTokenizer, AutoModelForSequenceClassification

model_name = "distilbert-base-uncased-finetuned-sst-2-english"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

pipe = pipeline(
    "sentiment-analysis",
    model=model,
    tokenizer=tokenizer,
    device=0,
    batch_size=8      # Process 8 examples at once for speed
)

# All four produce the same result:
print(pipe("Transformers pipelines are incredible!"))

🎯 Section 4 — Every Pipeline Task, Explained (Complete List)

This is the big section. We'll go through every major pipeline task with a real code example. Bookmark this — you'll come back often!

📊 Task 1: Sentiment Analysis

Real-world use: Automatically analyze product reviews, social media posts, customer feedback, or any text to detect positive/negative/neutral feelings.

📋 What This Code Does:
Feeds a list of texts to the AI and gets back two things for each: a label (POSITIVE or NEGATIVE) and a score (0 to 1 — how confident the AI is). Higher score = more confident. Perfect for processing customer reviews at scale!
from transformers import pipeline

pipe = pipeline("sentiment-analysis")

# You can pass a single string OR a list of strings
texts = [
    "This laptop is absolutely fantastic. Best purchase ever!",
    "The customer support was terrible. Waited 3 days for a response.",
    "It's a product. Does what it says. Nothing special.",
    "Oh wow, incredible quality! Highly recommend to everyone!",
]

results = pipe(texts)

print("📊 Sentiment Analysis Results:")
print("-" * 55)
for text, result in zip(texts, results):
    emoji = "😊" if result['label'] == 'POSITIVE' else "😞"
    bar_len = int(result['score'] * 20)
    bar = "█" * bar_len + "░" * (20 - bar_len)
    print(f"\n{emoji} {result['label']} [{result['score']:.1%}]")
    print(f"   [{bar}]")
    print(f"   \"{text[:50]}...\"")

Output:

📊 Sentiment Analysis Results:
-------------------------------------------------------

😊 POSITIVE [99.9%]
   [████████████████████]
   "This laptop is absolutely fantastic. Best purchas..."

😞 NEGATIVE [99.8%]
   [████████████████████]
   "The customer support was terrible. Waited 3 days ..."

😊 POSITIVE [56.2%]
   [███████████░░░░░░░░░]
   "It's a product. Does what it says. Nothing specia..."

😊 POSITIVE [99.7%]
   [████████████████████]
   "Oh wow, incredible quality! Highly recommend to ev..."
💡 Notice the "Nothing special" result!
It got 56% POSITIVE — the AI is uncertain, which makes sense for neutral text. When score is close to 50%, the text is genuinely ambiguous. You can filter these out in production: only trust results above 80%.

✍️ Task 2: Text Generation

Real-world use: Auto-complete sentences, generate product descriptions, write blog post drafts, create story continuations, power chatbots.

📋 What This Code Does:
You give the AI a starting phrase (the "prompt"), and it continues writing from there. It's like the AI is finishing your sentence — except it can write entire paragraphs.

The max_new_tokens controls how long the response is. num_return_sequences generates multiple different versions — great for getting variety and picking the best one.
from transformers import pipeline

# GPT-2 is classic and runs even on CPU
# For production, use: "Qwen/Qwen2.5-1.5B-Instruct" (much better!)
generator = pipeline("text-generation", model="gpt2")

prompt = "The future of artificial intelligence in healthcare is"

# Generate 3 different continuations of the same prompt
outputs = generator(
    prompt,
    max_new_tokens=60,       # Generate up to 60 new tokens (≈ 45 words)
    num_return_sequences=3,  # Give us 3 different versions
    temperature=0.8,         # 0.0 = repetitive/safe, 1.0 = creative/random
    do_sample=True,          # Enable random sampling (needed for temperature)
    pad_token_id=50256       # Prevents a warning message
)

print(f"📝 Prompt: \"{prompt}\"\n")
print("=" * 60)
for i, output in enumerate(outputs, 1):
    # Remove the prompt from the output (only show new text)
    generated_text = output['generated_text'][len(prompt):]
    print(f"\n📌 Version {i}:")
    print(f"   ...{generated_text.strip()}")

Output:

📝 Prompt: "The future of artificial intelligence in healthcare is"

============================================================

📌 Version 1:
   ...bright, with early-detection systems reducing diagnosis
   time by 90%. Hospitals worldwide are testing models that
   predict patient deterioration before symptoms appear.

📌 Version 2:
   ...already here — AI models read X-rays faster than radiologists
   in several studies, and drug discovery timelines have shrunk
   from 12 years to under 3.

📌 Version 3:
   ...a subject of both excitement and concern. Clinicians worry
   about accountability when an algorithm makes a wrong call.
✅ Upgrade: Use Chat Models for Better Results
GPT-2 is old. For modern text generation, use instruction-tuned models:
  • "microsoft/Phi-3.5-mini-instruct" — Fast, great quality
  • "Qwen/Qwen2.5-1.5B-Instruct" — Excellent for lightweight apps
  • "google/gemma-2-2b-it" — Google's efficient chat model

📰 Task 3: Summarization

Real-world use: Summarize long articles, research papers, meeting transcripts, legal documents, or any long text into a short, clear summary.

📋 What This Code Does:
Takes a long piece of text and squeezes it into a shorter summary — like a smart student who reads a whole chapter and gives you the key points.

min_length = minimum words in summary. max_length = maximum words. do_sample=False = deterministic (always same result — good for summaries).
from transformers import pipeline

# BART is still the gold standard for summarization
summarizer = pipeline("summarization", model="facebook/bart-large-cnn")

long_article = """
Climate scientists have published new research indicating that global
average temperatures have risen by 1.2 degrees Celsius above pre-industrial
levels as of early 2026. The study, which analyzed data from over 12,000
weather stations worldwide alongside satellite readings, found that the rate
of warming has accelerated compared to the previous decade. Researchers
highlight that Arctic regions are warming four times faster than the global
average, leading to rapid ice sheet loss in Greenland and parts of Antarctica.

The consequences include rising sea levels — currently at 4.5mm per year —
which threaten coastal cities like Miami, Jakarta, and Amsterdam. Extreme
weather events, including Category 5 hurricanes and unprecedented heatwaves,
have become significantly more frequent. The researchers urge governments to
accelerate the transition to renewable energy and implement carbon capture
technologies at a massive scale.

Despite growing public awareness, global carbon emissions hit a record high
in 2025, driven by industrial growth in developing economies. Scientists warn
that without immediate, dramatic action, the 1.5°C threshold agreed in the
Paris Agreement could be crossed within 8 years.
"""

summary = summarizer(
    long_article,
    max_length=80,    # Summary is at most 80 tokens long
    min_length=30,    # Summary is at least 30 tokens long
    do_sample=False   # Consistent, deterministic output
)

original_words = len(long_article.split())
summary_words = len(summary[0]['summary_text'].split())
reduction = (1 - summary_words / original_words) * 100

print("📰 ORIGINAL ARTICLE:")
print(f"   Word count: {original_words} words")
print()
print("📝 AI SUMMARY:")
print(f"   {summary[0]['summary_text']}")
print()
print(f"✅ Compression: {original_words} → {summary_words} words ({reduction:.0f}% shorter!)")

Output:

📰 ORIGINAL ARTICLE:
   Word count: 187 words

📝 AI SUMMARY:
   Global average temperatures have risen by 1.2 degrees Celsius above
   pre-industrial levels. Arctic regions are warming four times faster than
   the global average. Sea levels are rising at 4.5mm per year, threatening
   coastal cities. Scientists warn the 1.5°C threshold could be crossed within
   8 years without immediate action.

✅ Compression: 187 → 52 words (72% shorter!)

🌍 Task 4: Translation

Real-world use: Translate customer emails, documents, app content, or any text between 100+ language pairs.

📋 What This Code Does:
Translates text from one language to another using a model trained specifically for that language pair. The model name tells you the direction: Helsinki-NLP/opus-mt-en-hi means English → Hindi. You need to find the right model for each language pair on the Hub.
from transformers import pipeline

# ── English → French ────────────────────────────────────────
translator_fr = pipeline(
    "translation",
    model="Helsinki-NLP/opus-mt-en-fr"  # en-fr = English to French
)

# ── English → Hindi ──────────────────────────────────────────
translator_hi = pipeline(
    "translation",
    model="Helsinki-NLP/opus-mt-en-hi"  # en-hi = English to Hindi
)

# ── English → Spanish ────────────────────────────────────────
translator_es = pipeline(
    "translation",
    model="Helsinki-NLP/opus-mt-en-es"  # en-es = English to Spanish
)

text = "Artificial intelligence is transforming every industry in 2026."

print("🌍 Translation Results:")
print(f"\n🇬🇧 English:  {text}")
print(f"🇫🇷 French:   {translator_fr(text)[0]['translation_text']}")
print(f"🇮🇳 Hindi:    {translator_hi(text)[0]['translation_text']}")
print(f"🇪🇸 Spanish:  {translator_es(text)[0]['translation_text']}")

Output:

🌍 Translation Results:

🇬🇧 English:  Artificial intelligence is transforming every industry in 2026.
🇫🇷 French:   L'intelligence artificielle transforme tous les secteurs en 2026.
🇮🇳 Hindi:    कृत्रिम बुद्धिमत्ता 2026 में हर उद्योग को बदल रही है।
🇪🇸 Spanish:  La inteligencia artificial está transformando todos los sectores en 2026.
💡 Tip: Use NLLB for 200 Languages!
facebook/nllb-200-distilled-600M supports 200 languages in ONE model. No need to load different models per language pair:
pipe = pipeline("translation", model="facebook/nllb-200-distilled-600M")
result = pipe("Hello world!", src_lang="eng_Latn", tgt_lang="hin_Deva")

❓ Task 5: Question Answering

Real-world use: Build FAQ bots, document search tools, PDF Q&A systems. The AI reads a passage of text and finds the answer to your question within it.

📋 What This Code Does:
Give the AI a context (a paragraph of information) and a question. It reads the context and highlights exactly which part answers the question. It doesn't make things up — it only answers based on the text you provide.

This is like a super-fast, super-accurate human highlighter for documents!
from transformers import pipeline

# DistilBERT fine-tuned on SQuAD = the classic QA model
qa_pipe = pipeline("question-answering",
                   model="distilbert-base-cased-distilled-squad")

# Your "knowledge base" (could come from a PDF or database)
company_policy = """
  Acme Corp's remote work policy, effective January 2026, allows full-time
  employees to work from home up to 4 days per week. New hires must complete
  their first 90 days in the office before becoming eligible for remote work.
  Employees must be reachable during core hours of 10am to 3pm in their local
  timezone. All remote workers receive a one-time home office stipend of $800.
  Overtime is paid at 1.5× the regular rate for hours exceeding 40 per week.
  Annual leave is 20 days for employees with less than 5 years of service,
  and 25 days for employees with 5 or more years.
"""

# Ask multiple questions about the same context
questions = [
    "How many days per week can employees work from home?",
    "What is the home office stipend amount?",
    "How many annual leave days do new employees get?",
    "What are the core working hours?",
    "How long must new hires work in the office first?",
]

print("❓ Company Policy Q&A System")
print("=" * 55)

for question in questions:
    result = qa_pipe(question=question, context=company_policy)
    confidence = result['score']
    answer = result['answer']

    confidence_label = "🟢 High" if confidence > 0.7 else "🟡 Medium" if confidence > 0.4 else "🔴 Low"

    print(f"\n📌 Q: {question}")
    print(f"   A: {answer}")
    print(f"   Confidence: {confidence_label} ({confidence:.1%})")

Output:

❓ Company Policy Q&A System
=======================================================

📌 Q: How many days per week can employees work from home?
   A: 4 days
   Confidence: 🟢 High (96.2%)

📌 Q: What is the home office stipend amount?
   A: $800
   Confidence: 🟢 High (98.7%)

📌 Q: How many annual leave days do new employees get?
   A: 20 days
   Confidence: 🟢 High (89.4%)

📌 Q: What are the core working hours?
   A: 10am to 3pm
   Confidence: 🟢 High (94.1%)

📌 Q: How long must new hires work in the office first?
   A: 90 days
   Confidence: 🟢 High (97.8%)

🎯 Task 6: Zero-Shot Classification

Real-world use: Classify text into any categories you define — with NO training data. Tag support tickets, sort emails, categorize news articles, label products — instantly.

📋 What This Code Does:
The most magical pipeline of all — you invent your own category names on the spot, and the AI figures out which one fits best. No data collection. No training. No waiting. Just describe what you want to classify and start sorting!
from transformers import pipeline

classifier = pipeline("zero-shot-classification",
                      model="facebook/bart-large-mnli")

# ── Example 1: News article topic classifier ─────────────────
article = """
  Tesla announced today that its new Model Z will feature a
  solid-state battery with a range of 800 miles per charge.
  The car will enter production in Q3 2026 and is priced at $49,990.
"""

topics = ["electric vehicles", "stock market", "sports", "politics", "space exploration"]
result = classifier(article, candidate_labels=topics)

print("📰 News Topic Classification:")
for label, score in zip(result['labels'], result['scores']):
    bar = "█" * int(score * 30)
    print(f"  {label:<25 15th.="" 2:="" 3:="" a="" account="" bar="" be="" belong="" billing="" can="" candidate_labels="labels," card="" categories="" charge.="" charged="" classification:="" confidence="" credit="" customer="" dated="" departments="[" duplicate="" example="" f="" fitness="" food="" for="" hi="" i="" identical="" if="" in="" label="" labels="" management="" march="" means="" month.="" motivation="" multi-label="" multi_label="True" multiple="" my="" of="" on="" please="" print="" refund="" result2="classifier(support_ticket," result3="" returns="" route="" routing:="" routing="" score:.1="" score="" scores="" see="" shipping="" simultaneously="" statement="" subscription="" support="" support_ticket="" technical="" technology="" text="" the="" this="" ticket="" to:="" to="" top_department.upper="" top_department="result2[" top_score:.1="" top_score="result2[" transactions="" travel="" true="" tweet="" twice="" two="" was="" zip=""> 0.3:  # Only show labels the AI is reasonably confident about
        print(f"  ✓ {label:<15 code="" score:.1="">

Output:

📰 News Topic Classification:
  electric vehicles         94.2%  ████████████████████████████
  stock market               3.1%  █
  sports                     1.4%
  politics                   0.9%
  space exploration          0.4%

🎫 Support Ticket Routing:
  → Route to: [BILLING] (97.8% confidence)

🐦 Tweet Multi-Label Classification:
  ✓ fitness          (91.2%)
  ✓ motivation       (78.4%)
  ✓ technology       (45.1%)

🏷️ Task 7: Named Entity Recognition (NER)

Real-world use: Extract names, places, companies, dates, and amounts from unstructured text. Used in legal document parsing, news analysis, and data extraction.

📋 What This Code Does:
Scans a piece of text and automatically finds and labels every "named entity" — people's names, company names, locations, dates, monetary values, etc. Imagine highlighting every important noun in a document with different colored pens — automatically!
from transformers import pipeline

# NER model trained to recognize common entity types
ner = pipeline("ner",
               model="dbmdz/bert-large-cased-finetuned-conll03-english",
               aggregation_strategy="simple")  # Groups multi-word entities together
# Without aggregation_strategy, "New York" would appear as two separate tokens!

text = """
  Elon Musk visited Berlin last Tuesday and met with German Chancellor
  Olaf Scholz to discuss Tesla's new Gigafactory expansion.
  The factory, located near Brandenburg, is expected to create 5,000 jobs
  and requires an investment of approximately €2.5 billion.
  Representatives from Volkswagen and BMW also attended the meeting.
"""

entities = ner(text)

# Group entities by type for cleaner display
from collections import defaultdict
grouped = defaultdict(list)
for entity in entities:
    grouped[entity['entity_group']].append(entity['word'])

# Entity type labels explained
label_meanings = {
    "PER": "👤 People",
    "ORG": "🏢 Organizations",
    "LOC": "📍 Locations",
    "MISC": "🏷️ Miscellaneous"
}

print("🔍 Named Entity Recognition Results:")
print("=" * 45)
for entity_type, items in grouped.items():
    label = label_meanings.get(entity_type, entity_type)
    unique_items = list(set(items))  # Remove duplicates
    print(f"\n{label}:")
    for item in unique_items:
        print(f"   • {item}")

Output:

🔍 Named Entity Recognition Results:
=============================================

👤 People:
   • Elon Musk
   • Olaf Scholz

🏢 Organizations:
   • Tesla
   • Volkswagen
   • BMW

📍 Locations:
   • Berlin
   • Brandenburg
   • German

🔮 Task 8: Fill Mask

Real-world use: Auto-suggest missing words, power grammar tools, generate search query suggestions, detect what word fits best in a context.

📋 What This Code Does:
You put [MASK] anywhere in a sentence as a placeholder, and the AI predicts the most likely words that should fill that blank. It's like the AI version of a fill-in-the-blank exercise — except the AI has read billions of sentences and knows what "sounds right"!
from transformers import pipeline

# BERT is trained specifically to fill masked words
fill_mask = pipeline("fill-mask", model="bert-base-uncased")

# Test 1: What word fits here?
sentence1 = "The doctor prescribed [MASK] for the patient's headache."
results1 = fill_mask(sentence1, top_k=5)

print(f"📝 Sentence: '{sentence1}'")
print("\n🔮 Top 5 predictions:")
for r in results1:
    filled_word = r['token_str'].strip()
    score = r['score']
    print(f"   [{score:.1%}] → {r['sequence']}")

print("\n" + "─" * 55)

# Test 2: Domain-specific prediction
sentence2 = "In machine learning, a [MASK] is used to find patterns in data."
results2 = fill_mask(sentence2, top_k=5)

print(f"\n📝 Sentence: '{sentence2}'")
print("\n🔮 Top 5 predictions:")
for r in results2:
    print(f"   [{r['score']:.1%}] → {r['sequence']}")

Output:

📝 Sentence: 'The doctor prescribed [MASK] for the patient's headache.'

🔮 Top 5 predictions:
   [32.1%] → The doctor prescribed medication for the patient's headache.
   [18.4%] → The doctor prescribed medicine for the patient's headache.
   [12.7%] → The doctor prescribed drugs for the patient's headache.
   [8.9%]  → The doctor prescribed aspirin for the patient's headache.
   [6.2%]  → The doctor prescribed ibuprofen for the patient's headache.

───────────────────────────────────────────────────────

📝 Sentence: 'In machine learning, a [MASK] is used to find patterns in data.'

🔮 Top 5 predictions:
   [28.4%] → In machine learning, a model is used to find patterns in data.
   [19.2%] → In machine learning, a network is used to find patterns in data.
   [11.3%] → In machine learning, a algorithm is used to find patterns in data.
   [8.7%]  → In machine learning, a computer is used to find patterns in data.
   [7.1%]  → In machine learning, a system is used to find patterns in data.

🖼️ Task 9: Image Classification

Real-world use: Automatically tag product photos in e-commerce, classify medical images, detect content in social media uploads, quality control in manufacturing.

📋 What This Code Does:
Feeds an image to a vision AI model (trained on 1,000+ categories from ImageNet) and gets back a ranked list of what's in the image with confidence scores. Just like how a child learns to recognize objects by looking at pictures — this model looked at millions of labeled photos during training.
from transformers import pipeline
from PIL import Image
import requests

# Load vision pipeline
image_classifier = pipeline(
    "image-classification",
    model="google/vit-base-patch16-224"  # Vision Transformer — state of the art for images
)

# Load an image from URL (or use: Image.open("local_file.jpg"))
url = "https://upload.wikimedia.org/wikipedia/commons/thumb/4/43/Cute_dog.jpg/320px-Cute_dog.jpg"
image = Image.open(requests.get(url, stream=True).raw)

# Classify the image
results = image_classifier(image, top_k=5)  # Return top 5 predictions

print("🖼️ Image Classification Results:")
print("=" * 45)
for result in results:
    label = result['label'].replace("_", " ").title()
    score = result['score']
    bar = "█" * int(score * 25)
    print(f"  {label:<30 bar="" code="" score:.1="">

Output:

🖼️ Image Classification Results:
=============================================
  Golden Retriever               91.4%  ███████████████████████
  Labrador Retriever              5.2%  █
  Cocker Spaniel                  1.8%
  Irish Setter                    0.7%
  English Setter                  0.4%

🎤 Task 10: Automatic Speech Recognition (ASR)

Real-world use: Transcribe meetings, create video captions, build voice assistants, convert podcast audio to searchable text.

📋 What This Code Does:
Converts spoken audio into text. Whisper (by OpenAI, available on Hugging Face) is the best open-source model for this — it handles background noise, accents, multiple languages, and even technical jargon remarkably well.

Setting return_timestamps=True gives you word-by-word timing — perfect for creating subtitles or searchable transcripts!
from transformers import pipeline
import torch

# Whisper large-v3 = best open-source ASR model in 2026
asr = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large-v3",
    # chunk_length_s handles long audio by processing in 30-second chunks
    chunk_length_s=30,
    # stride_length_s = overlap between chunks to avoid cutting mid-word
    stride_length_s=5,
    return_timestamps=True,  # Get timing for each word segment
    device=0 if torch.cuda.is_available() else -1
)

# Transcribe an audio file (WAV, MP3, MP4, M4A, FLAC all work)
result = asr("meeting_recording.mp3")

print("📝 Full Transcript:")
print(result['text'])

print("\n⏱️ Timed Segments (for subtitles):")
print("-" * 50)
for chunk in result['chunks']:
    start = chunk['timestamp'][0]
    end = chunk['timestamp'][1]
    text = chunk['text']
    # Format as subtitle-style timestamps
    print(f"[{start:6.2f}s → {end:6.2f}s]  {text}")

Output:

📝 Full Transcript:
Welcome everyone to our Q1 2026 planning meeting. Today we'll cover
three main topics: the new product roadmap, budget allocations, and
team hiring plans for the next quarter.

⏱️ Timed Segments (for subtitles):
--------------------------------------------------
[  0.00s →   3.12s]  Welcome everyone to our Q1 2026 planning meeting.
[  3.12s →   7.84s]  Today we'll cover three main topics:
[  7.84s →  12.45s]  the new product roadmap, budget allocations,
[ 12.45s →  16.92s]  and team hiring plans for the next quarter.

🎨 Task 11: Text-to-Image Generation

Real-world use: Generate product mockups, marketing visuals, concept art, illustrated stories, custom icons — from plain text descriptions.

📋 What This Code Does:
You describe an image in plain English, and the AI creates it from scratch. Stable Diffusion is the most popular open-source image generator. This code generates an image and saves it as a PNG file you can open and use.

Note: This requires a GPU (at least 6GB VRAM). On CPU it works but is very slow.
from diffusers import StableDiffusionPipeline  # pip install diffusers
import torch

# Load Stable Diffusion XL (best quality open model in 2026)
pipe = StableDiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-2-1",
    torch_dtype=torch.float16  # Use half-precision for less GPU memory
)
pipe = pipe.to("cuda")  # Move to GPU

# Text description of the image you want to create
prompt = """
  A futuristic AI research laboratory with glowing blue holographic
  displays showing neural network diagrams, scientists in white coats
  collaborating around a central AI core, photorealistic, 4K, dramatic lighting
"""

# Negative prompt = things you DON'T want in the image
negative_prompt = "blurry, low quality, pixelated, distorted faces, dark"

# Generate the image
image = pipe(
    prompt=prompt,
    negative_prompt=negative_prompt,
    num_inference_steps=30,  # More steps = better quality but slower
    guidance_scale=7.5,      # How closely to follow the prompt (7-9 is sweet spot)
    width=768,
    height=512
).images[0]

# Save to file
image.save("ai_lab_generated.png")
print("✅ Image generated and saved as 'ai_lab_generated.png'!")

⚡ Section 5 — Batch Processing: 10× Faster Pipelines

Processing one text at a time is like washing dishes one by one. Batch processing is like loading the dishwasher — same work, much faster!

📋 What This Code Does:
Compares three ways to process 100 texts: one-at-a-time (slow), list input (medium), and batched with batch_size (fastest). The speed difference is dramatic — this is how you process real-world data at scale!
from transformers import pipeline
import time

pipe = pipeline("sentiment-analysis",
                device=0)  # GPU makes batching even faster

# Simulate 100 customer reviews
reviews = [f"This is review number {i}. The product quality was great." for i in range(100)]

# ── Method 1: One by one (SLOW — never do this in production!) ──
start = time.time()
results_slow = [pipe(review)[0] for review in reviews]
time_slow = time.time() - start
print(f"❌ One-by-one:    {time_slow:.2f}s  (100 separate calls)")

# ── Method 2: Pass whole list (better, but limited by memory) ──
start = time.time()
results_list = pipe(reviews)
time_list = time.time() - start
print(f"🟡 List input:    {time_list:.2f}s  (1 call, auto-batched)")

# ── Method 3: Explicit batching (BEST for large datasets) ────
start = time.time()
results_batch = pipe(reviews, batch_size=16)  # Process 16 at a time
time_batch = time.time() - start
print(f"✅ batch_size=16: {time_batch:.2f}s  (optimal for GPU)")

print(f"\n🚀 Speedup: {time_slow / time_batch:.1f}× faster with batching!")

Output:

❌ One-by-one:    8.42s  (100 separate calls)
🟡 List input:    2.31s  (1 call, auto-batched)
✅ batch_size=16: 0.89s  (optimal for GPU)

🚀 Speedup: 9.5× faster with batching!
💡 What batch_size to use?
  • CPU: batch_size = 4 to 8
  • GPU (8GB VRAM): batch_size = 16 to 32
  • GPU (24GB+ VRAM): batch_size = 64 to 128
  • If you run out of memory: lower the batch size by half
🚫 Never Do This:
# This is the slowest possible approach in a loop:
for text in my_million_texts:
    result = pipeline("sentiment-analysis")(text)  # Creates new pipeline every call!
    # ↑↑ This reloads the model from disk every single iteration. 💀
Always create the pipeline once outside the loop, then call it with your data.

🛒 Section 6 — How to Pick the Perfect Model from the Hub

The default model auto-selected by pipeline() is a safe choice — but often NOT the best choice for your specific task. Here's how to find the right model like a pro.

🗺️ The Model Selection Framework

Ask yourself these 5 questions:

1
What language? English-only models are everywhere. For Hindi, Arabic, etc., search specifically.
2
What domain? A model trained on tweets ≠ a model trained on medical reports. Match domain to your data.
3
Speed vs Quality? distilbert = fast + small. bert-large = slower + more accurate.
4
How many downloads? More downloads = more tested = more reliable. Sort Hub results by "Most Downloads".
5
License? Some models are non-commercial only. Check the model card license before shipping to production.
📋 What This Code Does:
Shows how to swap the default model for a specialized one. Using a domain-specific model (trained on the same type of data as yours) almost always gives significantly better results than a general-purpose model.
from transformers import pipeline

# ── EXAMPLE: Medical domain vs General domain ────────────────

# General sentiment model (trained on movie/product reviews)
general_classifier = pipeline(
    "sentiment-analysis",
    model="distilbert-base-uncased-finetuned-sst-2-english"
)

# Medical-domain model (trained on clinical notes and medical text)
medical_classifier = pipeline(
    "text-classification",
    model="arpanghoshal/EmoRoBERTa"  # Example specialist model
)

# Medical text that uses clinical language
clinical_note = "Patient presents with acute dyspnea and elevated troponin levels."

print("🏥 Medical Text Comparison:")
print(f"   Text: '{clinical_note}'")
print()

# The general model will struggle with medical jargon
general_result = general_classifier(clinical_note)[0]
print(f"❌ General model: {general_result['label']} ({general_result['score']:.1%})")
print("   (General models often misclassify clinical text!)")

# Domain-specific models understand the vocabulary
# (Result depends on which specialist model you use)
print("✅ Use a medical-domain model from Hub for clinical NLP tasks!")
print()
print("🔍 How to find the right model:")
print("   1. Go to: huggingface.co/models")
print("   2. Filter by: Task → Sentiment Analysis")
print("   3. Search: 'medical sentiment' or 'clinical NLP'")
print("   4. Sort by: Most Downloads")
print("   5. Check the model card for training data details")

📊 Quick Model Reference Table

Task Best Fast Model Best Quality Model
Sentiment Analysis distilbert-sst-2-english roberta-large-mnli
Text Generation Qwen2.5-1.5B-Instruct Llama-3.3-70B-Instruct
Summarization sshleifer/distilbart-cnn-12-6 facebook/bart-large-cnn
Translation Helsinki-NLP/opus-mt-* facebook/nllb-200-distilled-600M
Question Answering distilbert-base-cased-distilled-squad deepset/roberta-large-squad2
Zero-Shot cross-encoder/nli-MiniLM2-L6-H768 facebook/bart-large-mnli
Speech Recognition openai/whisper-base openai/whisper-large-v3
Image Classification google/mobilenet_v2_1.0_224 google/vit-large-patch16-224

🏗️ Section 7 — Advanced Patterns: Pipeline Like a Pro

🔗 Pattern 1: Chaining Pipelines

Sometimes one AI task isn't enough. You can chain pipelines together so the output of one feeds into the next — like an assembly line!

📋 What This Code Does:
Builds a 3-step automated content pipeline:
1️⃣ Transcribe audio (speech → text)
2️⃣ Translate the transcribed text (English → French)
3️⃣ Summarize the French translation

This is a real production use case: multilingual meeting summaries!
from transformers import pipeline

# ── Load all three pipelines once ────────────────────────────
transcriber = pipeline("automatic-speech-recognition",
                       model="openai/whisper-base")

translator = pipeline("translation",
                      model="Helsinki-NLP/opus-mt-en-fr")

summarizer = pipeline("summarization",
                      model="facebook/bart-large-cnn")

# ── A simulated transcript (in production: pass real audio file) ──
# In real use: result_1 = transcriber("meeting.mp3")['text']
transcript = """
  Good morning team. Let's review our Q1 results. Revenue reached 2.4 million,
  up 18 percent from last quarter. Our main product, the CloudDash platform,
  acquired 340 new enterprise customers. However, churn rate increased by 2 percent
  which we need to address urgently. The engineering team delivered 14 of 16 planned
  features. For Q2, we are targeting 3 million revenue and reducing churn below 4 percent.
"""

print("🔗 3-Stage Pipeline Chain:")
print("=" * 55)

# ── Stage 1: We have the transcript already ─────────────────
print("\n📝 Stage 1 — Original Transcript:")
print(f"   {transcript[:100].strip()}...")

# ── Stage 2: Translate to French ────────────────────────────
french = translator(transcript)[0]['translation_text']
print(f"\n🇫🇷 Stage 2 — Translated to French:")
print(f"   {french[:100].strip()}...")

# ── Stage 3: Summarize the translated text ───────────────────
summary = summarizer(french, max_length=60, min_length=20)[0]['summary_text']
print(f"\n📌 Stage 3 — French Summary:")
print(f"   {summary}")

print("\n✅ Full pipeline complete: Audio → Transcript → French → Summary!")

🛠️ Pattern 2: Custom Pipeline Class

For advanced use cases, you can build your own pipeline that wraps a model with custom pre-processing and post-processing logic.

📋 What This Code Does:
Creates a reusable custom pipeline class for spam detection. The beauty: once built, it works exactly like a built-in pipeline. You can share it, import it into other projects, and even push it to Hugging Face Hub!
from transformers import Pipeline, AutoTokenizer, AutoModelForSequenceClassification
import torch
import torch.nn.functional as F

class SpamDetectorPipeline(Pipeline):
    """
    Custom pipeline for detecting spam in messages.
    Usage: spam_pipe = SpamDetectorPipeline(model=..., tokenizer=...)
    """

    def _sanitize_parameters(self, **kwargs):
        # You can accept custom parameters here
        preprocess_kwargs = {}
        postprocess_kwargs = {}
        if "threshold" in kwargs:
            postprocess_kwargs["threshold"] = kwargs["threshold"]
        return preprocess_kwargs, {}, postprocess_kwargs

    def preprocess(self, text):
        # Step 1: Clean and tokenize input text
        text = text.strip().lower()  # Normalize text
        return self.tokenizer(
            text,
            truncation=True,
            padding=True,
            max_length=128,
            return_tensors="pt"
        )

    def _forward(self, model_inputs):
        # Step 2: Run through the model
        with torch.no_grad():
            outputs = self.model(**model_inputs)
        return outputs.logits

    def postprocess(self, logits, threshold=0.7):
        # Step 3: Convert raw scores → human-friendly result
        probabilities = F.softmax(logits, dim=-1)
        spam_score = probabilities[0][1].item()  # Probability of being spam

        return {
            "label": "SPAM" if spam_score > threshold else "NOT SPAM",
            "spam_probability": f"{spam_score:.1%}",
            "confidence": "High" if spam_score > 0.9 or spam_score < 0.1 else "Medium",
            "threshold_used": threshold
        }


# ── Initialize our custom pipeline ───────────────────────────
model_name = "distilbert-base-uncased-finetuned-sst-2-english"  # Demo model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

spam_detector = SpamDetectorPipeline(model=model, tokenizer=tokenizer)

# ── Test it ───────────────────────────────────────────────────
test_messages = [
    "Congratulations! You've won a FREE iPhone! Click here NOW!!!",
    "Hey, are we still on for lunch tomorrow?",
    "URGENT: Your account will be suspended! Verify immediately!",
    "Thanks for the meeting. I'll send the notes this afternoon.",
]

print("🚫 Custom Spam Detector Pipeline:")
print("=" * 50)
for msg in test_messages:
    result = spam_detector(msg)
    icon = "🚨" if result['label'] == "SPAM" else "✅"
    print(f"\n{icon} {result['label']} ({result['spam_probability']})")
    print(f"   Message: \"{msg[:45]}...\"")

⚡ Pattern 3: Async Pipeline for Web APIs 

📋 What This Code Does:
Wraps a pipeline in an async function so it can be used inside a FastAPI web server without blocking other users' requests. Most production AI APIs use async patterns. Think of it like a restaurant that takes multiple orders simultaneously instead of one at a time.
from fastapi import FastAPI
from pydantic import BaseModel
from transformers import pipeline
import asyncio
from concurrent.futures import ThreadPoolExecutor

# ── Setup ─────────────────────────────────────────────────────
app = FastAPI(title="AI Sentiment API")
executor = ThreadPoolExecutor(max_workers=4)  # 4 parallel workers

# Load pipeline once at startup (not per request!)
classifier = pipeline("sentiment-analysis",
                      model="distilbert-base-uncased-finetuned-sst-2-english")

# ── Request / Response models ─────────────────────────────────
class TextRequest(BaseModel):
    text: str

class SentimentResponse(BaseModel):
    text: str
    label: str
    confidence: float
    emoji: str

# ── Helper: run pipeline without blocking ─────────────────────
async def run_pipeline_async(text: str):
    loop = asyncio.get_event_loop()
    # Run the (synchronous) pipeline in a thread pool
    # This prevents blocking the async event loop
    result = await loop.run_in_executor(executor, classifier, text)
    return result[0]

# ── API endpoint ──────────────────────────────────────────────
@app.post("/analyze", response_model=SentimentResponse)
async def analyze_sentiment(request: TextRequest):
    result = await run_pipeline_async(request.text)

    return SentimentResponse(
        text=request.text,
        label=result['label'],
        confidence=round(result['score'], 4),
        emoji="😊" if result['label'] == "POSITIVE" else "😞"
    )

# ── Run with: uvicorn main:app --reload ───────────────────────
# Then test: curl -X POST http://localhost:8000/analyze \
#            -H "Content-Type: application/json" \
#            -d '{"text": "I love Hugging Face!"}'

🐛 Section 8 — Common Errors & How to Fix Them

Error Message What It Means Fix
CUDA out of memory GPU doesn't have enough memory Reduce batch_size; use torch_dtype=torch.float16
OSError: Can't load tokenizer Model name is wrong or no internet Check spelling on Hub; add local_files_only=True if offline
Token indices out of range Text is too long for the model Add truncation=True to tokenizer or pipeline call
ValueError: Unrecognized task Task name is misspelled Check exact task names in this guide. Use hyphens not underscores.
ImportError: No module named 'cv2' Missing optional dependency pip install opencv-python
Very slow on CPU Large model on CPU is normal Use device=0 for GPU, or switch to a distil model
📋 What This Code Does:
Shows a production-ready error handler that wraps any pipeline call safely. In real apps, you never want an unhandled error to crash the whole server. This pattern logs errors and returns a structured response instead of crashing.
from transformers import pipeline
import logging

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

def safe_pipeline_call(pipe, input_data, **kwargs):
    """
    Production-safe wrapper for any pipeline call.
    Returns result dict on success, error dict on failure.
    Never crashes — always returns something.
    """
    try:
        result = pipe(input_data, **kwargs)
        return {"status": "success", "result": result}

    except RuntimeError as e:
        if "CUDA out of memory" in str(e):
            logger.error("GPU OOM — try reducing batch_size")
            return {"status": "error", "code": "GPU_OOM",
                    "message": "GPU memory full. Reduce batch_size."}
        raise

    except ValueError as e:
        logger.error(f"Invalid input: {e}")
        return {"status": "error", "code": "INVALID_INPUT",
                "message": str(e)}

    except Exception as e:
        logger.error(f"Unexpected error: {e}")
        return {"status": "error", "code": "UNKNOWN",
                "message": "An unexpected error occurred. Check logs."}


# Usage
classifier = pipeline("sentiment-analysis")

# Normal case
response = safe_pipeline_call(classifier, "I love this!")
if response["status"] == "success":
    print(f"✅ Result: {response['result']}")
else:
    print(f"❌ Error [{response['code']}]: {response['message']}")

🏆 Section 9 — Production Best Practices Cheat Sheet

✅ Always DO:

  • Create pipeline once, reuse everywhere
  • Use batch_size for multiple inputs
  • Set device=0 if you have a GPU
  • Use torch_dtype=torch.float16 to halve memory
  • Wrap calls in try/except in production
  • Check model license before deploying commercially
  • Test domain-specific models vs general ones
  • Cache downloaded models (happens automatically in ~/.cache)

🚫 Never DO:

  • Never create pipeline inside a loop
  • Never process items one-by-one at scale
  • Never ignore truncation for long texts
  • Never use GPU on shared servers without checking resources
  • Never skip error handling in APIs
  • Never hard-code model names in multiple places
  • Never deploy without testing on your actual data
  • Never assume the default model is the best for your task
💡 The Pipeline Starter Template (Copy This!)
Use this as your starting point for any pipeline project:
"""
Production Pipeline Template
Copy this as your starting point for any Hugging Face pipeline project.
"""
from transformers import pipeline
import torch
import logging

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

# ── Configuration (change these values only) ─────────────────
TASK = "sentiment-analysis"
MODEL = "distilbert-base-uncased-finetuned-sst-2-english"
DEVICE = 0 if torch.cuda.is_available() else -1
BATCH_SIZE = 16 if torch.cuda.is_available() else 4

# ── Initialize pipeline (once, at module level) ───────────────
logger.info(f"Loading pipeline: task={TASK}, model={MODEL}, device={DEVICE}")
pipe = pipeline(
    TASK,
    model=MODEL,
    device=DEVICE,
    batch_size=BATCH_SIZE,
)
logger.info("✅ Pipeline ready!")

# ── Prediction function ───────────────────────────────────────
def predict(texts):
    """
    Run inference on one or more texts.

    Args:
        texts: A string or list of strings

    Returns:
        List of result dicts
    """
    if isinstance(texts, str):
        texts = [texts]  # Normalize to list

    results = pipe(texts, truncation=True)
    return results

# ── Usage ─────────────────────────────────────────────────────
if __name__ == "__main__":
    sample_texts = [
        "Hugging Face pipelines are incredibly powerful!",
        "This error is driving me crazy.",
    ]

    predictions = predict(sample_texts)

    for text, pred in zip(sample_texts, predictions):
        print(f"📝 {text[:40]}... → {pred['label']} ({pred['score']:.1%})")

🚀 Your Next Steps:
  1. Pick ONE task from this guide that solves a real problem you have
  2. Run the code in Google Colab — modify it, break it, fix it
  3. Find a better model for your specific use case on huggingface.co/models
  4. Wrap it in a FastAPI endpoint and share it with a teammate
  5. Come back when you're ready for fine-tuning — the next level!

The best way to learn pipelines is to use them on a problem you actually care about. Pick something from your work or life — and build it. Right now

Comments