Topic Clustering in Machine Learning: Group and Organize Text Data with Unsupervised Learning
Imagine you receive 100,000 customer support tickets every month. No human team can read all of them. But what if a machine could automatically group them into buckets like "Billing Issues", "Technical Errors", "Shipping Delays" — without anyone telling it what categories to look for?
That is exactly what Topic Clustering does. It finds hidden structure in large collections of text — automatically, intelligently, and at scale. In modern MLOps pipelines, it powers everything from data labelling to model monitoring to search engine organisation.
💡 Think of it like: A super-smart librarian who reads every book in a warehouse, then organises them into themed shelves — without being told what the themes are in advance. 📚
What Exactly Is Topic Clustering?
Topic Clustering is an unsupervised machine learning technique that groups a collection of text documents into clusters, where each cluster represents a common underlying theme or topic.
The key word here is unsupervised.
- Supervised learning → You label the data first, then train a model to predict those labels
- Unsupervised learning → You give the model raw, unlabelled data and it discovers structure on its own
- Topic Clustering is unsupervised — no labelling required before you start
This makes it incredibly powerful for exploring datasets you've never seen before — or for continuously monitoring incoming data streams in a live MLOps system.
Why Does Topic Clustering Matter in MLOps?
MLOps is the practice of building, deploying, and maintaining machine learning models in production. Topic Clustering plays a critical role across several stages of this lifecycle:
- Data Understanding → Before training any model, you need to know what topics your dataset covers. Topic clustering gives you an instant overview.
- Data Labelling → Cluster first, label one representative per cluster, then propagate labels. This reduces human labelling effort by 90%+.
- Data Quality → Clusters with very few members often signal noisy or out-of-distribution data that shouldn't be in your training set.
- Model Monitoring → After deployment, cluster incoming requests to detect if users are asking about topics the model was never trained on (concept drift).
- Instruction Dataset Creation → In LLM fine-tuning, cluster your instructions to ensure diverse topic coverage and avoid accidentally training on 80% coding samples and 20% everything else.
The Topic Clustering Pipeline — Step by Step
Topic clustering is not a single algorithm. It's a pipeline of steps that transform raw text into meaningful, interpretable clusters. Here is the complete flow:
📐 The Topic Clustering Pipeline
Let's walk through each stage with real code and real examples!
Step 1 — Raw Text Collection & Preprocessing
Before any machine learning can happen, you need clean text. Raw text from the real world is messy — it contains HTML tags, emojis, duplicate whitespace, special characters, and language noise.
Preprocessing transforms messy raw text into clean, consistent input that the next steps can work with reliably.
What Preprocessing Typically Includes
- Lowercasing → "Machine Learning" and "machine learning" should be treated as the same phrase
- Removing HTML tags → Strip
<b>,<p>and other markup from web-scraped text - Removing special characters → Punctuation and symbols that don't carry semantic meaning
- Stripping extra whitespace → Multiple spaces, newlines, tabs collapsed to single spaces
- Language filtering → Keep only samples in your target language if needed
Preprocessing Code Example
import re
def clean_text(text: str) -> str:
# Remove HTML tags
text = re.sub(r'<[^>]+>', ' ', text)
# Remove URLs
text = re.sub(r'http\S+|www\S+', ' ', text)
# Remove special characters (keep letters, digits, spaces)
text = re.sub(r'[^a-zA-Z0-9\s]', ' ', text)
# Collapse multiple whitespace into one
text = re.sub(r'\s+', ' ', text)
# Lowercase and strip leading/trailing whitespace
return text.strip().lower()
# Example usage
raw = "We're having issues with our ORDER #12345!! Contact: help@example.com
"
print(clean_text(raw))
Output:
we are having issues with our order 12345 contact help example com
Clean, simple, ready for embedding! 🎯
Step 2 — Text Embeddings (Turning Words into Numbers)
Computers don't understand words — they understand numbers. Embeddings convert text into dense numerical vectors (lists of numbers) where similar texts produce vectors that are close together in mathematical space.
💡 Think of it like: A map of the world 🗺️. Cities in the same country are geographically close. Similarly, documents about the same topic produce embeddings that are close to each other in a high-dimensional mathematical space.
Traditional vs. Modern Embedding Approaches
- TF-IDF (Traditional) → Represents each document as a sparse vector based on word frequencies. Fast and interpretable, but misses meaning and context entirely. "Bank" (river) and "Bank" (finance) get the same representation.
- Word2Vec / GloVe (2013–2016) → Learn word-level meaning from co-occurrence patterns. Better than TF-IDF, but still operate at the word level. A sentence embedding is just an average of its word embeddings — losing sentence structure.
- Sentence Transformers (Modern) → Encode entire sentences and paragraphs into rich, context-aware embeddings. Two sentences with the same meaning but different words will have very similar embeddings. This is the state-of-the-art approach for topic clustering.
Generating Embeddings with Sentence-Transformers
from sentence_transformers import SentenceTransformer
# Load a pre-trained embedding model
# 'all-MiniLM-L6-v2' is fast, small, and surprisingly powerful
model = SentenceTransformer('all-MiniLM-L6-v2')
documents = [
"My credit card was charged twice for the same order.",
"The app keeps crashing when I open the settings menu.",
"Where is my package? It has been 10 days since I ordered.",
"I was billed incorrectly on my last invoice.",
"The mobile app freezes every time I try to log in.",
"My shipment still hasn't arrived after two weeks.",
"There is a duplicate charge on my bank statement.",
"The software crashes immediately after launch.",
]
# Generate embeddings — each document becomes a 384-dimensional vector
embeddings = model.encode(documents, show_progress_bar=True)
print(f"Number of documents: {len(documents)}")
print(f"Embedding shape: {embeddings.shape}")
print(f"First embedding (first 5 values): {embeddings[0][:5]}")
Output:
Number of documents: 8
Embedding shape: (8, 384)
First embedding (first 5 values): [ 0.0234 -0.1123 0.0891 0.2341 -0.0567]
Each document is now a list of 384 numbers that capture its meaning! Documents about billing will have similar numbers. Documents about crashes will have similar numbers. The machine can now measure similarity mathematically. 🔢
all-MiniLM-L6-v2 (fast, 80MB) or all-mpnet-base-v2 (more accurate, 420MB).
For multilingual data, use paraphrase-multilingual-MiniLM-L12-v2.
All are free and available on the Hugging Face Hub.
Step 3 — Dimensionality Reduction
Our embeddings have 384 dimensions. Running clustering directly on 384-dimensional vectors is problematic — this is known as the "Curse of Dimensionality".
In very high dimensions, all points start to look equally far apart from each other. Clustering algorithms lose their effectiveness because distance measurements stop being meaningful.
💡 Think of it like: Trying to group people by similarity using 384 different personality traits simultaneously. It's overwhelming and ineffective. But if you reduce it to the top 5 most distinguishing traits, grouping becomes obvious and accurate.
Two Popular Dimensionality Reduction Tools
- PCA (Principal Component Analysis) → Classic, fast, linear method. Preserves global structure well. Good for initial exploration. Can be lossy if your data has complex non-linear structure.
- UMAP (Uniform Manifold Approximation and Projection) → Modern, non-linear method. Preserves both local and global structure far better than PCA. The current gold standard for dimensionality reduction before clustering. Also produces beautiful 2D visualisations.
Applying UMAP to Our Embeddings
import umap
# Reduce 384 dimensions → 5 dimensions for clustering
# (Keep 5D for clustering, reduce to 2D separately just for visualisation)
reducer = umap.UMAP(
n_components=5, # target dimensions for clustering
n_neighbors=15, # controls local vs global structure balance
min_dist=0.0, # tighter clusters (0.0 is best for clustering, not viz)
metric='cosine', # cosine similarity works best for text embeddings
random_state=42 # for reproducibility
)
reduced_embeddings = reducer.fit_transform(embeddings)
print(f"Original shape: {embeddings.shape}")
print(f"Reduced shape: {reduced_embeddings.shape}")
Output:
Original shape: (8, 384)
Reduced shape: (8, 5)
random_state=42 in UMAP (and all other stochastic algorithms).
Without it, you get different cluster assignments every run — which makes debugging, reproducibility,
and comparing runs nearly impossible in a production MLOps pipeline.
Step 4 — Clustering Algorithms
Now we have our reduced embeddings. It's time to group similar documents together using a clustering algorithm.
There are three main families of clustering algorithms used in topic clustering. Each has different strengths and is suited to different situations.
Algorithm 1 — K-Means Clustering
K-Means partitions your data into exactly K clusters (you specify K upfront). It works by placing K "centroids" (imaginary cluster centres) and assigning each point to the nearest centroid, then adjusting the centroids until the assignments stabilise.
💡 Analogy: You drop K pins on a map, then everyone travels to the nearest pin. You then move each pin to the centre of its group. Repeat until no one moves. 📍
When to use K-Means:
- You know roughly how many topics you expect
- Your dataset is large (K-Means scales well)
- Your clusters are roughly similar in size
from sklearn.cluster import KMeans
# Cluster our reduced embeddings into 3 clusters
kmeans = KMeans(
n_clusters=3,
n_init=10, # run 10 times with different initialisations, keep best result
random_state=42
)
kmeans.fit(reduced_embeddings)
labels = kmeans.labels_
# Print each document with its assigned cluster
for doc, label in zip(documents, labels):
print(f"Cluster {label}: {doc[:60]}...")
Output:
Cluster 0: My credit card was charged twice for the same order....
Cluster 1: The app keeps crashing when I open the settings menu....
Cluster 2: Where is my package? It has been 10 days since I order...
Cluster 0: I was billed incorrectly on my last invoice....
Cluster 1: The mobile app freezes every time I try to log in....
Cluster 2: My shipment still hasn't arrived after two weeks....
Cluster 0: There is a duplicate charge on my bank statement....
Cluster 1: The software crashes immediately after launch....
Cluster 0 = Billing issues. Cluster 1 = App crashes. Cluster 2 = Shipping delays. It found the three real topics with zero labelling! 🎉
How to Choose K — The Elbow Method
The biggest challenge with K-Means is choosing the right K. The Elbow Method helps by plotting the "inertia" (sum of distances from each point to its cluster centre) for different values of K. The optimal K is usually at the "elbow" of the curve where inertia stops decreasing rapidly.
import matplotlib.pyplot as plt
inertias = []
k_range = range(2, 15)
for k in k_range:
km = KMeans(n_clusters=k, n_init=10, random_state=42)
km.fit(reduced_embeddings)
inertias.append(km.inertia_)
plt.figure(figsize=(8, 4))
plt.plot(k_range, inertias, 'bo-', markersize=8)
plt.xlabel('Number of Clusters (K)')
plt.ylabel('Inertia')
plt.title('Elbow Method — Finding the Optimal K')
plt.grid(True, alpha=0.3)
plt.axvline(x=3, color='red', linestyle='--', label='Elbow at K=3')
plt.legend()
plt.show()
Look for the "elbow" — the point where adding more clusters gives diminishing returns! 📉
Algorithm 2 — HDBSCAN (The Modern Favourite)
HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) is the current state-of-the-art for topic clustering in NLP and LLM pipelines.
Unlike K-Means, HDBSCAN does not require you to specify the number of clusters upfront. It discovers clusters based on density — areas where documents are tightly packed together. Documents in low-density regions are classified as noise (outliers) rather than forced into a cluster.
Key advantages over K-Means:
- No need to specify K in advance
- Handles clusters of different sizes and shapes naturally
- Explicitly marks outliers as noise (label = -1) instead of forcing them into a cluster
- Much more robust on real-world, messy text data
import hdbscan
clusterer = hdbscan.HDBSCAN(
min_cluster_size=5, # minimum documents to form a cluster
min_samples=3, # controls how conservative cluster detection is
metric='euclidean', # use euclidean on UMAP-reduced embeddings
cluster_selection_method='eom' # 'eom' generally gives more clusters than 'leaf'
)
labels = clusterer.fit_predict(reduced_embeddings)
# Count how many documents ended up in each cluster
from collections import Counter
print(Counter(labels))
# -1 means 'noise' — documents that didn't fit any cluster
Output:
Counter({0: 3, 1: 3, 2: 3, -1: 2})
Algorithm 3 — Agglomerative (Hierarchical) Clustering
Hierarchical clustering builds a tree structure (called a dendrogram) showing how documents group together at different levels of similarity.
💡 Analogy: Think of a family tree 🌳. At the bottom, every person is their own node. As you move up, siblings merge into families, families into extended families, and so on. You can cut the tree at any level to get different granularities of grouping.
This is useful when you want to explore topic hierarchies — for example, "Technology" at a high level that breaks down into "Mobile Apps", "Desktop Software", and "Web Platforms" at a more detailed level.
from sklearn.cluster import AgglomerativeClustering
from scipy.cluster.hierarchy import dendrogram, linkage
import matplotlib.pyplot as plt
# Build linkage matrix for visualisation
Z = linkage(reduced_embeddings, method='ward')
plt.figure(figsize=(12, 5))
dendrogram(
Z,
labels=[doc[:30] + "..." for doc in documents],
leaf_rotation=45,
leaf_font_size=9
)
plt.title('Hierarchical Clustering Dendrogram')
plt.tight_layout()
plt.show()
# Cut the tree at 3 clusters
agg = AgglomerativeClustering(n_clusters=3, linkage='ward')
labels = agg.fit_predict(reduced_embeddings)
print(labels)
Output:
[0 1 2 0 1 2 0 1]
Step 5 — Topic Labelling
After clustering, each cluster has a number (0, 1, 2...) but no human-readable name. Topic labelling is the process of assigning a meaningful description to each cluster.
There are two approaches:
Approach 1 — Keyword Extraction (Traditional)
Extract the most distinctive words from each cluster using TF-IDF or KeyBERT. The top keywords become the cluster's label.
from sklearn.feature_extraction.text import TfidfVectorizer
import numpy as np
def get_cluster_keywords(documents, labels, n_keywords=5):
unique_labels = sorted(set(labels))
cluster_keywords = {}
for label in unique_labels:
if label == -1:
continue # skip noise cluster
# Collect all documents in this cluster
cluster_docs = [doc for doc, lbl in zip(documents, labels) if lbl == label]
# TF-IDF across cluster documents
vectorizer = TfidfVectorizer(stop_words='english', max_features=50)
tfidf_matrix = vectorizer.fit_transform(cluster_docs)
# Sum TF-IDF scores across all docs in cluster
scores = np.array(tfidf_matrix.sum(axis=0)).flatten()
top_indices = scores.argsort()[-n_keywords:][::-1]
top_words = [vectorizer.get_feature_names_out()[i] for i in top_indices]
cluster_keywords[label] = top_words
print(f"Cluster {label}: {', '.join(top_words)}")
return cluster_keywords
keywords = get_cluster_keywords(documents, labels)
Output:
Cluster 0: charged, billing, invoice, duplicate, bank
Cluster 1: app, crashes, freezes, software, mobile
Cluster 2: package, shipment, order, arrived, shipping
From these keywords, a human immediately reads: Cluster 0 = Billing, Cluster 1 = Tech Issues, Cluster 2 = Delivery! 📦
Approach 2 — LLM-Based Labelling (Modern & Powerful)
Send the top documents from each cluster to an LLM and ask it to generate a descriptive label. This produces much more natural, contextually aware topic names than keyword extraction alone.
from openai import OpenAI
client = OpenAI()
def label_cluster_with_llm(cluster_documents: list, n_examples: int = 5) -> str:
# Use the first N documents as representative examples
examples = cluster_documents[:n_examples]
examples_text = "\n".join(f"- {doc}" for doc in examples)
prompt = f"""You are analysing a cluster of customer support tickets.
Below are representative examples from this cluster:
{examples_text}
In 3-5 words, what is the single best topic label for this cluster?
Return ONLY the label text, nothing else."""
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0
)
return response.choices[0].message.content.strip()
# Label all clusters
from collections import defaultdict
cluster_docs = defaultdict(list)
for doc, label in zip(documents, labels):
if label != -1:
cluster_docs[label].append(doc)
for cluster_id, docs in sorted(cluster_docs.items()):
label_text = label_cluster_with_llm(docs)
print(f"Cluster {cluster_id}: '{label_text}'")
Output:
Cluster 0: 'Billing & Payment Issues'
Cluster 1: 'App Crashes & Technical Errors'
Cluster 2: 'Shipping & Delivery Delays'
Human-quality topic names, generated automatically! 🏆
Step 6 — Visualisation
Numbers and labels are useful, but humans understand patterns best when they can see them. Visualising your clusters helps you quickly validate whether the clustering makes sense and communicate results to stakeholders.
2D UMAP Visualisation with Matplotlib
import matplotlib.pyplot as plt
import umap
import numpy as np
# Create a 2D reduction just for visualisation (separate from the 5D used for clustering)
viz_reducer = umap.UMAP(
n_components=2,
n_neighbors=15,
min_dist=0.1,
metric='cosine',
random_state=42
)
coords_2d = viz_reducer.fit_transform(embeddings)
# Plot
topic_labels = {
0: 'Billing Issues',
1: 'App Crashes',
2: 'Shipping Delays',
-1: 'Noise'
}
colors = {0: '#2196F3', 1: '#4CAF50', 2: '#FF9800', -1: '#9E9E9E'}
plt.figure(figsize=(10, 7))
for cluster_id in sorted(set(labels)):
mask = np.array(labels) == cluster_id
plt.scatter(
coords_2d[mask, 0],
coords_2d[mask, 1],
c=colors[cluster_id],
label=topic_labels[cluster_id],
alpha=0.8,
s=100,
edgecolors='white',
linewidth=0.5
)
plt.title('Topic Clusters — 2D UMAP Projection', fontsize=14, fontweight='bold')
plt.xlabel('UMAP Dimension 1')
plt.ylabel('UMAP Dimension 2')
plt.legend(title='Topics', loc='best')
plt.grid(True, alpha=0.2)
plt.tight_layout()
plt.savefig('topic_clusters.png', dpi=150, bbox_inches='tight')
plt.show()
The resulting plot shows three clearly separated clouds of points — one for each topic! Dots of the same colour belong to the same cluster. Nearby dots are semantically similar documents. 🎨
The Hero Tool — BERTopic (All Steps in One Package) 🚀
What if you could run the entire pipeline — embeddings, dimensionality reduction, clustering, and topic labelling — in just 5 lines of code?
That's exactly what BERTopic does. It is the most popular, actively maintained, and production-ready topic modelling library in the Python ecosystem. Built and maintained by Maarten Grootendorst, it packages the entire modern pipeline into a clean, modular API.
BERTopic — Minimal Working Example
from bertopic import BERTopic
# Load your documents
documents = [
"My credit card was charged twice for the same order.",
"The app keeps crashing when I open the settings menu.",
"Where is my package? It has been 10 days since I ordered.",
"I was billed incorrectly on my last invoice.",
"The mobile app freezes every time I try to log in.",
"My shipment still hasn't arrived after two weeks.",
"There is a duplicate charge on my bank statement.",
"The software crashes immediately after launch.",
"Why was I charged a cancellation fee I did not agree to?",
"My order tracking shows delivered but I received nothing.",
"The application throws an error when saving settings.",
"I need a refund for the double payment on my account.",
]
# Create and fit the model — that's it!
topic_model = BERTopic(language="english", calculate_probabilities=True, verbose=True)
topics, probabilities = topic_model.fit_transform(documents)
# See what topics were discovered
topic_info = topic_model.get_topic_info()
print(topic_info[['Topic', 'Count', 'Name']])
Output:
Topic Count Name
0 -1 0 -1_noise
1 0 4 0_charged_billing_bank_invoice
2 1 4 1_app_crashes_error_software
3 2 4 2_package_shipment_order_delivered
Four lines of real ML work done automatically! 🎯
BERTopic — Getting Topic Keywords and Representative Documents
# Get the top keywords for each topic
for topic_id in range(3):
keywords = topic_model.get_topic(topic_id)
print(f"\nTopic {topic_id}:")
for word, score in keywords[:5]:
print(f" '{word}' (score: {score:.4f})")
# Get the most representative document for each topic
representative_docs = topic_model.get_representative_docs()
for topic_id, docs in representative_docs.items():
print(f"\nTopic {topic_id} representative:")
print(f" → {docs[0]}")
Output:
Topic 0:
'charged' (score: 0.4821)
'billing' (score: 0.4103)
'bank' (score: 0.3892)
'invoice' (score: 0.3541)
'duplicate' (score: 0.3201)
Topic 1:
'app' (score: 0.5021)
'crashes' (score: 0.4782)
'error' (score: 0.4101)
'software' (score: 0.3901)
'freezes' (score: 0.3672)
Topic 2:
'package' (score: 0.4932)
'shipment' (score: 0.4621)
'order' (score: 0.4210)
'delivered' (score: 0.3891)
'arrived' (score: 0.3701)
Topic 0 representative:
→ There is a duplicate charge on my bank statement.
Topic 1 representative:
→ The software crashes immediately after launch.
Topic 2 representative:
→ My order tracking shows delivered but I received nothing.
BERTopic — Built-In Visualisations
# Interactive topic overview (in Jupyter / Colab)
topic_model.visualize_topics()
# Barchart of top keywords per topic
topic_model.visualize_barchart(top_n_topics=5)
# Documents projected in 2D, coloured by topic
topic_model.visualize_documents(documents)
# Topic similarity heatmap
topic_model.visualize_heatmap()
# Hierarchical topic tree
topic_model.visualize_hierarchy()
.show() or save with .write_html("clusters.html")
to view in a browser.
BERTopic — Advanced Configuration (Swapping Components)
One of BERTopic's biggest strengths is its modular design. You can swap out any component with a different implementation:
from bertopic import BERTopic
from bertopic.representation import KeyBERTInspired, OpenAI
from sentence_transformers import SentenceTransformer
from umap import UMAP
import hdbscan
# Step 1: Custom embedding model
embedding_model = SentenceTransformer("all-mpnet-base-v2") # higher accuracy
# Step 2: Custom UMAP settings
umap_model = UMAP(
n_neighbors=15,
n_components=5,
min_dist=0.0,
metric='cosine',
random_state=42
)
# Step 3: Custom HDBSCAN settings
hdbscan_model = hdbscan.HDBSCAN(
min_cluster_size=10,
min_samples=5,
metric='euclidean',
prediction_data=True
)
# Step 4: LLM-powered topic representation
representation_model = OpenAI(
client=client, # your OpenAI client
model="gpt-4o-mini",
chat=True
)
# Assemble the full custom pipeline
topic_model = BERTopic(
embedding_model=embedding_model,
umap_model=umap_model,
hdbscan_model=hdbscan_model,
representation_model=representation_model,
top_n_words=10,
verbose=True
)
topics, probs = topic_model.fit_transform(documents)
This gives you full control over every stage of the pipeline while still using BERTopic's clean API and visualisation tools! 💪
Complete Tool Reference — Everything You Need
Here is the full landscape of tools used in topic clustering pipelines, from embedding to visualisation:
Embedding Models
-
sentence-transformers (
pip install sentence-transformers) → The go-to library for text embeddings. Dozens of pre-trained models for English, multilingual, domain-specific use cases. Models hosted on Hugging Face Hub. -
OpenAI Embeddings API (
text-embedding-3-small,text-embedding-3-large) → Extremely high-quality embeddings via API. No local GPU required. Best quality but costs money per token. - Cohere Embed API → Strong multilingual embedding API. Competitive with OpenAI for many clustering tasks.
- FastText → Lightweight, fast word-level embeddings. Good for CPU-only environments with very large datasets.
Dimensionality Reduction Tools
-
UMAP (
pip install umap-learn) → State-of-the-art for pre-clustering reduction and visualisation. Preserves local and global structure. Recommended default choice. -
t-SNE (
sklearn.manifold.TSNE) → Classic visualisation tool. Excellent at revealing local cluster structure. Slow on large datasets and not suitable for out-of-sample projection. Use for visualisation only, not for the clustering step itself. -
PCA (
sklearn.decomposition.PCA) → Fast, linear, deterministic. Good first pass on large datasets. Less effective than UMAP for non-linear text embedding spaces.
Clustering Algorithms
-
HDBSCAN (
pip install hdbscan) → Best choice for text topic clustering. No K required, handles noise, works well with UMAP-reduced embeddings. -
K-Means (
sklearn.cluster.KMeans) → Fast and scalable. Use when you know K or need consistent cluster boundaries. Good for large datasets where HDBSCAN is too slow. -
Agglomerative Clustering (
sklearn.cluster.AgglomerativeClustering) → Produces topic hierarchies (dendrograms). Use when you want to explore topics at multiple levels of granularity. -
Spectral Clustering (
sklearn.cluster.SpectralClustering) → Good for non-convex clusters. More expensive than K-Means. Rarely first choice for text.
Topic Modelling Frameworks
-
BERTopic (
pip install bertopic) → The current gold standard for modern neural topic modelling. Modular, actively maintained, production-ready. Start here for any new project. -
Top2Vec (
pip install top2vec) → Earlier neural topic model. Simple API. Less flexible than BERTopic but still useful for quick exploration. -
LDA (Latent Dirichlet Allocation) (via
gensimorsklearn) → The traditional probabilistic topic model. Interpretable but does not use neural embeddings. Produces weaker topic quality than BERTopic on modern datasets. Use only when you need probabilistic topic distributions rather than hard cluster assignments. -
NMF (Non-negative Matrix Factorisation) (
sklearn.decomposition.NMF) → Classic TF-IDF based topic model. Fast and interpretable. Still used in some industry pipelines where simplicity is valued over accuracy.
Keyword Extraction Tools
-
KeyBERT (
pip install keybert) → Extracts keywords using BERT embeddings. Far more semantically accurate than TF-IDF keyword extraction. Integrates natively into BERTopic's representation step. -
YAKE (
pip install yake) → Fast, unsupervised, language-independent keyword extraction. No pre-trained model needed. Good for multilingual pipelines.
Visualisation Tools
-
Plotly (
pip install plotly) → Interactive charts used internally by BERTopic. Excellent for interactive cluster exploration in notebooks. -
Matplotlib (
pip install matplotlib) → Static chart generation. Good for publication-quality figures and saving PNG/PDF outputs. -
Datamapplot (
pip install datamapplot) → Beautiful static and interactive 2D maps of clusters. Increasingly popular for visualising large embedding spaces. Works natively with BERTopic outputs.
Topic Clustering in a Real MLOps Pipeline
Let's look at how topic clustering slots into a real, end-to-end MLOps workflow — not just as a one-off analysis, but as a continuous, automated process.
🔄 Topic Clustering in an LLM MLOps Pipeline
1️⃣ Data Ingestion
Raw text arrives: user queries, documents, support tickets, feedback
2️⃣ Topic Clustering
BERTopic assigns each document to a topic cluster automatically
3️⃣ Monitoring
Track cluster size over time — sudden growth = emerging topic or drift
4️⃣ Alert & Action
New clusters trigger retraining or prompt updates. Noise clusters trigger review.
Use Case 1 — LLM Dataset Coverage Audit
Before fine-tuning an LLM, cluster your instruction dataset. Check whether the topic distribution matches your intended use case.
from bertopic import BERTopic
import pandas as pd
# Load your instruction dataset
df = pd.read_json("instructions.json")
# Cluster all instructions
topic_model = BERTopic(verbose=False)
topics, _ = topic_model.fit_transform(df["instruction"].tolist())
# Add topic info back to dataframe
df["topic"] = topics
df["topic_name"] = df["topic"].map(
lambda t: topic_model.get_topic_info().set_index("Topic").loc[t, "Name"]
if t != -1 else "Noise"
)
# Check distribution
coverage = df["topic_name"].value_counts(normalize=True) * 100
print(coverage.head(10))
Output (example):
Python coding tasks 42.3%
SQL database queries 18.1%
Text summarisation 14.7%
Creative writing 8.2%
Math reasoning 6.1%
Data analysis 5.9%
Other topics 4.7%
42% coding tasks! If you're building a general assistant, this dataset will produce a model heavily biased towards coding. You now know exactly where to add more data. 🔍
Use Case 2 — Production Concept Drift Detection
from bertopic import BERTopic
from collections import Counter
import datetime
# Load the fitted model from training time
topic_model = BERTopic.load("production_topic_model")
# Incoming production queries (last 24 hours)
new_queries = load_recent_queries(hours=24)
# Assign topics to new queries
new_topics, _ = topic_model.transform(new_queries)
topic_counts = Counter(new_topics)
# Flag any new topic (-1 noise) that exceeds threshold
noise_count = topic_counts.get(-1, 0)
noise_rate = noise_count / len(new_queries)
if noise_rate > 0.15: # more than 15% of queries don't match known topics
send_alert(
title="⚠️ Concept Drift Detected",
message=f"{noise_rate:.1%} of today's queries don't match any known topic. "
f"Model may need retraining. Review noise cluster samples."
)
print(f"ALERT: {noise_rate:.1%} noise rate detected at {datetime.datetime.now()}")
topic_model.transform() (not fit_transform()) on new incoming data.
This way, new documents are assigned to the same topic space as your training data,
making drift detection meaningful and consistent.
# Save the trained model
topic_model.save("production_topic_model")
# Later, load and use for inference only
loaded_model = BERTopic.load("production_topic_model")
topics, probs = loaded_model.transform(new_documents)
Common Mistakes and How to Avoid Them 🪲
- Running clustering directly on raw 384D embeddings → Always apply UMAP first. High-dimensional distance metrics are unreliable and clustering results will be poor.
-
Ignoring the noise cluster →
HDBSCAN's
-1cluster is valuable signal, not garbage. Always inspect it — it often contains the most interesting, emerging topics. -
Using random_state inconsistently →
Set
random_state=42everywhere. UMAP, HDBSCAN, and K-Means are all stochastic. Without a fixed seed, you get different clusters every run — impossible to compare results. - Choosing K blindly in K-Means → Always use the Elbow Method or Silhouette Score to guide your choice. Picking K=10 because it "sounds right" is not a strategy.
- Using LDA on short texts → LDA assumes each document has many words to estimate topic proportions from. On short texts like tweets, search queries, or product descriptions, it performs poorly. Use BERTopic instead.
- Not versioning topic models → In an MLOps pipeline, always version your trained BERTopic model alongside your training data. If you retrain without saving the old model, drift detection becomes impossible to calibrate.
Evaluating Your Topic Model
How do you know if your clusters are actually good? Here are the key metrics to compute:
Silhouette Score — Cluster Compactness
from sklearn.metrics import silhouette_score
# Exclude noise points (label -1) from evaluation
non_noise_mask = [l != -1 for l in labels]
filtered_embeddings = reduced_embeddings[non_noise_mask]
filtered_labels = [l for l in labels if l != -1]
score = silhouette_score(filtered_embeddings, filtered_labels, metric='euclidean')
print(f"Silhouette Score: {score:.4f}")
# Score ranges from -1 to 1
# Above 0.5 = strong clustering
# 0.25-0.5 = reasonable clustering
# Below 0.25 = weak clustering — consider adjusting parameters
Output:
Silhouette Score: 0.6823
0.68 is a strong result! Our three clusters are well-separated and internally compact. ✅
Topic Diversity — Are Topics Distinct?
def compute_topic_diversity(topic_model, top_n=10):
"""
Topic Diversity measures how many unique words appear across all topics.
High diversity (close to 1.0) means topics are well-differentiated.
Low diversity means topics share too many words and may overlap.
"""
all_words = []
unique_words = set()
for topic_id in topic_model.get_topic_info()["Topic"]:
if topic_id == -1:
continue
keywords = [word for word, _ in topic_model.get_topic(topic_id)[:top_n]]
all_words.extend(keywords)
unique_words.update(keywords)
diversity = len(unique_words) / len(all_words) if all_words else 0
print(f"Topic Diversity: {diversity:.4f} ({len(unique_words)} unique / {len(all_words)} total)")
return diversity
diversity = compute_topic_diversity(topic_model)
Output:
Topic Diversity: 0.9333 (28 unique / 30 total)
93% diversity — almost every keyword is unique to one topic. Excellent topic separation! 🎯
Quick Summary 📝
What we learned today:
- What Topic Clustering Is → Unsupervised grouping of text documents by theme, requiring no pre-labelled data
- Why It Matters in MLOps → Data understanding, labelling efficiency, model monitoring, and concept drift detection
- The Pipeline → Preprocessing → Embeddings → Dimensionality Reduction → Clustering → Labelling → Visualisation
- Embeddings → Use
sentence-transformersfor local, free embeddings. OpenAI for maximum quality. - Dimensionality Reduction → UMAP is the gold standard. Reduce to 5D for clustering, 2D for visualisation.
- Clustering → HDBSCAN for unknown K, K-Means for known K, Agglomerative for topic hierarchies
- BERTopic → The one library that wraps the entire modern pipeline with a clean API and beautiful visualisations
- MLOps Integration → Save models, use
transform()for inference, monitor noise rates for drift - Evaluation → Silhouette Score for cluster quality, Topic Diversity for topic distinctiveness
Topic Clustering is one of those techniques that, once you understand it, you start seeing opportunities to apply it everywhere — in your data pipelines, your model monitoring, your LLM datasets, and your production systems 🧠✨
Comments
Post a Comment