Skip to main content

Topic Clustering in Machine Learning: Group and Organize Text Data with Unsupervised Learning

Calculating read time…

Imagine you receive 100,000 customer support tickets every month. No human team can read all of them. But what if a machine could automatically group them into buckets like "Billing Issues", "Technical Errors", "Shipping Delays" — without anyone telling it what categories to look for?






That is exactly what Topic Clustering does. It finds hidden structure in large collections of text — automatically, intelligently, and at scale. In modern MLOps pipelines, it powers everything from data labelling to model monitoring to search engine organisation.

💡 Think of it like: A super-smart librarian who reads every book in a warehouse, then organises them into themed shelves — without being told what the themes are in advance. 📚

What Exactly Is Topic Clustering?

Topic Clustering is an unsupervised machine learning technique that groups a collection of text documents into clusters, where each cluster represents a common underlying theme or topic.

The key word here is unsupervised.

  • Supervised learning → You label the data first, then train a model to predict those labels
  • Unsupervised learning → You give the model raw, unlabelled data and it discovers structure on its own
  • Topic Clustering is unsupervised — no labelling required before you start

This makes it incredibly powerful for exploring datasets you've never seen before — or for continuously monitoring incoming data streams in a live MLOps system.

Why Does Topic Clustering Matter in MLOps?

MLOps is the practice of building, deploying, and maintaining machine learning models in production. Topic Clustering plays a critical role across several stages of this lifecycle:

  • Data Understanding → Before training any model, you need to know what topics your dataset covers. Topic clustering gives you an instant overview.
  • Data Labelling → Cluster first, label one representative per cluster, then propagate labels. This reduces human labelling effort by 90%+.
  • Data Quality → Clusters with very few members often signal noisy or out-of-distribution data that shouldn't be in your training set.
  • Model Monitoring → After deployment, cluster incoming requests to detect if users are asking about topics the model was never trained on (concept drift).
  • Instruction Dataset Creation → In LLM fine-tuning, cluster your instructions to ensure diverse topic coverage and avoid accidentally training on 80% coding samples and 20% everything else.
🟢 DO: Run topic clustering on your dataset before any model training. Understanding the topic landscape of your data is one of the most valuable 30-minute investments you can make. It often reveals surprising imbalances and gaps that would silently hurt model quality.

The Topic Clustering Pipeline — Step by Step

Topic clustering is not a single algorithm. It's a pipeline of steps that transform raw text into meaningful, interpretable clusters. Here is the complete flow:

📐 The Topic Clustering Pipeline

📄 Raw Text
→
🧹 Preprocessing
→
🔢 Embeddings
→
📉 Dimensionality Reduction
→
🔵 Clustering
→
🏷️ Topic Labelling
→
📊 Visualisation

Let's walk through each stage with real code and real examples!

Step 1 — Raw Text Collection & Preprocessing

Before any machine learning can happen, you need clean text. Raw text from the real world is messy — it contains HTML tags, emojis, duplicate whitespace, special characters, and language noise.

Preprocessing transforms messy raw text into clean, consistent input that the next steps can work with reliably.

What Preprocessing Typically Includes

  • Lowercasing → "Machine Learning" and "machine learning" should be treated as the same phrase
  • Removing HTML tags → Strip <b>, <p> and other markup from web-scraped text
  • Removing special characters → Punctuation and symbols that don't carry semantic meaning
  • Stripping extra whitespace → Multiple spaces, newlines, tabs collapsed to single spaces
  • Language filtering → Keep only samples in your target language if needed
🟡 TIP: With modern sentence embedding models (like those from Sentence-Transformers), aggressive preprocessing like stop word removal or stemming is not recommended. These models handle language nuance natively. Simple cleaning — remove HTML, fix encoding, strip excess whitespace — is usually enough.

Preprocessing Code Example

import re

def clean_text(text: str) -> str:
    # Remove HTML tags
    text = re.sub(r'<[^>]+>', ' ', text)

    # Remove URLs
    text = re.sub(r'http\S+|www\S+', ' ', text)

    # Remove special characters (keep letters, digits, spaces)
    text = re.sub(r'[^a-zA-Z0-9\s]', ' ', text)

    # Collapse multiple whitespace into one
    text = re.sub(r'\s+', ' ', text)

    # Lowercase and strip leading/trailing whitespace
    return text.strip().lower()


# Example usage
raw = "

We're having issues with our ORDER #12345!! Contact: help@example.com

" print(clean_text(raw))

Output:

we are having issues with our order  12345   contact  help example com

Clean, simple, ready for embedding! 🎯

Step 2 — Text Embeddings (Turning Words into Numbers)

Computers don't understand words — they understand numbers. Embeddings convert text into dense numerical vectors (lists of numbers) where similar texts produce vectors that are close together in mathematical space.

💡 Think of it like: A map of the world 🗺️. Cities in the same country are geographically close. Similarly, documents about the same topic produce embeddings that are close to each other in a high-dimensional mathematical space.

Traditional vs. Modern Embedding Approaches

  • TF-IDF (Traditional) → Represents each document as a sparse vector based on word frequencies. Fast and interpretable, but misses meaning and context entirely. "Bank" (river) and "Bank" (finance) get the same representation.
  • Word2Vec / GloVe (2013–2016) → Learn word-level meaning from co-occurrence patterns. Better than TF-IDF, but still operate at the word level. A sentence embedding is just an average of its word embeddings — losing sentence structure.
  • Sentence Transformers (Modern) → Encode entire sentences and paragraphs into rich, context-aware embeddings. Two sentences with the same meaning but different words will have very similar embeddings. This is the state-of-the-art approach for topic clustering.

Generating Embeddings with Sentence-Transformers

from sentence_transformers import SentenceTransformer

# Load a pre-trained embedding model
# 'all-MiniLM-L6-v2' is fast, small, and surprisingly powerful
model = SentenceTransformer('all-MiniLM-L6-v2')

documents = [
    "My credit card was charged twice for the same order.",
    "The app keeps crashing when I open the settings menu.",
    "Where is my package? It has been 10 days since I ordered.",
    "I was billed incorrectly on my last invoice.",
    "The mobile app freezes every time I try to log in.",
    "My shipment still hasn't arrived after two weeks.",
    "There is a duplicate charge on my bank statement.",
    "The software crashes immediately after launch.",
]

# Generate embeddings — each document becomes a 384-dimensional vector
embeddings = model.encode(documents, show_progress_bar=True)

print(f"Number of documents: {len(documents)}")
print(f"Embedding shape: {embeddings.shape}")
print(f"First embedding (first 5 values): {embeddings[0][:5]}")

Output:

Number of documents: 8
Embedding shape: (8, 384)
First embedding (first 5 values): [ 0.0234 -0.1123  0.0891  0.2341 -0.0567]

Each document is now a list of 384 numbers that capture its meaning! Documents about billing will have similar numbers. Documents about crashes will have similar numbers. The machine can now measure similarity mathematically. 🔢

🟢 DO: For most topic clustering tasks, start with all-MiniLM-L6-v2 (fast, 80MB) or all-mpnet-base-v2 (more accurate, 420MB). For multilingual data, use paraphrase-multilingual-MiniLM-L12-v2. All are free and available on the Hugging Face Hub.
🔴 DON'T: Use TF-IDF for topic clustering on modern datasets. It treats "car accident" and "vehicle collision" as completely different topics because they share no words — even though they mean the same thing. Sentence embeddings understand meaning, not just word overlap.

Step 3 — Dimensionality Reduction

Our embeddings have 384 dimensions. Running clustering directly on 384-dimensional vectors is problematic — this is known as the "Curse of Dimensionality".

In very high dimensions, all points start to look equally far apart from each other. Clustering algorithms lose their effectiveness because distance measurements stop being meaningful.

💡 Think of it like: Trying to group people by similarity using 384 different personality traits simultaneously. It's overwhelming and ineffective. But if you reduce it to the top 5 most distinguishing traits, grouping becomes obvious and accurate.

Two Popular Dimensionality Reduction Tools

  • PCA (Principal Component Analysis) → Classic, fast, linear method. Preserves global structure well. Good for initial exploration. Can be lossy if your data has complex non-linear structure.
  • UMAP (Uniform Manifold Approximation and Projection) → Modern, non-linear method. Preserves both local and global structure far better than PCA. The current gold standard for dimensionality reduction before clustering. Also produces beautiful 2D visualisations.

Applying UMAP to Our Embeddings

import umap

# Reduce 384 dimensions → 5 dimensions for clustering
# (Keep 5D for clustering, reduce to 2D separately just for visualisation)
reducer = umap.UMAP(
    n_components=5,       # target dimensions for clustering
    n_neighbors=15,       # controls local vs global structure balance
    min_dist=0.0,         # tighter clusters (0.0 is best for clustering, not viz)
    metric='cosine',      # cosine similarity works best for text embeddings
    random_state=42       # for reproducibility
)

reduced_embeddings = reducer.fit_transform(embeddings)

print(f"Original shape: {embeddings.shape}")
print(f"Reduced shape:  {reduced_embeddings.shape}")

Output:

Original shape: (8, 384)
Reduced shape:  (8, 5)
🟡 TIP: Always set random_state=42 in UMAP (and all other stochastic algorithms). Without it, you get different cluster assignments every run — which makes debugging, reproducibility, and comparing runs nearly impossible in a production MLOps pipeline.

Step 4 — Clustering Algorithms

Now we have our reduced embeddings. It's time to group similar documents together using a clustering algorithm.

There are three main families of clustering algorithms used in topic clustering. Each has different strengths and is suited to different situations.

Algorithm 1 — K-Means Clustering

K-Means partitions your data into exactly K clusters (you specify K upfront). It works by placing K "centroids" (imaginary cluster centres) and assigning each point to the nearest centroid, then adjusting the centroids until the assignments stabilise.

💡 Analogy: You drop K pins on a map, then everyone travels to the nearest pin. You then move each pin to the centre of its group. Repeat until no one moves. 📍

When to use K-Means:

  • You know roughly how many topics you expect
  • Your dataset is large (K-Means scales well)
  • Your clusters are roughly similar in size
from sklearn.cluster import KMeans

# Cluster our reduced embeddings into 3 clusters
kmeans = KMeans(
    n_clusters=3,
    n_init=10,        # run 10 times with different initialisations, keep best result
    random_state=42
)

kmeans.fit(reduced_embeddings)
labels = kmeans.labels_

# Print each document with its assigned cluster
for doc, label in zip(documents, labels):
    print(f"Cluster {label}: {doc[:60]}...")

Output:

Cluster 0: My credit card was charged twice for the same order....
Cluster 1: The app keeps crashing when I open the settings menu....
Cluster 2: Where is my package? It has been 10 days since I order...
Cluster 0: I was billed incorrectly on my last invoice....
Cluster 1: The mobile app freezes every time I try to log in....
Cluster 2: My shipment still hasn't arrived after two weeks....
Cluster 0: There is a duplicate charge on my bank statement....
Cluster 1: The software crashes immediately after launch....

Cluster 0 = Billing issues. Cluster 1 = App crashes. Cluster 2 = Shipping delays. It found the three real topics with zero labelling! 🎉

How to Choose K — The Elbow Method

The biggest challenge with K-Means is choosing the right K. The Elbow Method helps by plotting the "inertia" (sum of distances from each point to its cluster centre) for different values of K. The optimal K is usually at the "elbow" of the curve where inertia stops decreasing rapidly.

import matplotlib.pyplot as plt

inertias = []
k_range = range(2, 15)

for k in k_range:
    km = KMeans(n_clusters=k, n_init=10, random_state=42)
    km.fit(reduced_embeddings)
    inertias.append(km.inertia_)

plt.figure(figsize=(8, 4))
plt.plot(k_range, inertias, 'bo-', markersize=8)
plt.xlabel('Number of Clusters (K)')
plt.ylabel('Inertia')
plt.title('Elbow Method — Finding the Optimal K')
plt.grid(True, alpha=0.3)
plt.axvline(x=3, color='red', linestyle='--', label='Elbow at K=3')
plt.legend()
plt.show()

Look for the "elbow" — the point where adding more clusters gives diminishing returns! 📉

Algorithm 2 — HDBSCAN (The Modern Favourite)

HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) is the current state-of-the-art for topic clustering in NLP and LLM pipelines.

Unlike K-Means, HDBSCAN does not require you to specify the number of clusters upfront. It discovers clusters based on density — areas where documents are tightly packed together. Documents in low-density regions are classified as noise (outliers) rather than forced into a cluster.

Key advantages over K-Means:

  • No need to specify K in advance
  • Handles clusters of different sizes and shapes naturally
  • Explicitly marks outliers as noise (label = -1) instead of forcing them into a cluster
  • Much more robust on real-world, messy text data
import hdbscan

clusterer = hdbscan.HDBSCAN(
    min_cluster_size=5,    # minimum documents to form a cluster
    min_samples=3,         # controls how conservative cluster detection is
    metric='euclidean',    # use euclidean on UMAP-reduced embeddings
    cluster_selection_method='eom'   # 'eom' generally gives more clusters than 'leaf'
)

labels = clusterer.fit_predict(reduced_embeddings)

# Count how many documents ended up in each cluster
from collections import Counter
print(Counter(labels))
# -1 means 'noise' — documents that didn't fit any cluster

Output:

Counter({0: 3, 1: 3, 2: 3, -1: 2})
🟢 DO: For real-world MLOps and LLM dataset work, default to HDBSCAN over K-Means. You rarely know the "correct" number of topics in advance, and real data almost always contains outlier documents that don't belong to any clean cluster. HDBSCAN handles both situations naturally.
🔴 DON'T: Ignore the noise cluster (label = -1). In a production system, noise documents are worth inspecting manually — they often represent emerging new topics that your model hasn't seen before, which is exactly the signal you want for concept drift detection.

Algorithm 3 — Agglomerative (Hierarchical) Clustering

Hierarchical clustering builds a tree structure (called a dendrogram) showing how documents group together at different levels of similarity.

💡 Analogy: Think of a family tree 🌳. At the bottom, every person is their own node. As you move up, siblings merge into families, families into extended families, and so on. You can cut the tree at any level to get different granularities of grouping.

This is useful when you want to explore topic hierarchies — for example, "Technology" at a high level that breaks down into "Mobile Apps", "Desktop Software", and "Web Platforms" at a more detailed level.

from sklearn.cluster import AgglomerativeClustering
from scipy.cluster.hierarchy import dendrogram, linkage
import matplotlib.pyplot as plt

# Build linkage matrix for visualisation
Z = linkage(reduced_embeddings, method='ward')

plt.figure(figsize=(12, 5))
dendrogram(
    Z,
    labels=[doc[:30] + "..." for doc in documents],
    leaf_rotation=45,
    leaf_font_size=9
)
plt.title('Hierarchical Clustering Dendrogram')
plt.tight_layout()
plt.show()

# Cut the tree at 3 clusters
agg = AgglomerativeClustering(n_clusters=3, linkage='ward')
labels = agg.fit_predict(reduced_embeddings)
print(labels)

Output:

[0 1 2 0 1 2 0 1]

Step 5 — Topic Labelling

After clustering, each cluster has a number (0, 1, 2...) but no human-readable name. Topic labelling is the process of assigning a meaningful description to each cluster.

There are two approaches:

Approach 1 — Keyword Extraction (Traditional)

Extract the most distinctive words from each cluster using TF-IDF or KeyBERT. The top keywords become the cluster's label.

from sklearn.feature_extraction.text import TfidfVectorizer
import numpy as np

def get_cluster_keywords(documents, labels, n_keywords=5):
    unique_labels = sorted(set(labels))
    cluster_keywords = {}

    for label in unique_labels:
        if label == -1:
            continue  # skip noise cluster

        # Collect all documents in this cluster
        cluster_docs = [doc for doc, lbl in zip(documents, labels) if lbl == label]

        # TF-IDF across cluster documents
        vectorizer = TfidfVectorizer(stop_words='english', max_features=50)
        tfidf_matrix = vectorizer.fit_transform(cluster_docs)

        # Sum TF-IDF scores across all docs in cluster
        scores = np.array(tfidf_matrix.sum(axis=0)).flatten()
        top_indices = scores.argsort()[-n_keywords:][::-1]
        top_words = [vectorizer.get_feature_names_out()[i] for i in top_indices]

        cluster_keywords[label] = top_words
        print(f"Cluster {label}: {', '.join(top_words)}")

    return cluster_keywords

keywords = get_cluster_keywords(documents, labels)

Output:

Cluster 0: charged, billing, invoice, duplicate, bank
Cluster 1: app, crashes, freezes, software, mobile
Cluster 2: package, shipment, order, arrived, shipping

From these keywords, a human immediately reads: Cluster 0 = Billing, Cluster 1 = Tech Issues, Cluster 2 = Delivery! 📦

Approach 2 — LLM-Based Labelling (Modern & Powerful)

Send the top documents from each cluster to an LLM and ask it to generate a descriptive label. This produces much more natural, contextually aware topic names than keyword extraction alone.

from openai import OpenAI

client = OpenAI()

def label_cluster_with_llm(cluster_documents: list, n_examples: int = 5) -> str:
    # Use the first N documents as representative examples
    examples = cluster_documents[:n_examples]
    examples_text = "\n".join(f"- {doc}" for doc in examples)

    prompt = f"""You are analysing a cluster of customer support tickets.
Below are representative examples from this cluster:

{examples_text}

In 3-5 words, what is the single best topic label for this cluster?
Return ONLY the label text, nothing else."""

    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        temperature=0
    )

    return response.choices[0].message.content.strip()


# Label all clusters
from collections import defaultdict

cluster_docs = defaultdict(list)
for doc, label in zip(documents, labels):
    if label != -1:
        cluster_docs[label].append(doc)

for cluster_id, docs in sorted(cluster_docs.items()):
    label_text = label_cluster_with_llm(docs)
    print(f"Cluster {cluster_id}: '{label_text}'")

Output:

Cluster 0: 'Billing & Payment Issues'
Cluster 1: 'App Crashes & Technical Errors'
Cluster 2: 'Shipping & Delivery Delays'

Human-quality topic names, generated automatically! 🏆

🟢 DO: Combine both approaches — use keyword extraction for quick reference and LLM labelling for final, human-readable topic names. LLM labels are especially valuable when cluster topics are nuanced or domain-specific.

Step 6 — Visualisation

Numbers and labels are useful, but humans understand patterns best when they can see them. Visualising your clusters helps you quickly validate whether the clustering makes sense and communicate results to stakeholders.

2D UMAP Visualisation with Matplotlib

import matplotlib.pyplot as plt
import umap
import numpy as np

# Create a 2D reduction just for visualisation (separate from the 5D used for clustering)
viz_reducer = umap.UMAP(
    n_components=2,
    n_neighbors=15,
    min_dist=0.1,
    metric='cosine',
    random_state=42
)

coords_2d = viz_reducer.fit_transform(embeddings)

# Plot
topic_labels = {
    0: 'Billing Issues',
    1: 'App Crashes',
    2: 'Shipping Delays',
    -1: 'Noise'
}

colors = {0: '#2196F3', 1: '#4CAF50', 2: '#FF9800', -1: '#9E9E9E'}

plt.figure(figsize=(10, 7))

for cluster_id in sorted(set(labels)):
    mask = np.array(labels) == cluster_id
    plt.scatter(
        coords_2d[mask, 0],
        coords_2d[mask, 1],
        c=colors[cluster_id],
        label=topic_labels[cluster_id],
        alpha=0.8,
        s=100,
        edgecolors='white',
        linewidth=0.5
    )

plt.title('Topic Clusters — 2D UMAP Projection', fontsize=14, fontweight='bold')
plt.xlabel('UMAP Dimension 1')
plt.ylabel('UMAP Dimension 2')
plt.legend(title='Topics', loc='best')
plt.grid(True, alpha=0.2)
plt.tight_layout()
plt.savefig('topic_clusters.png', dpi=150, bbox_inches='tight')
plt.show()

The resulting plot shows three clearly separated clouds of points — one for each topic! Dots of the same colour belong to the same cluster. Nearby dots are semantically similar documents. 🎨

The Hero Tool — BERTopic (All Steps in One Package) 🚀

What if you could run the entire pipeline — embeddings, dimensionality reduction, clustering, and topic labelling — in just 5 lines of code?

That's exactly what BERTopic does. It is the most popular, actively maintained, and production-ready topic modelling library in the Python ecosystem. Built and maintained by Maarten Grootendorst, it packages the entire modern pipeline into a clean, modular API.

BERTopic — Minimal Working Example

from bertopic import BERTopic

# Load your documents
documents = [
    "My credit card was charged twice for the same order.",
    "The app keeps crashing when I open the settings menu.",
    "Where is my package? It has been 10 days since I ordered.",
    "I was billed incorrectly on my last invoice.",
    "The mobile app freezes every time I try to log in.",
    "My shipment still hasn't arrived after two weeks.",
    "There is a duplicate charge on my bank statement.",
    "The software crashes immediately after launch.",
    "Why was I charged a cancellation fee I did not agree to?",
    "My order tracking shows delivered but I received nothing.",
    "The application throws an error when saving settings.",
    "I need a refund for the double payment on my account.",
]

# Create and fit the model — that's it!
topic_model = BERTopic(language="english", calculate_probabilities=True, verbose=True)
topics, probabilities = topic_model.fit_transform(documents)

# See what topics were discovered
topic_info = topic_model.get_topic_info()
print(topic_info[['Topic', 'Count', 'Name']])

Output:

   Topic  Count                           Name
0     -1      0                         -1_noise
1      0      4    0_charged_billing_bank_invoice
2      1      4    1_app_crashes_error_software
3      2      4    2_package_shipment_order_delivered

Four lines of real ML work done automatically! 🎯

BERTopic — Getting Topic Keywords and Representative Documents

# Get the top keywords for each topic
for topic_id in range(3):
    keywords = topic_model.get_topic(topic_id)
    print(f"\nTopic {topic_id}:")
    for word, score in keywords[:5]:
        print(f"  '{word}' (score: {score:.4f})")

# Get the most representative document for each topic
representative_docs = topic_model.get_representative_docs()
for topic_id, docs in representative_docs.items():
    print(f"\nTopic {topic_id} representative:")
    print(f"  → {docs[0]}")

Output:

Topic 0:
  'charged' (score: 0.4821)
  'billing' (score: 0.4103)
  'bank' (score: 0.3892)
  'invoice' (score: 0.3541)
  'duplicate' (score: 0.3201)

Topic 1:
  'app' (score: 0.5021)
  'crashes' (score: 0.4782)
  'error' (score: 0.4101)
  'software' (score: 0.3901)
  'freezes' (score: 0.3672)

Topic 2:
  'package' (score: 0.4932)
  'shipment' (score: 0.4621)
  'order' (score: 0.4210)
  'delivered' (score: 0.3891)
  'arrived' (score: 0.3701)

Topic 0 representative:
  → There is a duplicate charge on my bank statement.

Topic 1 representative:
  → The software crashes immediately after launch.

Topic 2 representative:
  → My order tracking shows delivered but I received nothing.

BERTopic — Built-In Visualisations

# Interactive topic overview (in Jupyter / Colab)
topic_model.visualize_topics()

# Barchart of top keywords per topic
topic_model.visualize_barchart(top_n_topics=5)

# Documents projected in 2D, coloured by topic
topic_model.visualize_documents(documents)

# Topic similarity heatmap
topic_model.visualize_heatmap()

# Hierarchical topic tree
topic_model.visualize_hierarchy()
🟡 TIP: BERTopic's visualisations are interactive Plotly charts. They work best in Jupyter notebooks or Google Colab. If you're running a script, call .show() or save with .write_html("clusters.html") to view in a browser.

BERTopic — Advanced Configuration (Swapping Components)

One of BERTopic's biggest strengths is its modular design. You can swap out any component with a different implementation:

from bertopic import BERTopic
from bertopic.representation import KeyBERTInspired, OpenAI
from sentence_transformers import SentenceTransformer
from umap import UMAP
import hdbscan

# Step 1: Custom embedding model
embedding_model = SentenceTransformer("all-mpnet-base-v2")  # higher accuracy

# Step 2: Custom UMAP settings
umap_model = UMAP(
    n_neighbors=15,
    n_components=5,
    min_dist=0.0,
    metric='cosine',
    random_state=42
)

# Step 3: Custom HDBSCAN settings
hdbscan_model = hdbscan.HDBSCAN(
    min_cluster_size=10,
    min_samples=5,
    metric='euclidean',
    prediction_data=True
)

# Step 4: LLM-powered topic representation
representation_model = OpenAI(
    client=client,             # your OpenAI client
    model="gpt-4o-mini",
    chat=True
)

# Assemble the full custom pipeline
topic_model = BERTopic(
    embedding_model=embedding_model,
    umap_model=umap_model,
    hdbscan_model=hdbscan_model,
    representation_model=representation_model,
    top_n_words=10,
    verbose=True
)

topics, probs = topic_model.fit_transform(documents)

This gives you full control over every stage of the pipeline while still using BERTopic's clean API and visualisation tools! 💪

Complete Tool Reference — Everything You Need

Here is the full landscape of tools used in topic clustering pipelines, from embedding to visualisation:

Embedding Models

  • sentence-transformers (pip install sentence-transformers) → The go-to library for text embeddings. Dozens of pre-trained models for English, multilingual, domain-specific use cases. Models hosted on Hugging Face Hub.
  • OpenAI Embeddings API (text-embedding-3-small, text-embedding-3-large) → Extremely high-quality embeddings via API. No local GPU required. Best quality but costs money per token.
  • Cohere Embed API → Strong multilingual embedding API. Competitive with OpenAI for many clustering tasks.
  • FastText → Lightweight, fast word-level embeddings. Good for CPU-only environments with very large datasets.

Dimensionality Reduction Tools

  • UMAP (pip install umap-learn) → State-of-the-art for pre-clustering reduction and visualisation. Preserves local and global structure. Recommended default choice.
  • t-SNE (sklearn.manifold.TSNE) → Classic visualisation tool. Excellent at revealing local cluster structure. Slow on large datasets and not suitable for out-of-sample projection. Use for visualisation only, not for the clustering step itself.
  • PCA (sklearn.decomposition.PCA) → Fast, linear, deterministic. Good first pass on large datasets. Less effective than UMAP for non-linear text embedding spaces.

Clustering Algorithms

  • HDBSCAN (pip install hdbscan) → Best choice for text topic clustering. No K required, handles noise, works well with UMAP-reduced embeddings.
  • K-Means (sklearn.cluster.KMeans) → Fast and scalable. Use when you know K or need consistent cluster boundaries. Good for large datasets where HDBSCAN is too slow.
  • Agglomerative Clustering (sklearn.cluster.AgglomerativeClustering) → Produces topic hierarchies (dendrograms). Use when you want to explore topics at multiple levels of granularity.
  • Spectral Clustering (sklearn.cluster.SpectralClustering) → Good for non-convex clusters. More expensive than K-Means. Rarely first choice for text.

Topic Modelling Frameworks

  • BERTopic (pip install bertopic) → The current gold standard for modern neural topic modelling. Modular, actively maintained, production-ready. Start here for any new project.
  • Top2Vec (pip install top2vec) → Earlier neural topic model. Simple API. Less flexible than BERTopic but still useful for quick exploration.
  • LDA (Latent Dirichlet Allocation) (via gensim or sklearn) → The traditional probabilistic topic model. Interpretable but does not use neural embeddings. Produces weaker topic quality than BERTopic on modern datasets. Use only when you need probabilistic topic distributions rather than hard cluster assignments.
  • NMF (Non-negative Matrix Factorisation) (sklearn.decomposition.NMF) → Classic TF-IDF based topic model. Fast and interpretable. Still used in some industry pipelines where simplicity is valued over accuracy.

Keyword Extraction Tools

  • KeyBERT (pip install keybert) → Extracts keywords using BERT embeddings. Far more semantically accurate than TF-IDF keyword extraction. Integrates natively into BERTopic's representation step.
  • YAKE (pip install yake) → Fast, unsupervised, language-independent keyword extraction. No pre-trained model needed. Good for multilingual pipelines.

Visualisation Tools

  • Plotly (pip install plotly) → Interactive charts used internally by BERTopic. Excellent for interactive cluster exploration in notebooks.
  • Matplotlib (pip install matplotlib) → Static chart generation. Good for publication-quality figures and saving PNG/PDF outputs.
  • Datamapplot (pip install datamapplot) → Beautiful static and interactive 2D maps of clusters. Increasingly popular for visualising large embedding spaces. Works natively with BERTopic outputs.

Topic Clustering in a Real MLOps Pipeline

Let's look at how topic clustering slots into a real, end-to-end MLOps workflow — not just as a one-off analysis, but as a continuous, automated process.

🔄 Topic Clustering in an LLM MLOps Pipeline

1️⃣ Data Ingestion

Raw text arrives: user queries, documents, support tickets, feedback

→

2️⃣ Topic Clustering

BERTopic assigns each document to a topic cluster automatically

→

3️⃣ Monitoring

Track cluster size over time — sudden growth = emerging topic or drift

→

4️⃣ Alert & Action

New clusters trigger retraining or prompt updates. Noise clusters trigger review.

Use Case 1 — LLM Dataset Coverage Audit

Before fine-tuning an LLM, cluster your instruction dataset. Check whether the topic distribution matches your intended use case.

from bertopic import BERTopic
import pandas as pd

# Load your instruction dataset
df = pd.read_json("instructions.json")

# Cluster all instructions
topic_model = BERTopic(verbose=False)
topics, _ = topic_model.fit_transform(df["instruction"].tolist())

# Add topic info back to dataframe
df["topic"] = topics
df["topic_name"] = df["topic"].map(
    lambda t: topic_model.get_topic_info().set_index("Topic").loc[t, "Name"]
    if t != -1 else "Noise"
)

# Check distribution
coverage = df["topic_name"].value_counts(normalize=True) * 100
print(coverage.head(10))

Output (example):

Python coding tasks          42.3%
SQL database queries         18.1%
Text summarisation           14.7%
Creative writing              8.2%
Math reasoning                6.1%
Data analysis                 5.9%
Other topics                  4.7%

42% coding tasks! If you're building a general assistant, this dataset will produce a model heavily biased towards coding. You now know exactly where to add more data. 🔍

Use Case 2 — Production Concept Drift Detection

from bertopic import BERTopic
from collections import Counter
import datetime

# Load the fitted model from training time
topic_model = BERTopic.load("production_topic_model")

# Incoming production queries (last 24 hours)
new_queries = load_recent_queries(hours=24)

# Assign topics to new queries
new_topics, _ = topic_model.transform(new_queries)

topic_counts = Counter(new_topics)

# Flag any new topic (-1 noise) that exceeds threshold
noise_count = topic_counts.get(-1, 0)
noise_rate = noise_count / len(new_queries)

if noise_rate > 0.15:    # more than 15% of queries don't match known topics
    send_alert(
        title="⚠️ Concept Drift Detected",
        message=f"{noise_rate:.1%} of today's queries don't match any known topic. "
                f"Model may need retraining. Review noise cluster samples."
    )
    print(f"ALERT: {noise_rate:.1%} noise rate detected at {datetime.datetime.now()}")
🟡 TIP — Saving and Loading BERTopic Models: In production, always save your trained topic model after fitting on your training data. Then use topic_model.transform() (not fit_transform()) on new incoming data. This way, new documents are assigned to the same topic space as your training data, making drift detection meaningful and consistent.
# Save the trained model
topic_model.save("production_topic_model")

# Later, load and use for inference only
loaded_model = BERTopic.load("production_topic_model")
topics, probs = loaded_model.transform(new_documents)

Common Mistakes and How to Avoid Them 🪲

  • Running clustering directly on raw 384D embeddings → Always apply UMAP first. High-dimensional distance metrics are unreliable and clustering results will be poor.
  • Ignoring the noise cluster → HDBSCAN's -1 cluster is valuable signal, not garbage. Always inspect it — it often contains the most interesting, emerging topics.
  • Using random_state inconsistently → Set random_state=42 everywhere. UMAP, HDBSCAN, and K-Means are all stochastic. Without a fixed seed, you get different clusters every run — impossible to compare results.
  • Choosing K blindly in K-Means → Always use the Elbow Method or Silhouette Score to guide your choice. Picking K=10 because it "sounds right" is not a strategy.
  • Using LDA on short texts → LDA assumes each document has many words to estimate topic proportions from. On short texts like tweets, search queries, or product descriptions, it performs poorly. Use BERTopic instead.
  • Not versioning topic models → In an MLOps pipeline, always version your trained BERTopic model alongside your training data. If you retrain without saving the old model, drift detection becomes impossible to calibrate.
🔴 DON'T: Evaluate your topic model purely by eyeballing the clusters. Always compute objective metrics too — Silhouette Score for compactness/separation, Topic Coherence (C_v) for human interpretability, and Topic Diversity for measuring how distinct your clusters are from each other.

Evaluating Your Topic Model

How do you know if your clusters are actually good? Here are the key metrics to compute:

Silhouette Score — Cluster Compactness

from sklearn.metrics import silhouette_score

# Exclude noise points (label -1) from evaluation
non_noise_mask = [l != -1 for l in labels]
filtered_embeddings = reduced_embeddings[non_noise_mask]
filtered_labels = [l for l in labels if l != -1]

score = silhouette_score(filtered_embeddings, filtered_labels, metric='euclidean')
print(f"Silhouette Score: {score:.4f}")
# Score ranges from -1 to 1
# Above 0.5 = strong clustering
# 0.25-0.5 = reasonable clustering
# Below 0.25 = weak clustering — consider adjusting parameters

Output:

Silhouette Score: 0.6823

0.68 is a strong result! Our three clusters are well-separated and internally compact. ✅

Topic Diversity — Are Topics Distinct?

def compute_topic_diversity(topic_model, top_n=10):
    """
    Topic Diversity measures how many unique words appear across all topics.
    High diversity (close to 1.0) means topics are well-differentiated.
    Low diversity means topics share too many words and may overlap.
    """
    all_words = []
    unique_words = set()

    for topic_id in topic_model.get_topic_info()["Topic"]:
        if topic_id == -1:
            continue
        keywords = [word for word, _ in topic_model.get_topic(topic_id)[:top_n]]
        all_words.extend(keywords)
        unique_words.update(keywords)

    diversity = len(unique_words) / len(all_words) if all_words else 0
    print(f"Topic Diversity: {diversity:.4f} ({len(unique_words)} unique / {len(all_words)} total)")
    return diversity

diversity = compute_topic_diversity(topic_model)

Output:

Topic Diversity: 0.9333 (28 unique / 30 total)

93% diversity — almost every keyword is unique to one topic. Excellent topic separation! 🎯

Quick Summary 📝

What we learned today:

  • What Topic Clustering Is → Unsupervised grouping of text documents by theme, requiring no pre-labelled data
  • Why It Matters in MLOps → Data understanding, labelling efficiency, model monitoring, and concept drift detection
  • The Pipeline → Preprocessing → Embeddings → Dimensionality Reduction → Clustering → Labelling → Visualisation
  • Embeddings → Use sentence-transformers for local, free embeddings. OpenAI for maximum quality.
  • Dimensionality Reduction → UMAP is the gold standard. Reduce to 5D for clustering, 2D for visualisation.
  • Clustering → HDBSCAN for unknown K, K-Means for known K, Agglomerative for topic hierarchies
  • BERTopic → The one library that wraps the entire modern pipeline with a clean API and beautiful visualisations
  • MLOps Integration → Save models, use transform() for inference, monitor noise rates for drift
  • Evaluation → Silhouette Score for cluster quality, Topic Diversity for topic distinctiveness

Topic Clustering is one of those techniques that, once you understand it, you start seeing opportunities to apply it everywhere — in your data pipelines, your model monitoring, your LLM datasets, and your production systems 🧠✨

Comments