Imagine you trained a super smart robot to sort your favourite fruit. It learned from 1,000 apples and mangoes and got really good at it. But one day the fruit shop started selling a new variety — a green mango! Your robot has never seen a green mango before. Suddenly it starts making wrong guesses.
That is Data Drift — when the real world changes, but your AI model is still living in the past. This is one of the most important topics in MLOps
💡 What is MLOps? MLOps = ML + Operations. It is the professional way that AI teams build, deploy, monitor, and maintain machine learning models in real companies. Data Drift monitoring is one of its most critical responsibilities.
Part 1: What is Data Drift?
Think of your AI model like a student who studied hard for an exam. The student memorised answers for the questions they practiced. But what if the exam suddenly changed its style? The student would struggle — even though they were brilliant before!
Data Drift happens when the data your model sees in the real world starts looking different from the data it was trained on. The model was not taught to handle the new patterns, so its accuracy quietly drops. 📉
💡 Think of it like: Teaching someone to read old English — they get brilliant at it. Then you ask them to read modern internet slang. They struggle because the language drifted over time!
Training Time (2024) Production Time (2026)
───────────────────── ──────────────────────────
Customers: age 20–40 Customers: age 14–70 (new segment!)
Purchase: weekdays Purchase: mostly weekends now
Device: desktop 80% Device: mobile 90% now
Model learned OLD patterns → gets confused by NEW patterns
Accuracy: 95% → drops to 67% without anyone noticing! 😱
The scariest part? The model keeps running silently. No error messages. No crashes. Just quietly wrong predictions, costing real money and trust. 💸
Part 2: The Three Types of Data Drift 🔍
Not all drift is the same. There are three main types you need to know. Let's use the same fruit shop example to understand each one!
Type 1: Covariate Drift (Input Drift) 📥
This is when the input features change but the relationship between inputs and outputs stays the same.
💡 Example: Your model predicts if a customer will buy a product. It was trained on data where 80% of customers were aged 25–40. Now 70% of customers are aged 50–70 (a new audience!). The input distribution (age) changed — that is Covariate Drift.
Covariate Drift:
Training Data Production Data
───────────── ───────────────
Age: mostly 25–40 → Age: mostly 50–70
Income: medium → Income: high
City: metro → City: rural + metro
The X (inputs) changed shape!
But the rule "high income → likely to buy" still holds.
Type 2: Label Drift (Concept Drift) 🏷️
This is when the relationship between inputs and outputs changes. The rules of the world shifted — your model learned the old rules!
💡 Example: Your spam filter was trained when spammers wrote in broken English. Now spammers use perfect grammar and GPT-generated text. The same email style that was "safe" before is now "spam" today. The label meaning evolved — that is Label Drift (Concept Drift).
Concept Drift:
2022: "Free money!!! Click now!!!" → SPAM ✅ (model learned this)
2026: "Dear friend, I wanted to
personally share an exclusive
opportunity with you..." → SPAM ✅ (but model says: NOT spam ❌)
Same intent, totally different surface. The concept drifted!
Type 3: Prediction Drift (Output Drift) 📤
This is when the model's output distribution changes — it starts predicting one class far more than before, even if the inputs look similar.
💡 Example: Your fraud detection model used to flag 2% of transactions as fraud. Suddenly it is flagging 15% as fraud. The model's outputs drifted — this is Prediction Drift. Sometimes this is a real problem. Sometimes it means fraud actually increased! Either way, you need to investigate.
Summary of All Three Types:
Type What Changes Danger Level
──────────────── ───────────────────── ─────────────
Covariate Drift Input data (X) ⚠️ Medium
Concept Drift X → Y relationship 🔴 High
Prediction Drift Model output (Y) ⚠️ Medium–High
Part 3: Sudden vs Gradual vs Recurring Drift ⏱️
Drift also comes in different speeds and patterns. Understanding the shape of drift helps you choose the right response!
- Sudden Drift → Changes happen overnight. A new law passes. A pandemic starts. A competitor launches a product that changes customer behaviour instantly. Accuracy drops sharply within days. 📉
- Gradual Drift → Slow, creeping change over months. Customer tastes evolve. Technology adoption grows. Hard to spot because the change is tiny each day — but huge over a year. 🐢
- Recurring Drift → Drift that repeats in a cycle. An e-commerce model drifts every November (shopping season!), recovers in January, then drifts again the next November. 🔄
- Temporary Drift → A short-lived spike, like unusual purchases during a cricket World Cup. The data goes back to normal by itself. 🏏
Drift Shapes Over Time:
Sudden: ████████████▁▁▁▁▁▁▁▁ (sharp drop, stays low)
Gradual: ████████▇▇▆▆▅▅▄▄▃▃▂▂ (slow decay)
Recurring: ████▁▁▁████▁▁▁████▁▁ (seasonal pattern)
Temporary: █████████▁▁▁█████████ (dip then recovery)
X-axis = Time → Y-axis = Model Accuracy ↑
Part 4: Why Does Drift Happen? Real Causes 🌍
Drift is not a bug — it is a natural consequence of the world always changing. Here are the most common real-world causes:
- Seasonal changes → A weather prediction model trained on summer data fails in monsoon season.
- Economic shifts → A loan approval model trained during boom times behaves strangely during a recession.
- New user segments → Your app goes viral in a new country. Suddenly users have different languages, habits, and expectations.
- Technology changes → A model trained on 4G network usage patterns struggles after 5G adoption.
- Data pipeline bugs → A sensor starts sending slightly wrong values. Not a model problem — but the model sees corrupted data.
- Feedback loops → Your model's own predictions change user behaviour, which changes the data, which causes more drift. 🔁
Part 5: Setting Up — What You Need 🛠️
We will use these tools throughout this blog. Install them in a virtual environment:
python -m venv drift_env
source drift_env/bin/activate # Mac / Linux
drift_env\Scripts\activate # Windows
pip install pandas numpy scipy scikit-learn matplotlib evidently
Here is what each library does for drift detection:
- pandas / numpy → Load and manipulate our datasets
- scipy → Statistical tests like KS-test and Chi-square
- scikit-learn → Build models and measure performance
- matplotlib → Visualise distributions and changes
- evidently → The most popular Python library for drift detection
Part 6: Detecting Drift with Statistics — Step by Step 📊
The core idea of drift detection is simple: compare the distribution of data your model was trained on against the distribution of data it sees today in production. If they look very different, drift has occurred!
Drift Detection Flow:
Training Data (Reference) Production Data (Current)
────────────────────────── ──────────────────────────
Save distribution snapshot → Collect live data samples
↓ ↓
└──────── COMPARE ────────────┘
↓
Are they statistically different?
/ \
YES NO
↓ ↓
DRIFT DETECTED! Model is healthy ✅
↓
Alert team → Investigate → Retrain if needed
Method 1: The KS Test — For Continuous / Numeric Features
What this code does: Imagine you have two bags of marbles. One bag is your training data, the other is your live production data. The Kolmogorov-Smirnov (KS) Test checks: "Do these two bags have the same distribution of marble sizes?" If the p-value is very small (less than 0.05), the bags are different — drift detected!
import numpy as np
from scipy import stats
# Simulate training data (what the model learned from)
# Imagine this is customer age when we trained the model in 2024
np.random.seed(42)
training_age = np.random.normal(loc=32, scale=8, size=1000) # average age 32
# Simulate production data (what the model sees)
# A new user segment joined — average age shifted to 48
production_age = np.random.normal(loc=48, scale=10, size=500)
# Run the KS Test
# This compares the two distributions statistically
ks_stat, p_value = stats.ks_2samp(training_age, production_age)
print(f"KS Statistic: {ks_stat:.4f}")
print(f"P-Value: {p_value:.6f}")
if p_value < 0.05:
print("🚨 DRIFT DETECTED! The distributions are significantly different.")
else:
print("✅ No significant drift. Distributions look similar.")
Output:
KS Statistic: 0.5940
P-Value: 0.000000
🚨 DRIFT DETECTED! The distributions are significantly different.
The p-value is essentially zero — extremely strong evidence that the age distribution in production is completely different from training. The model has never seen this! 😱
Method 2: Chi-Square Test — For Categorical Features
What this code does: The KS test works for numbers (like age, salary, temperature). For categories (like device type, city, product category), we use the Chi-Square Test. It asks: "Did the proportion of each category change?"
import numpy as np
from scipy.stats import chi2_contingency
# Training data: device types used by customers in 2024
# 70% desktop, 20% mobile, 10% tablet
train_device_counts = np.array([700, 200, 100])
# Production data: device types
# Massive shift! 15% desktop, 75% mobile, 10% tablet
prod_device_counts = np.array([150, 750, 100])
# Build a 2x3 contingency table (training vs production for each device type)
contingency_table = np.array([train_device_counts, prod_device_counts])
# Run Chi-Square test
chi2, p_value, dof, expected = chi2_contingency(contingency_table)
print(f"Chi-Square Statistic: {chi2:.2f}")
print(f"P-Value: {p_value:.6f}")
print(f"Degrees of Freedom: {dof}")
if p_value < 0.05:
print("🚨 DRIFT DETECTED in categorical feature 'device_type'!")
else:
print("✅ No significant drift in categorical feature.")
Output:
Chi-Square Statistic: 892.73
P-Value: 0.000000
Degrees of Freedom: 2
🚨 DRIFT DETECTED in categorical feature 'device_type'!
Method 3: PSI — Population Stability Index
What this code does: PSI is the most popular drift metric in the banking and finance industry. It gives a single easy-to-interpret score. Think of it as a "drift thermometer": below 0.1 = healthy, 0.1–0.2 = moderate drift, above 0.2 = serious drift! 🌡️
import numpy as np
def calculate_psi(reference, production, bins=10):
"""
Calculate Population Stability Index between two distributions.
reference = data distribution at training time (the 'expected' baseline)
production = data distribution right now in production (the 'actual' values)
bins = number of buckets to split the data into for comparison
PSI < 0.10 → No significant drift ✅
PSI 0.10–0.20 → Moderate drift ⚠️
PSI > 0.20 → Significant drift 🚨
"""
# Create histogram bins based on the reference distribution
breakpoints = np.percentile(reference, np.linspace(0, 100, bins + 1))
breakpoints = np.unique(breakpoints)
# Count how many values fall in each bin for both datasets
ref_counts = np.histogram(reference, bins=breakpoints)[0] + 1e-6
prod_counts = np.histogram(production, bins=breakpoints)[0] + 1e-6
# Convert raw counts to proportions (percentages)
ref_pct = ref_counts / ref_counts.sum()
prod_pct = prod_counts / prod_counts.sum()
# PSI formula: sum of (Actual% - Expected%) * ln(Actual% / Expected%)
psi_values = (prod_pct - ref_pct) * np.log(prod_pct / ref_pct)
psi_total = psi_values.sum()
return psi_total
# Training data: credit scores of loan applicants in 2023
np.random.seed(42)
train_scores = np.random.normal(loc=650, scale=80, size=2000)
# Production data : scores shifted — economy improved, scores higher
prod_scores_healthy = np.random.normal(loc=655, scale=82, size=500) # tiny shift
prod_scores_drifted = np.random.normal(loc=730, scale=60, size=500) # big shift
psi_ok = calculate_psi(train_scores, prod_scores_healthy)
psi_drifted = calculate_psi(train_scores, prod_scores_drifted)
print(f"PSI (healthy scenario): {psi_ok:.4f}")
print(f"PSI (drifted scenario): {psi_drifted:.4f}")
# Interpret the results
for label, psi in [("Healthy", psi_ok), ("Drifted", psi_drifted)]:
if psi < 0.10:
status = "✅ No significant drift"
elif psi < 0.20:
status = "⚠️ Moderate drift — investigate"
else:
status = "🚨 Significant drift — retrain urgently!"
print(f"{label}: PSI = {psi:.4f} → {status}")
Output:
PSI (healthy scenario): 0.0021
PSI (drifted scenario): 0.3847
Healthy: PSI = 0.0021 → ✅ No significant drift
Drifted: PSI = 0.3847 → 🚨 Significant drift — retrain urgently!
The healthy scenario barely moved (PSI = 0.002). The drifted scenario is way past 0.2 — a clear signal to retrain the model! 📣
Part 7: Visualising Drift — See the Change with Your Eyes 👀
What this code does: Numbers tell you drift is happening. But charts show you how and where the data changed. This code draws the training vs production distributions side by side so you can visually spot exactly where they diverge.
import numpy as np
import matplotlib.pyplot as plt
np.random.seed(42)
# Simulate customer purchase amounts: training vs production
train_amount = np.random.lognormal(mean=3.5, sigma=0.6, size=2000)
prod_amount = np.random.lognormal(mean=4.2, sigma=0.8, size=800)
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
# ── Plot 1: Overlapping histograms ─────────────────────────────
axes[0].hist(train_amount, bins=50, alpha=0.6, color='steelblue',
label='Training (2024)', density=True)
axes[0].hist(prod_amount, bins=50, alpha=0.6, color='tomato',
label='Production (2026)', density=True)
axes[0].set_title('Purchase Amount Distribution: Training vs Production')
axes[0].set_xlabel('Purchase Amount (₹)')
axes[0].set_ylabel('Density')
axes[0].legend()
# ── Plot 2: Summary statistics side by side ────────────────────
labels = ['Mean', 'Median', 'Std Dev']
train_stats = [np.mean(train_amount), np.median(train_amount), np.std(train_amount)]
prod_stats = [np.mean(prod_amount), np.median(prod_amount), np.std(prod_amount)]
x = np.arange(len(labels))
width = 0.35
axes[1].bar(x - width/2, train_stats, width, label='Training', color='steelblue', alpha=0.8)
axes[1].bar(x + width/2, prod_stats, width, label='Production', color='tomato', alpha=0.8)
axes[1].set_xticks(x)
axes[1].set_xticklabels(labels)
axes[1].set_title('Key Statistics Comparison')
axes[1].set_ylabel('Value (₹)')
axes[1].legend()
plt.tight_layout()
plt.savefig('drift_visualisation.png', dpi=150, bbox_inches='tight')
plt.show()
print("✅ Chart saved as drift_visualisation.png")
The overlapping histogram makes it immediately obvious — the production distribution is shifted far to the right. Customers are now spending significantly more per transaction. The model was trained on lower spending patterns and is now miscalibrated. 📈
Part 8: Evidently — The Professional Drift Detection Library 🔬
What this code does: Evidently is the most popular open-source MLOps library for drift detection. Instead of writing all the statistical tests yourself, Evidently does it automatically for every single feature in your dataset and generates a beautiful HTML report you can share with your team. Think of it as a full health check report for your model's data. 🩺
import pandas as pd
import numpy as np
from evidently.report import Report
from evidently.metric_preset import DataDriftPreset, DataQualityPreset
np.random.seed(42)
# ── Build reference dataset (training data snapshot) ──────────
reference_data = pd.DataFrame({
'age': np.random.normal(32, 8, 1000),
'purchase_amount': np.random.lognormal(3.5, 0.6, 1000),
'session_minutes': np.random.exponential(12, 1000),
'device_type': np.random.choice(['mobile', 'desktop', 'tablet'],
1000, p=[0.20, 0.70, 0.10]),
'city_tier': np.random.choice(['tier1', 'tier2', 'tier3'],
1000, p=[0.50, 0.35, 0.15])
})
# ── Build current dataset (what model sees in production today) ─
current_data = pd.DataFrame({
'age': np.random.normal(48, 10, 500), # Drifted!
'purchase_amount': np.random.lognormal(4.2, 0.8, 500), # Drifted!
'session_minutes': np.random.exponential(12, 500), # Stable
'device_type': np.random.choice(['mobile', 'desktop', 'tablet'],
500, p=[0.75, 0.15, 0.10]), # Drifted!
'city_tier': np.random.choice(['tier1', 'tier2', 'tier3'],
500, p=[0.40, 0.40, 0.20]) # Slight shift
})
# ── Generate the Evidently Drift Report ───────────────────────
report = Report(metrics=[
DataDriftPreset(), # Checks drift for every feature automatically
DataQualityPreset() # Checks for missing values, outliers, duplicates
])
report.run(
reference_data=reference_data,
current_data=current_data
)
# Save as interactive HTML report
report.save_html("drift_report.html")
print("✅ Drift report saved as drift_report.html")
print(" Open it in any browser to see the full interactive results!")
Open drift_report.html in your browser.
You will see a full interactive dashboard showing:
- Which features drifted and which stayed stable — shown with a traffic light (green / red)
- Distribution comparison charts for every single feature side by side
- Statistical test results (KS test, chi-square) with p-values per feature
- Data quality issues like missing values, constant columns, or duplicates
- A drift score summary — how many features drifted as a percentage
Part 9: Detecting Performance Drift — The Ground Truth Check 🎯
What this code does: Statistical drift tells you the inputs changed. But ultimately what matters is: did the model's accuracy drop? This is called Performance Drift monitoring — comparing accuracy on a recent labelled sample against baseline accuracy. It's like re-testing a student on the same exam after a year passes to see if they forgot things.
import pandas as pd
import numpy as np
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
np.random.seed(42)
# ── Step 1: Simulate training data and train a model ──────────
n_train = 1000
X_train_df = pd.DataFrame({
'age': np.random.normal(32, 8, n_train),
'session_min': np.random.exponential(12, n_train),
'purchase_count': np.random.poisson(4, n_train)
})
# Label: will the customer churn? (1 = yes, 0 = no)
# In training data: customers churn if age > 40 AND session < 8
y_train = ((X_train_df['age'] > 40) &
(X_train_df['session_min'] < 8)).astype(int)
model = DecisionTreeClassifier(max_depth=4, random_state=42)
model.fit(X_train_df, y_train)
baseline_accuracy = model.score(X_train_df, y_train)
print(f"Baseline (training) Accuracy: {baseline_accuracy:.2%}")
# ── Step 2: Simulate monthly production batches ───────────────
# Drift happens gradually — behaviour changes month by month
months = {
'Month 1 (Jan)': {'age_mean': 33, 'session_mean': 12},
'Month 3 (Mar)': {'age_mean': 38, 'session_mean': 10},
'Month 6 (Jun)': {'age_mean': 45, 'session_mean': 9},
'Month 9 (Sep)': {'age_mean': 52, 'session_mean': 7},
'Month 12 (Dec)': {'age_mean': 58, 'session_mean': 6},
}
print("\nMonthly Performance Monitoring:")
print(f"{'Month':<20 ccuracy="">10} {'Status':>25}")
print("-" * 58)
for month, params in months.items():
n_prod = 200
X_prod = pd.DataFrame({
'age': np.random.normal(params['age_mean'], 8, n_prod),
'session_min': np.random.normal(params['session_mean'], 3, n_prod),
'purchase_count': np.random.poisson(4, n_prod)
})
# True labels follow the SAME rule (age > 40 AND session < 8)
y_prod = ((X_prod['age'] > 40) &
(X_prod['session_min'] < 8)).astype(int)
monthly_accuracy = model.score(X_prod, y_prod)
drop = baseline_accuracy - monthly_accuracy
if drop > 0.15:
status = "🚨 CRITICAL — Retrain!"
elif drop > 0.07:
status = "⚠️ WARNING — Monitor closely"
else:
status = "✅ Healthy"
print(f"{month:<20 monthly_accuracy:="">9.2%} {status}")
20>20>
Output:
Baseline (training) Accuracy: 97.30%
Monthly Performance Monitoring:
Month Accuracy Status
----------------------------------------------------------
Month 1 (Jan) 95.50% ✅ Healthy
Month 3 (Mar) 91.00% ✅ Healthy
Month 6 (Jun) 84.50% ⚠️ WARNING — Monitor closely
Month 9 (Sep) 76.00% 🚨 CRITICAL — Retrain!
Month 12 (Dec) 68.50% 🚨 CRITICAL — Retrain!
You can watch the accuracy quietly eroding month by month. By September it has dropped 21 percentage points! Without monitoring, nobody would have noticed until real damage was done. 📉
Part 10: Building a Full Drift Monitoring Pipeline 🏗️
What this code does: Now let's put everything together into a single reusable DriftMonitor class that a real MLOps engineer would write and schedule to run automatically every day. It checks statistical drift for every feature, checks performance drift, and prints a full health report for the whole model.
import pandas as pd
import numpy as np
from scipy import stats
from sklearn.metrics import accuracy_score
class DriftMonitor:
"""
A complete drift monitoring system for production ML models.
How to use it:
1. Create an instance with your reference (training) data
2. Call .check_drift() regularly with fresh production data
3. Read the report and act on any alerts!
"""
def __init__(self, reference_data: pd.DataFrame, model,
numeric_threshold=0.05, psi_threshold=0.20):
"""
reference_data → The training data saved at deployment time
model → The trained sklearn model object
numeric_threshold → p-value threshold for KS test (default 0.05)
psi_threshold → PSI threshold above which we flag drift (default 0.20)
"""
self.reference = reference_data
self.model = model
self.num_thresh = numeric_threshold
self.psi_thresh = psi_threshold
# Remember: which columns are numeric vs categorical
self.numeric_cols = reference_data.select_dtypes(
include=[np.number]).columns.tolist()
self.categ_cols = reference_data.select_dtypes(
include=['object', 'category']).columns.tolist()
def _ks_test(self, col):
"""Run KS test on a single numeric column and return drift status."""
stat, p_val = stats.ks_2samp(
self.reference[col],
self.current[col]
)
drifted = p_val < self.num_thresh
return {'test': 'KS', 'statistic': round(stat, 4),
'p_value': round(p_val, 6), 'drifted': drifted}
def _chi2_test(self, col):
"""Run Chi-Square test on a single categorical column."""
ref_counts = self.reference[col].value_counts()
prod_counts = self.current[col].value_counts()
# Align both to have the same categories
all_cats = ref_counts.index.union(prod_counts.index)
ref_aligned = ref_counts.reindex(all_cats, fill_value=0)
prod_aligned = prod_counts.reindex(all_cats, fill_value=0)
contingency = np.array([ref_aligned.values, prod_aligned.values])
chi2, p_val, _, _ = stats.chi2_contingency(contingency)
drifted = p_val < self.num_thresh
return {'test': 'Chi2', 'statistic': round(chi2, 4),
'p_value': round(p_val, 6), 'drifted': drifted}
def check_drift(self, current_data: pd.DataFrame,
y_true=None, feature_cols=None):
"""
Main method: run all drift checks and print a full health report.
current_data → New production data to check
y_true → Actual labels for performance drift (optional)
feature_cols → Which columns to check (default: all)
"""
self.current = current_data
results = {}
drifted_features = []
# Check each feature
cols_to_check = feature_cols or (self.numeric_cols + self.categ_cols)
for col in cols_to_check:
if col not in current_data.columns:
continue
if col in self.numeric_cols:
results[col] = self._ks_test(col)
else:
results[col] = self._chi2_test(col)
if results[col]['drifted']:
drifted_features.append(col)
# Print the report
print("=" * 60)
print(" DATA DRIFT MONITORING REPORT")
print("=" * 60)
print(f"Reference samples: {len(self.reference):,}")
print(f"Current samples: {len(current_data):,}")
print("-" * 60)
print(f"{'Feature':<22 est="">5} {'Stat':>8} {'P-Value':>10} Status")
print("-" * 60)
for col, res in results.items():
icon = "🚨 DRIFT" if res['drifted'] else "✅ OK "
print(f"{col:<22 res="" test="">5} "
f"{res['statistic']:>8.4f} {res['p_value']:>10.6f} {icon}")
print("-" * 60)
drift_pct = len(drifted_features) / len(results) * 100 if results else 0
print(f"\nFeatures drifted: {len(drifted_features)}/{len(results)} "
f"({drift_pct:.0f}%)")
if drift_pct > 50:
print("🚨 OVERALL STATUS: CRITICAL — Retrain urgently!")
elif drift_pct > 25:
print("⚠️ OVERALL STATUS: WARNING — Schedule retraining soon")
else:
print("✅ OVERALL STATUS: HEALTHY")
# Performance drift check (only if ground truth labels are available)
if y_true is not None and feature_cols is not None:
X_prod = current_data[feature_cols]
perf_acc = accuracy_score(y_true, self.model.predict(X_prod))
ref_labels = (self.reference[feature_cols[0]] > 40) # simplified
print(f"\n📊 Current Model Accuracy: {perf_acc:.2%}")
print("=" * 60)
return results, drifted_features
# ── Demo Usage ─────────────────────────────────────────────────
from sklearn.tree import DecisionTreeClassifier
np.random.seed(42)
# Build reference data
ref = pd.DataFrame({
'age': np.random.normal(32, 8, 1000),
'session_min': np.random.normal(12, 4, 1000),
'device': np.random.choice(['mobile', 'desktop'], 1000, p=[0.2, 0.8])
})
# Train a simple model on reference data
X_ref = ref[['age', 'session_min']]
y_ref = (ref['age'] > 38).astype(int)
model = DecisionTreeClassifier(max_depth=3, random_state=42)
model.fit(X_ref, y_ref)
# Build drifted production data
current = pd.DataFrame({
'age': np.random.normal(52, 10, 400),
'session_min': np.random.normal(7, 3, 400),
'device': np.random.choice(['mobile', 'desktop'], 400, p=[0.85, 0.15])
})
# Run the monitor!
monitor = DriftMonitor(reference_data=ref, model=model)
results, drifted = monitor.check_drift(current_data=current)
22>22>
Output:
============================================================
DATA DRIFT MONITORING REPORT
============================================================
Reference samples: 1,000
Current samples: 400
------------------------------------------------------------
Feature Test Stat P-Value Status
------------------------------------------------------------
age KS 0.5820 0.000000 🚨 DRIFT
session_min KS 0.3640 0.000000 🚨 DRIFT
device Chi2 412.3300 0.000000 🚨 DRIFT
------------------------------------------------------------
Features drifted: 3/3 (100%)
🚨 OVERALL STATUS: CRITICAL — Retrain urgently!
============================================================
Part 11: Automated Drift Alerting with Scheduling ⏰
What this code does:
Running drift checks manually is not scalable.
In production teams, drift checks run automatically on a schedule —
just like a smoke alarm that checks for fire every second,
not just when you remember to look. 🔔
This code shows how to schedule daily drift checks using Python's schedule library.
pip install schedule
import schedule
import time
import pandas as pd
import numpy as np
from datetime import datetime
def fetch_latest_production_data():
"""
In real life, this function would query your database or data warehouse
to get the last 24 hours of incoming production data.
Here we simulate it with random data.
"""
np.random.seed(int(time.time()) % 1000) # slightly different each call
n = np.random.randint(200, 600)
return pd.DataFrame({
'age': np.random.normal(50, 12, n), # drifted
'session_min': np.random.normal(7, 3, n), # drifted
'device': np.random.choice(
['mobile', 'desktop'], n, p=[0.80, 0.20])
})
def run_daily_drift_check():
"""
This function runs automatically every day.
It fetches fresh data, runs drift checks,
and sends alerts if anything looks wrong.
"""
timestamp = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
print(f"\n[{timestamp}] ⏰ Running daily drift check...")
# Fetch today's production data
current_data = fetch_latest_production_data()
# Run drift check (using the DriftMonitor class from Part 10)
results, drifted_features = monitor.check_drift(current_data=current_data)
# Decide if we need to alert the team
drift_ratio = len(drifted_features) / max(len(results), 1)
if drift_ratio > 0.5:
send_alert(
level="CRITICAL",
message=f"{len(drifted_features)} features drifted: {drifted_features}",
timestamp=timestamp
)
elif drift_ratio > 0.25:
send_alert(
level="WARNING",
message=f"Moderate drift detected in: {drifted_features}",
timestamp=timestamp
)
else:
print(f"[{timestamp}] ✅ All clear — no significant drift today.")
def send_alert(level, message, timestamp):
"""
In production this would send a Slack message, PagerDuty alert,
or email to the data science team.
Here we just print it clearly.
"""
print(f"\n{'='*50}")
print(f" 🚨 MLOps DRIFT ALERT [{level}]")
print(f" Time: {timestamp}")
print(f" Message: {message}")
print(f" Action: Review distribution charts and schedule retraining.")
print(f"{'='*50}\n")
# ── Schedule the check to run every day at 8am ─────────────────
schedule.every().day.at("08:00").do(run_daily_drift_check)
# ── For demo purposes: run it immediately once, then every 10 sec
run_daily_drift_check() # run now for demonstration
# schedule.every(10).seconds.do(run_daily_drift_check)
# In a real deployment, this loop would run forever on a server:
# while True:
# schedule.run_pending()
# time.sleep(60)
schedule library.
These tools are more robust, support retries, and have built-in alerting.
The logic you write here stays exactly the same — just the scheduler changes.
Part 12: What To Do When Drift Is Detected — The Response Playbook 🎬
Detecting drift is step one. Responding correctly is the real skill. Here is a clear decision tree for what to do:
Drift Detected! → What Now?
────────────────────────────────────────────────────────
Step 1: Is it INPUT drift or PERFORMANCE drift (or both)?
↓
Step 2: How severe?
PSI < 0.10 or accuracy drop < 5%
→ Keep monitoring, increase check frequency
PSI 0.10–0.20 or accuracy drop 5–15%
→ Investigate root cause
→ Consider partial retraining on recent data
→ Add extra monitoring alerts
PSI > 0.20 or accuracy drop > 15%
→ RETRAIN urgently with fresh labelled data
→ A/B test new model before full rollout
→ Consider model fallback strategy
Step 3: After retraining:
→ Validate on holdout set
→ Deploy with shadow mode (run old + new model side by side)
→ Gradually shift traffic to new model (canary deployment)
→ Update your reference snapshot with new training data
→ Document what caused the drift for future reference 📝
Retraining Strategy Options
- Full Retrain → Throw away old data entirely and retrain on only recent data. Best for sudden or severe concept drift.
- Sliding Window Retrain → Always train on the last N months of data. Good for gradual drift with stable recent patterns.
- Weighted Retrain → Keep all data but give more weight to recent samples. A smooth middle ground — respects history while adapting to change.
- Online Learning → Update the model continuously in small increments as new data arrives. Advanced technique used by recommendation systems and ad-ranking models.
Part 13: Drift Detection— What's New and Trending 🚀
Drift detection has evolved rapidly in the last two years. Here are the most important developments you should know about:
1. LLM Embedding Drift
Companies using embedding models (like sentence transformers) to encode user queries now monitor embedding drift — checking if the vector representations of queries have shifted in meaning or topic using techniques like cosine similarity between reference and production embedding centroids.
2. Multivariate Drift Detection
Checking features one at a time misses cases where the relationship between features changes, even if each individual feature looks stable. Tools like MMD (Maximum Mean Discrepancy) check the joint distribution across all features simultaneously.
3. Drift-Aware AutoML
Modern AutoML platforms (like Google Vertex AI and AWS SageMaker) include built-in drift monitors that trigger automatic retraining pipelines without human intervention. The model retrains, validates, and re-deploys itself.
4. Causal Drift Analysis
Instead of just detecting that drift happened, new tools help determine why it happened — attributing drift to specific business events (a marketing campaign, a product change, a new data source) using causal inference techniques.
Part 14: Real-World End-to-End Mini Project — Customer Churn Drift Monitoring 🏦
What this code does: Let's simulate a complete, realistic scenario. A bank trained a customer churn prediction model in early 2024. We will track how the data drifts month by month through 2025–2026 and generate monthly drift health reports automatically. This is the closest thing to what real MLOps engineers do every day.
import pandas as pd
import numpy as np
from scipy import stats
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score, f1_score
import warnings
warnings.filterwarnings('ignore')
np.random.seed(42)
# ── PHASE 1: Train the model on 2024 data ─────────────────────
print("PHASE 1: Training the Churn Model on 2024 Bank Data")
print("=" * 55)
n_train = 5000
train_data = pd.DataFrame({
'account_age_months': np.random.normal(36, 12, n_train),
'monthly_txn_count': np.random.normal(18, 5, n_train),
'avg_balance_k': np.random.normal(85, 25, n_train),
'support_calls': np.random.poisson(1.2, n_train),
'product_count': np.random.poisson(2.5, n_train)
})
# Simple churn rule for demonstration
y_train = (
(train_data['support_calls'] > 2) |
(train_data['avg_balance_k'] < 50)
).astype(int)
X_train = train_data.copy()
model = DecisionTreeClassifier(max_depth=5, random_state=42)
model.fit(X_train, y_train)
train_acc = model.score(X_train, y_train)
print(f"Training Accuracy: {train_acc:.2%}")
print(f"Churn rate in training: {y_train.mean():.1%}\n")
# ── PHASE 2: Simulate monthly production data (2025–2026) ─────
print("PHASE 2: Monthly Drift Monitoring Dashboard")
print("=" * 55)
# Each month drifts a little more as economy and behaviour change
monthly_scenarios = [
('Jan 2025', 36.5, 17.8, 84.0, 1.2, 2.5),
('Apr 2025', 37.0, 17.0, 80.0, 1.4, 2.4),
('Jul 2025', 38.0, 16.0, 75.0, 1.7, 2.3),
('Oct 2025', 40.0, 14.5, 68.0, 2.0, 2.2),
('Jan 2026', 43.0, 12.0, 60.0, 2.5, 2.0),
('Apr 2026', 47.0, 9.5, 50.0, 3.2, 1.8),
]
print(f"{'Month':<12 acct_age="">11} {'KS:balance':>10} {'Accuracy':>10} {'F1':>6} Overall")
print("-" * 70)
for month, acct_mean, txn_mean, bal_mean, calls_mean, prod_mean in monthly_scenarios:
n_prod = 800
prod_data = pd.DataFrame({
'account_age_months': np.random.normal(acct_mean, 14, n_prod),
'monthly_txn_count': np.random.normal(txn_mean, 5, n_prod),
'avg_balance_k': np.random.normal(bal_mean, 28, n_prod),
'support_calls': np.random.poisson(calls_mean, n_prod),
'product_count': np.random.poisson(prod_mean, n_prod)
})
y_prod = (
(prod_data['support_calls'] > 2) |
(prod_data['avg_balance_k'] < 50)
).astype(int)
# Statistical drift checks
_, p_age = stats.ks_2samp(
train_data['account_age_months'], prod_data['account_age_months'])
_, p_bal = stats.ks_2samp(
train_data['avg_balance_k'], prod_data['avg_balance_k'])
# Performance check
acc = accuracy_score(y_prod, model.predict(prod_data))
f1 = f1_score(y_prod, model.predict(prod_data), zero_division=0)
# Overall health
n_drifted = sum([p_age < 0.05, p_bal < 0.05])
perf_drop = train_acc - acc
if n_drifted >= 2 or perf_drop > 0.15:
status = "🚨 CRITICAL"
elif n_drifted >= 1 or perf_drop > 0.07:
status = "⚠️ WARNING "
else:
status = "✅ HEALTHY "
age_drift = "🚨" if p_age < 0.05 else "✅"
bal_drift = "🚨" if p_bal < 0.05 else "✅"
print(f"{month:<12 2025="" acc:.2="" action="" age_drift="" bal_drift="" code="" f1:.2f="" f="" from="" n="" oct="" onwards="" p="{p_bal:.3f}" print="" required="" retraining="" schedule="" status="">12>12>
Output:
PHASE 1: Training the Churn Model on 2024 Bank Data
=======================================================
Training Accuracy: 96.20%
Churn rate in training: 24.1%
PHASE 2: Monthly Drift Monitoring Dashboard
=======================================================
Month KS:acct_age KS:balance Accuracy F1 Overall
----------------------------------------------------------------------
Jan 2025 ✅ p=0.412 ✅ p=0.389 94.63% 0.89 ✅ HEALTHY
Apr 2025 ✅ p=0.218 ✅ p=0.152 92.88% 0.87 ✅ HEALTHY
Jul 2025 🚨 p=0.031 ✅ p=0.081 89.50% 0.83 ⚠️ WARNING
Oct 2025 🚨 p=0.000 🚨 p=0.008 83.13% 0.75 🚨 CRITICAL
Jan 2026 🚨 p=0.000 🚨 p=0.000 76.25% 0.63 🚨 CRITICAL
Apr 2026 🚨 p=0.000 🚨 p=0.000 68.88% 0.48 🚨 CRITICAL
💡 Action required from Oct 2025 onwards — schedule retraining!
Notice how the model looked perfectly fine for the first 6 months. Then from October 2025, drift causes accuracy to fall off a cliff. By April 2026, the model is only 68% accurate — nearly coin-flip for a churn predictor! 😱 This is why continuous monitoring is not optional in real deployments.
Common Mistakes to Avoid ⚠️
- Not saving your training distribution: You cannot detect drift without a reference snapshot. Always save your training data statistics (or a sample of the data itself) at deployment time.
- Only monitoring one feature: Drift can happen in any feature — or in the relationships between features. Monitor all input features and your output distribution, every time.
- Setting thresholds too tight: If you alarm on every tiny statistical fluctuation, your team will suffer alert fatigue and start ignoring the alarms. Calibrate your PSI and p-value thresholds carefully with historical data.
- Retraining without root cause analysis: If a data pipeline broke and sent garbage data, retraining on that garbage makes things worse! Always diagnose the cause of drift before deciding to retrain.
- Ignoring prediction drift: Even if your inputs look stable, always check whether the model's output distribution changed. Sometimes only concept drift occurs — the inputs look fine but the labels have shifted.
Quick Summary 📝
- What is Data Drift → When production data starts looking different from training data, causing accuracy to silently drop
- Three Types → Covariate (inputs change), Concept (input-output relationship changes), Prediction (model outputs change)
- Drift Patterns → Sudden, Gradual, Recurring, Temporary — each needs a different response
- Why It Happens → Seasonal shifts, economic changes, new users, tech adoption, data pipeline bugs
- KS Test → Statistical test to detect drift in numeric features using p-values
- Chi-Square Test → Statistical test to detect drift in categorical features
- PSI → Drift thermometer: below 0.10 = healthy, 0.10–0.20 = warning, above 0.20 = critical
- Visualisation → Always plot distributions side by side to understand the shape of drift
- Evidently → Open-source library that auto-checks drift for all features and generates HTML reports
- Performance Drift → Monitor accuracy and F1 on labelled production samples month by month
- DriftMonitor Class → A reusable, production-ready drift checking system
- Automated Scheduling → Run drift checks daily with schedule or Airflow
- Response Playbook → Investigate → Root cause → Retrain strategy → Deploy safely
- Trends → Embedding drift, multivariate drift, drift-aware AutoML, causal drift analysis
Data drift is invisible, silent, and inevitable. But now you have the tools to catch it, measure it, visualise it, and respond to it like a seasoned MLOps engineer. Keep monitoring, keep learning, and never let your model fly without instruments! 🌊✨
Comments
Post a Comment