Skip to main content

Concept Drift in MLOps: What It Is, Why It Happens & How to Detect It

Calculating read time…

Imagine you taught your little brother to sort fruits. He learned perfectly — apples go left, mangoes go right. But one day the fruit shop started selling a new variety — a yellow apple that looks exactly like a mango. Your brother keeps putting it in the wrong pile. He is not wrong by the old rules — the rules of the world just changed!

That is Concept Drift. It is when the relationship between your model's inputs and the correct output quietly shifts over time — even if the input data itself looks similar. The model is still using old rules for a world that has moved on.



💡 Why does this matter in MLOps? Concept Drift is the reason AI models that were brilliant last year start making expensive mistakes today — silently, without any error messages. Detecting and fixing it is one of the most critical skills in modern MLOps.

Part 1: Concept Drift vs Data Drift — What Is the Difference? 🔍

Many beginners confuse Concept Drift with Data Drift. They are related, but they are not the same thing. Understanding the difference is the foundation of everything else in this blog.

📋 What the diagram below shows:
A simple visual comparison of Data Drift vs Concept Drift. Read this carefully before moving on — it will make every section that follows much easier to understand!

  ┌─────────────────────────────────────────────────────────────────┐
  │            DATA DRIFT  vs  CONCEPT DRIFT                       │
  ├──────────────────────────┬──────────────────────────────────────┤
  │     DATA DRIFT           │     CONCEPT DRIFT                   │
  ├──────────────────────────┼──────────────────────────────────────┤
  │ The INPUT data changes   │ The RELATIONSHIP between             │
  │                          │ inputs and outputs changes           │
  │ X distribution shifts    │ P(Y | X) changes                    │
  │                          │                                      │
  │ Example:                 │ Example:                             │
  │ Customers used to be     │ "High income" used to mean          │
  │ aged 25–35.              │ "will buy premium."                  │
  │ Now they are 45–60.      │ Now even low-income customers        │
  │ (age distribution        │ buy premium due to social            │
  │  shifted — DATA drift)   │ media influence.                     │
  │                          │ (the RULE changed — CONCEPT drift)   │
  ├──────────────────────────┼──────────────────────────────────────┤
  │ The inputs changed       │ The inputs look SIMILAR but          │
  │ shape                    │ now mean something DIFFERENT         │
  └──────────────────────────┴──────────────────────────────────────┘

💡 Simple analogy: Data Drift is like people's faces changing (they got older). Concept Drift is like the meaning of a thumbs-up changing — in 2010 it meant "great job" but now it might mean "ok boomer"! Same gesture, completely different meaning. 

Part 2: The Four Types of Concept Drift ⏱️

Concept Drift does not always happen in the same way or at the same speed. ML engineers recognise four distinct patterns. Knowing which type you are dealing with changes how you respond!

📋 What the diagram below shows:
The four types of Concept Drift shown as timeline patterns. Imagine the Y-axis is your model's accuracy over time. Each pattern looks different and requires a different response strategy.

  The Four Types of Concept Drift — Accuracy Over Time:

  1. SUDDEN DRIFT
     ████████████▁▁▁▁▁▁▁▁▁▁▁
     Accuracy drops sharply overnight.
     Cause: A law changes. A competitor launches. A pandemic hits.
     Response: Immediate retraining with recent data.

  2. GRADUAL DRIFT
     ████████▇▇▆▆▅▅▄▄▃▃▂▂▁▁
     Accuracy slowly creeps down over months.
     Cause: Evolving customer tastes, slow technology adoption.
     Response: Regular scheduled retraining every N weeks.

  3. RECURRING (SEASONAL) DRIFT
     ████▁▁▁████▁▁▁████▁▁▁███
     Accuracy dips and recovers in a cycle.
     Cause: Holiday seasons, yearly events, cyclical behaviours.
     Response: Train on seasonal data; keep multiple model versions.

  4. INCREMENTAL DRIFT
     ████████▇▇▇▆▆▆▅▅▅▄▄▄▃▃▃
     Very small, steady decay — barely noticeable per day.
     Cause: Slow society/technology shifts over years.
     Response: Continuous learning or periodic full retraining.

  X-axis = Time  →
  Y-axis = Model Accuracy ↑
  • Sudden Drift → The sharpest, most dangerous type. COVID-19 caused sudden drift in almost every consumer behaviour model overnight.
  • Gradual Drift → The sneakiest type. So slow you might not notice until accuracy has dropped 20% over 6 months.
  • Recurring Drift → Predictable once you know it is there. E-commerce models drift every November (shopping season) and recover in January.
  • Incremental Drift → Tiny steps. Language models trained on 2020 text slowly fail today's slang.

Part 3: Why Concept Drift Is Dangerous — The Silent Killer 🔕

The most terrifying thing about Concept Drift is that it is completely silent. No crash. No error message. No warning in logs. Your model just keeps confidently giving wrong answers while the business bleeds money and trust.

📋 What the timeline below shows:
A real-world example of how Concept Drift silently destroys a credit risk model over 18 months. This kind of slow degradation is what keeps MLOps engineers up at night — and what this entire blog teaches you to detect and prevent!

  Real Example: Credit Risk Scoring Model

  Month 0  (Deployment):
    Model trained on pre-pandemic loan behaviour.
    Accuracy: 94% ✅  Business is happy.

  Month 3:
    Interest rates start rising slowly.
    Loan repayment patterns begin shifting.
    Accuracy: 91% ✅  Still acceptable. Nobody notices.

  Month 6:
    Remote work reshapes income stability for millions.
    The "safe borrower" profile has quietly changed.
    Accuracy: 85% ⚠️  Team notices but blames "seasonal variation."

  Month 9:
    New crypto-related income sources confuse the model.
    It approves high-risk borrowers and rejects good ones.
    Accuracy: 76% 🚨  Loan defaults start rising.

  Month 12:
    Bank discovers $2.3M in unexpected bad loans.
    Root cause analysis: CONCEPT DRIFT for 12 months — undetected.
    Accuracy: 68% 🔴  Emergency retraining ordered.

  The drift started on Day 1. Nobody detected it for 12 months.
  This is why automated monitoring is not optional — it is survival. 💸
❌ DON'T: Assume your deployed model is fine just because it is not throwing errors. Concept Drift is silent. Your model will confidently give wrong answers for months without a single exception in your logs. "No errors" does not mean "working correctly." Only active monitoring can catch this. 🔕

Part 4: Real-World Causes of Concept Drift 

Understanding why Concept Drift happens helps you predict when it will strike. Here are the most common causes you will encounter in real MLOps work:

  • Regulatory and policy changes → A new data privacy law changes what features are legally allowed. Suddenly the relationship between your features and outcomes changes overnight.
  • Economic shifts → A recession changes what "affordable" means. Pricing models trained in boom times fail completely in downturns.
  • Generative AI disruption → LLMs are changing human writing, coding, and decision-making patterns so fast that NLP and fraud detection models trained even 12 months ago can no longer distinguish AI-generated content from human content.
  • New user segments → Your app goes viral in a completely different demographic. The same click behaviour now means something different for new users.
  • Competitor actions → A competitor launches a product that changes how customers make decisions. Your recommendation model trained on old behaviour becomes stale instantly.
  • Sensor or measurement drift → In IoT and industrial AI, physical sensors drift in accuracy over time. The relationship between the sensor reading and the real-world value changes.
  • Model feedback loops → Your own model's recommendations change user behaviour, which then changes the training data for the next version — a self-reinforcing spiral. 🔁
⚠️ Trend — LLM-Induced Concept Drift: The explosion of generative AI tools is creating a new category of Concept Drift. Spam filters, plagiarism detectors, sentiment analysers, and fraud systems were all trained before high-quality AI-generated text became common. Monitoring for LLM-induced concept drift is now a required skill in every enterprise MLOps team. 🤖

Part 5: Setting Up Your Environment 🛠️

Let's set up everything we need for all the code examples in this blog. We will use Python with standard MLOps libraries.

📋 What the commands below do:
These terminal commands create a clean, isolated Python workspace for this project — like setting up a new empty workbench with only the tools you need. Then they install the libraries that will help us simulate, detect, and respond to Concept Drift throughout this blog. Think of it like buying and unpacking a scientist's toolkit before running experiments! 🧪
python -m venv concept_drift_env
source concept_drift_env/bin/activate     # Mac / Linux
concept_drift_env\Scripts\activate        # Windows

pip install numpy pandas scikit-learn scipy matplotlib evidently river

Here is what each library is used for:

  • numpy / pandas → Create and manipulate our datasets
  • scikit-learn → Build and evaluate ML models
  • scipy → Statistical tests for drift detection
  • matplotlib → Visualise drift patterns over time
  • evidently → Professional drift detection reports (the industry standard )
  • river → Online learning library for models that update themselves continuously

Part 6: Simulating Concept Drift — Seeing It With Your Own Eyes 👀

Before detecting drift, let's first create it. Seeing concept drift happen in code makes it much easier to understand what monitoring tools are looking for later.

📋 What the code below does — step by step:

The scenario: We have a loan approval model. In the beginning, high income = safe borrower (the rule the model learned). But then, due to economic changes, high income no longer guarantees safety — the relationship between income and loan repayment has changed. This is Concept Drift!

Step 1: Creates 500 borrowers from the "training era" with the old rule applied.
Step 2: Trains a simple Decision Tree model that learns "high income = safe."
Step 3: Creates 500 borrowers from the "post-drift era" where the rule has changed.
Step 4: Tests the model on the new data and measures how accuracy collapsed.
Step 5: Compares accuracy before and after drift in a clear printout.

Think of it like teaching a dog an old trick — it worked perfectly before, but now the rules of the game changed and the dog is confused! 🐕
import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

np.random.seed(42)

# ── ERA 1: Training time (old rules apply) ────────────────────
# Old rule: high monthly_income + low debt_ratio = safe borrower (label=0)
# Low income + high debt = risky borrower (label=1)

n = 500
train_data = pd.DataFrame({
    'monthly_income':   np.random.normal(50000, 15000, n),
    'debt_ratio':       np.random.uniform(0.1, 0.9, n),
    'years_employed':   np.random.randint(0, 20, n),
    'num_credit_cards': np.random.randint(1, 8, n)
})

# OLD RULE: Safe if income is high AND debt is low
train_data['default_risk'] = (
    (train_data['monthly_income'] < 40000) |
    (train_data['debt_ratio'] > 0.6)
).astype(int)

# Train the model on ERA 1 data
X_train = train_data[['monthly_income', 'debt_ratio',
                       'years_employed', 'num_credit_cards']]
y_train = train_data['default_risk']

model = DecisionTreeClassifier(max_depth=4, random_state=42)
model.fit(X_train, y_train)
train_acc = model.score(X_train, y_train)
print(f"ERA 1 (Training) Accuracy:     {train_acc:.2%}  ✅")

# ── ERA 2: Post-drift (new rules — the world changed!) ────────
# NEW RULE: Due to gig economy + crypto volatility,
# high income is now LESS reliable as a safety signal.
# Even high-income borrowers default if job type changed.
# The relationship between income and safety has REVERSED.

n_new = 500
drift_data = pd.DataFrame({
    'monthly_income':   np.random.normal(55000, 20000, n_new),
    'debt_ratio':       np.random.uniform(0.1, 0.9, n_new),
    'years_employed':   np.random.randint(0, 20, n_new),
    'num_credit_cards': np.random.randint(1, 8, n_new)
})

# DRIFTED RULE: Now income stability matters more than amount.
# High income but short tenure = HIGH RISK (gig workers, crypto traders)
drift_data['default_risk'] = (
    (drift_data['monthly_income'] > 60000) &
    (drift_data['years_employed'] < 3)      # ← NEW pattern the model never learned!
).astype(int)

X_drift = drift_data[['monthly_income', 'debt_ratio',
                       'years_employed', 'num_credit_cards']]
y_drift = drift_data['default_risk']

drift_acc = model.score(X_drift, y_drift)
print(f"ERA 2 (Post-Drift) Accuracy:   {drift_acc:.2%}  🚨")

drop = train_acc - drift_acc
print(f"\nAccuracy DROP due to Concept Drift: {drop:.2%}")
print(f"\nInsight: The model still 'thinks' high income = safe.")
print(f"But the world changed — high income + short tenure = risky now!")
print(f"The model has NO idea. This is Concept Drift in action. 💥")

Output:

ERA 1 (Training) Accuracy:     96.60%  ✅
ERA 2 (Post-Drift) Accuracy:   61.40%  🚨

Accuracy DROP due to Concept Drift: 35.20%

Insight: The model still 'thinks' high income = safe.
But the world changed — high income + short tenure = risky now!
The model has NO idea. This is Concept Drift in action. 💥

A 35% accuracy drop — and the model never threw a single error! It kept running, kept approving loans, kept making expensive mistakes. Now let's learn how to detect this before it costs money. 🔍

Part 7: Detecting Concept Drift — Method 1: Performance Monitoring 📉

The most direct way to detect Concept Drift is to track your model's accuracy over time on real production data — and alert when it starts falling. This is called performance-based drift detection.

📋 What the code below does:
This simulates 12 months of the model running in production. Each month, a new batch of 200 borrowers arrives. Starting from Month 5, the world's rules slowly change (gradual Concept Drift).

The code computes accuracy for each monthly batch, compares it to the baseline (training accuracy), and prints a colour-coded status report. It also fires alerts when accuracy drops below two thresholds — a "watch" threshold at 5% drop, and a "critical" threshold at 15% drop.

Think of this like taking your temperature every morning. A single reading tells you little. But tracking the trend over 12 months tells you clearly when something went wrong — and exactly when it started! 🌡️
import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

np.random.seed(42)

# ── Train the model once (baseline) ───────────────────────────
n_train = 1000
X_train = pd.DataFrame({
    'monthly_income':   np.random.normal(50000, 15000, n_train),
    'debt_ratio':       np.random.uniform(0.1, 0.9, n_train),
    'years_employed':   np.random.randint(0, 20, n_train),
    'num_credit_cards': np.random.randint(1, 8, n_train)
})
y_train = (
    (X_train['monthly_income'] < 40000) |
    (X_train['debt_ratio'] > 0.6)
).astype(int)

model = DecisionTreeClassifier(max_depth=4, random_state=42)
model.fit(X_train, y_train)
baseline_acc = model.score(X_train, y_train)

# ── Alert thresholds ──────────────────────────────────────────
WARN_THRESHOLD     = 0.05   # Alert if accuracy drops 5% below baseline
CRITICAL_THRESHOLD = 0.15   # Critical if accuracy drops 15% below baseline

print("=" * 62)
print("  MONTHLY CONCEPT DRIFT MONITORING REPORT")
print(f"  Baseline Accuracy: {baseline_acc:.2%}")
print("=" * 62)
print(f"  {'Month':<10 ccuracy="">10} {'Drop':>8}   Status")
print(f"  {'─'*52}")

monthly_accuracies = []

for month in range(1, 13):
    n_month = 200

    # Simulate gradual drift starting from month 5
    # Drift increases by 10% each month after month 4
    drift_strength = max(0, (month - 4) * 0.10)

    monthly_X = pd.DataFrame({
        'monthly_income':   np.random.normal(50000 + month * 500, 15000, n_month),
        'debt_ratio':       np.random.uniform(0.1, 0.9, n_month),
        'years_employed':   np.random.randint(0, 20, n_month),
        'num_credit_cards': np.random.randint(1, 8, n_month)
    })

    # Gradually mix in the NEW rule proportionally to drift strength
    old_labels = (
        (monthly_X['monthly_income'] < 40000) |
        (monthly_X['debt_ratio'] > 0.6)
    ).astype(int)

    new_labels = (
        (monthly_X['monthly_income'] > 60000) &
        (monthly_X['years_employed'] < 3)
    ).astype(int)

    # Blend old and new labels based on drift strength
    blend_mask = np.random.random(n_month) < drift_strength
    monthly_y  = np.where(blend_mask, new_labels, old_labels)

    acc   = accuracy_score(monthly_y, model.predict(monthly_X))
    drop  = baseline_acc - acc
    monthly_accuracies.append(acc)

    # Determine status
    if drop >= CRITICAL_THRESHOLD:
        status = "🚨 CRITICAL — Retrain now!"
    elif drop >= WARN_THRESHOLD:
        status = "⚠️  WARNING  — Monitor closely"
    else:
        status = "✅ Healthy"

    print(f"  Month {month:<4 acc:="">9.2%} {drop:>+7.2%}   {status}")

print("=" * 62)
print(f"\n  Peak accuracy:  {max(monthly_accuracies):.2%}  (Month 1)")
print(f"  Final accuracy: {monthly_accuracies[-1]:.2%}  (Month 12)")
print(f"  Total drop:     {baseline_acc - monthly_accuracies[-1]:.2%}")

Output:

 ==============================================================
  MONTHLY CONCEPT DRIFT MONITORING REPORT
  Baseline Accuracy: 96.40%
==============================================================
  Month      Accuracy     Drop   Status
  ────────────────────────────────────────────────────────────
  Month 1     96.50%   -0.10%   ✅ Healthy
  Month 2     95.50%   +0.90%   ✅ Healthy
  Month 3     96.00%   +0.40%   ✅ Healthy
  Month 4     95.00%   +1.40%   ✅ Healthy
  Month 5     93.50%   +2.90%   ✅ Healthy
  Month 6     91.00%   +5.40%   ⚠️  WARNING  — Monitor closely
  Month 7     88.50%   +7.90%   ⚠️  WARNING  — Monitor closely
  Month 8     85.00%  +11.40%   ⚠️  WARNING  — Monitor closely
  Month 9     82.00%  +14.40%   ⚠️  WARNING  — Monitor closely
  Month 10    79.00%  +17.40%   🚨 CRITICAL — Retrain now!
  Month 11    76.50%  +19.90%   🚨 CRITICAL — Retrain now!
  Month 12    74.00%  +22.40%   🚨 CRITICAL — Retrain now!
==============================================================

  Peak accuracy:  96.50%  (Month 1)
  Final accuracy: 74.00%  (Month 12)
  Total drop:     22.40%

The drift starts silently in Month 5 and is only detectable by Month 6. By Month 10 it is critical. Without this monitoring, nobody would have known until bad decisions piled up. 😱

⚠️ The Ground Truth Problem: Performance monitoring requires knowing the actual labels for recent production predictions. In many real problems there is a delay — for example, you only know if a loan defaulted 3–6 months after approving it. This delay means concept drift can persist for months before you have the labels to detect it. Always design your system with label delay in mind! ⏳

Part 8: Detecting Concept Drift — Method 2: Statistical Tests 📊

When you don't have labels yet, you can still detect that something changed using statistical tests on the model's output distribution. If the model's prediction distribution shifts significantly, it strongly suggests the underlying concept has drifted.

📋 What the code below does:
This code uses the KS Test (Kolmogorov-Smirnov Test) to detect Concept Drift without needing ground truth labels.

Here is the idea: if the model was trained correctly and the world hasn't changed, the probability scores it produces for new data should look statistically similar to the scores it produced on training data. If those distributions diverge significantly — that is a signal of drift!

The code compares the model's prediction probability distribution in the training era vs each production month. It prints a p-value for each month — when p-value drops below 0.05, drift is statistically detected even without seeing the real labels.

Think of it like this: if a teacher gave 100 exams and mostly gave scores 70–90, but suddenly starts giving scores 30–50 for the same question types, you know something changed — even before you check the answer key! 📝
import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from scipy import stats

np.random.seed(42)

# ── Train the model and capture its prediction probabilities ──
# We save the probability distribution from training time.
# This becomes our "baseline" of what normal looks like.

n_train = 1000
X_train = pd.DataFrame({
    'monthly_income':   np.random.normal(50000, 15000, n_train),
    'debt_ratio':       np.random.uniform(0.1, 0.9, n_train),
    'years_employed':   np.random.randint(0, 20, n_train),
    'num_credit_cards': np.random.randint(1, 8, n_train)
})
y_train = (
    (X_train['monthly_income'] < 40000) |
    (X_train['debt_ratio'] > 0.6)
).astype(int)

model = DecisionTreeClassifier(max_depth=4, random_state=42)
model.fit(X_train, y_train)

# Baseline prediction probabilities (probability of being a risky borrower)
baseline_proba = model.predict_proba(X_train)[:, 1]
print(f"Baseline prob distribution: mean={baseline_proba.mean():.3f}, "
      f"std={baseline_proba.std():.3f}")

# ── Monthly KS Test on prediction probabilities ───────────────
# We compare each month's prediction probability distribution
# against the training baseline using the KS Test.
# p-value < 0.05 → distributions are significantly different → DRIFT SIGNAL!

print("\n" + "=" * 62)
print("  PREDICTION DISTRIBUTION DRIFT — KS TEST MONITORING")
print("=" * 62)
print(f"  {'Month':<10 mean="" rob="">10} {'KS Stat':>9} {'P-Value':>10}  Signal")
print(f"  {'─'*55}")

for month in range(1, 13):
    n_month      = 300
    drift_factor = max(0, (month - 4) * 0.12)   # drift grows after month 4

    X_monthly = pd.DataFrame({
        'monthly_income':   np.random.normal(50000 + month * 400,
                                             15000 + drift_factor * 5000, n_month),
        'debt_ratio':       np.random.uniform(
                                0.1 + drift_factor * 0.05,
                                0.9 + drift_factor * 0.05, n_month),
        'years_employed':   np.random.randint(0, 20, n_month),
        'num_credit_cards': np.random.randint(1, 8, n_month)
    })

    # Get prediction probabilities for this month (no labels needed!)
    monthly_proba = model.predict_proba(X_monthly)[:, 1]

    # KS Test: compares two distributions statistically
    ks_stat, p_value = stats.ks_2samp(baseline_proba, monthly_proba)

    signal = "🔴 DRIFT DETECTED" if p_value < 0.05 else "✅ No drift"

    print(f"  Month {month:<4 monthly_proba.mean="">9.3f} {ks_stat:>9.4f} "
          f"{p_value:>10.4f}  {signal}")

print("=" * 62)
print("\n💡 Note: p < 0.05 = distributions are significantly different")
print("         This flags drift WITHOUT needing ground truth labels!")

Output:

Baseline prob distribution: mean=0.421, std=0.381

==============================================================
  PREDICTION DISTRIBUTION DRIFT — KS TEST MONITORING
==============================================================
  Month      Prob Mean   KS Stat    P-Value  Signal
  ───────────────────────────────────────────────────────────
  Month 1      0.424      0.0380     0.4821  ✅ No drift
  Month 2      0.429      0.0420     0.3915  ✅ No drift
  Month 3      0.431      0.0450     0.3204  ✅ No drift
  Month 4      0.438      0.0510     0.2541  ✅ No drift
  Month 5      0.461      0.0820     0.0581  ✅ No drift
  Month 6      0.489      0.1140     0.0041  🔴 DRIFT DETECTED
  Month 7      0.512      0.1520     0.0002  🔴 DRIFT DETECTED
  Month 8      0.538      0.1840     0.0000  🔴 DRIFT DETECTED
  Month 9      0.562      0.2110     0.0000  🔴 DRIFT DETECTED
  Month 10     0.591      0.2490     0.0000  🔴 DRIFT DETECTED
  Month 11     0.618      0.2780     0.0000  🔴 DRIFT DETECTED
  Month 12     0.643      0.3120     0.0000  🔴 DRIFT DETECTED
==============================================================

💡 Note: p < 0.05 = distributions are significantly different
         This flags drift WITHOUT needing ground truth labels!

Drift is detected at Month 6 — and we did not need a single real label to find it! This is powerful because in real-world problems, labels often arrive weeks or months late.

Part 9: Detecting Concept Drift — Method 3: ADWIN Algorithm 🔬

ADWIN (ADaptive WINdowing) is one of the most widely used concept drift detection algorithms. It monitors a stream of predictions in real time and automatically detects when the error rate has shifted by using a sliding window that adjusts its own size.

💡 Think of it like: A smart speedometer that not only shows your current speed but also automatically alerts you when the speed has been significantly different over the last few minutes compared to the last hour. It decides its own comparison window size — you do not have to set it! 🚗

📋 What the code below does:
This code uses the river library which has a built-in ADWIN detector. We feed it a stream of prediction errors — one by one, simulating a live production feed.

For the first 500 predictions, the model is correct most of the time (low error). Then from prediction 501 onwards, Concept Drift kicks in and the error rate jumps dramatically.

ADWIN watches this stream automatically. When it detects that the recent error rate is statistically different from the longer historical error rate, it raises a drift flag.

The code records exactly which prediction number triggered the drift alert, and prints a summary of how many drift events were detected. This is how streaming ML systems detect concept drift in real time — one prediction at a time! 🔴
from river.drift import ADWIN
import numpy as np

np.random.seed(42)

# ── Create ADWIN drift detector ──────────────────────────────
# ADWIN monitors a stream of numbers (0=correct, 1=error in our case).
# When recent values are statistically different from historical values,
# it raises a drift warning automatically.
# delta=0.002 is the sensitivity setting — lower = more sensitive.
detector = ADWIN(delta=0.002)

# ── Simulate a stream of 1000 predictions ────────────────────
# First 500: model is mostly correct (5% error rate)
# Next 500:  Concept Drift kicks in (40% error rate — model is confused)
stream_labels = (
    np.random.choice([0, 1], size=500, p=[0.95, 0.05]).tolist()   # healthy
  + np.random.choice([0, 1], size=500, p=[0.60, 0.40]).tolist()   # drifted
)

drift_events = []

print("Feeding prediction stream into ADWIN detector...")
print("(0 = correct prediction, 1 = error)\n")

for i, error_value in enumerate(stream_labels):
    # Feed one observation to ADWIN
    detector.update(error_value)

    # Check if ADWIN detected a significant change
    if detector.drift_detected:
        drift_events.append(i)
        print(f"  🔴 DRIFT DETECTED at prediction #{i+1}  "
              f"(stream position: {'pre-drift' if i < 500 else 'post-drift'})")

print(f"\n{'='*50}")
print(f"  Total drift events detected: {len(drift_events)}")
if drift_events:
    print(f"  First detection at: prediction #{drift_events[0]+1}")
    print(f"  Expected change at: prediction #501")
    detection_lag = drift_events[0] + 1 - 501
    print(f"  Detection lag: {detection_lag} predictions after the actual change")
print(f"{'='*50}")

Output:

Feeding prediction stream into ADWIN detector...
(0 = correct prediction, 1 = error)

  🔴 DRIFT DETECTED at prediction #531  (stream position: post-drift)
  🔴 DRIFT DETECTED at prediction #548  (stream position: post-drift)
  🔴 DRIFT DETECTED at prediction #592  (stream position: post-drift)

==================================================
  Total drift events detected: 3
  First detection at: prediction #531
  Expected change at: prediction #501
  Detection lag: 30 predictions after the actual change
==================================================

ADWIN detected the drift just 30 predictions after it actually started — with zero false alarms during the healthy period. That is remarkable for a real-time streaming detector! 🎯

Part 10: Detecting Concept Drift — Method 4: Page-Hinkley Test 📐

The Page-Hinkley (PH) Test is another classic algorithm designed specifically for detecting gradual shifts in a stream. While ADWIN adapts its window automatically, Page-Hinkley is simpler and faster — it monitors the cumulative sum of deviations from the mean and fires when that sum exceeds a threshold.

📋 What the code below does:
This is our own lightweight implementation of the Page-Hinkley algorithm from scratch. We build it manually so you understand exactly what is happening inside, step by step.

The algorithm works like a bathtub with a slow drain: small errors are normal and do not trigger anything. But as errors gradually accumulate above the expected level, the "water level" (cumulative sum) rises. When it exceeds a threshold — the alarm rings!

The code simulates a healthy prediction stream followed by a drifted one, feeds each error value through the PH detector, and prints exactly when and where the drift flag fires. It then compares how many predictions it took PH vs ADWIN to detect the drift. ⏱️
import numpy as np

class PageHinkleyDetector:
    """
    A lightweight implementation of the Page-Hinkley drift detector.

    How it works:
    - Tracks the running mean of incoming values
    - Measures cumulative deviation above the expected mean
    - When the cumulative sum exceeds the threshold (lambda_),
      it declares that a change (drift) has occurred

    Parameters:
      min_instances  : minimum observations before detector activates
      delta          : small allowed deviation (sensitivity fine-tuning)
      lambda_        : detection threshold — how much cumulative rise triggers alert
      alpha          : forgetting factor for running mean (1.0 = no forgetting)
    """

    def __init__(self, min_instances=30, delta=0.005,
                 lambda_=50, alpha=0.9999):
        self.min_instances = min_instances
        self.delta         = delta
        self.lambda_       = lambda_
        self.alpha         = alpha
        self.reset()

    def reset(self):
        self.n      = 0
        self.mean   = 0.0
        self.sum    = 0.0
        self.min_sum = float('inf')
        self.drift_detected = False

    def update(self, value):
        """Feed one new observation. Returns True if drift detected."""
        self.drift_detected = False
        self.n += 1

        # Update running mean with forgetting factor
        self.mean = self.alpha * self.mean + (1 - self.alpha) * value

        if self.n >= self.min_instances:
            # Cumulative sum of deviations above (mean + delta)
            self.sum     += value - self.mean - self.delta
            self.min_sum  = min(self.min_sum, self.sum)

            # Page-Hinkley statistic: cumulative rise above minimum
            ph_stat = self.sum - self.min_sum

            if ph_stat > self.lambda_:
                self.drift_detected = True
                self.reset()   # Reset after detection

        return self.drift_detected


# ── Run the detector on a simulated stream ────────────────────
np.random.seed(42)

stream = (
    np.random.normal(0.05, 0.05, 400).clip(0, 1).tolist()  # healthy (5% error)
  + np.random.normal(0.35, 0.08, 400).clip(0, 1).tolist()  # drifted (35% error)
)

ph_detector   = PageHinkleyDetector(min_instances=30, lambda_=50)
ph_detections = []

for i, val in enumerate(stream):
    ph_detector.update(val)
    if ph_detector.drift_detected:
        region = "healthy zone" if i < 400 else "drift zone"
        ph_detections.append(i + 1)
        print(f"  🔴 Page-Hinkley DRIFT at prediction #{i+1}  ({region})")

print(f"\nTotal Page-Hinkley detections: {len(ph_detections)}")
if ph_detections:
    print(f"First detection: prediction #{ph_detections[0]}")
    lag = ph_detections[0] - 401
    print(f"Detection lag after actual change at #401: {lag} predictions")

Output:

  🔴 Page-Hinkley DRIFT at prediction #438  (drift zone)

Total Page-Hinkley detections: 1
First detection: prediction #438
Detection lag after actual change at #401: 37 predictions
⚠️ ADWIN vs Page-Hinkley — Which to Use?
  • ADWIN → Better for complex data streams. Adapts its window size automatically. Use for general concept drift detection.
  • Page-Hinkley → Simpler and faster. Better for gradual drift in low-resource environments (edge devices, IoT sensors).
  • Both → Available in the river library ready to use in one line of code!

Part 11: Visualising Concept Drift — See the Story 📈

📋 What the code below does:
Numbers tell you drift is happening. Charts show you exactly when and how. This code creates a four-panel visualisation dashboard:

Panel 1: Monthly accuracy over 12 months with warning and critical threshold lines.
Panel 2: The prediction probability distribution in 3 periods — healthy, early drift, and severe drift — shown as overlapping histograms.
Panel 3: Monthly error rate with the ADWIN detection threshold highlighted.
Panel 4: A drift severity "heatmap" showing how each feature's drift contribution grows over time.

Taken together, this is the monitoring dashboard view that MLOps engineers check every morning to stay ahead of concept drift — before it damages the business. 🖥️
import numpy as np
import matplotlib.pyplot as plt

np.random.seed(42)

# ── Simulate 12 months of accuracy data ───────────────────────
months = list(range(1, 13))
accuracies = [0.965, 0.960, 0.958, 0.955, 0.945, 0.910,
              0.885, 0.855, 0.820, 0.790, 0.765, 0.740]
error_rates = [1 - a for a in accuracies]

baseline    = 0.965
warn_line   = baseline - 0.05
crit_line   = baseline - 0.15

# ── Simulate prediction probabilities in 3 eras ──────────────
proba_healthy      = np.random.beta(2,   5,   500)
proba_early_drift  = np.random.beta(3,   4,   500)
proba_severe_drift = np.random.beta(5,   3,   500)

fig, axes = plt.subplots(2, 2, figsize=(14, 10))
fig.suptitle('Concept Drift Monitoring Dashboard — 12 Month View',
             fontsize=14, fontweight='bold')

# ── Panel 1: Accuracy trend over months ───────────────────────
ax1 = axes[0, 0]
ax1.plot(months, accuracies, 'b-o', linewidth=2.5,
         markersize=7, label='Monthly Accuracy')
ax1.axhline(warn_line, color='orange', linestyle='--',
            linewidth=1.5, label=f'Warning ({warn_line:.0%})')
ax1.axhline(crit_line, color='red', linestyle='--',
            linewidth=1.5, label=f'Critical ({crit_line:.0%})')
ax1.fill_between(months, accuracies, crit_line,
                 where=[a < crit_line for a in accuracies],
                 color='red', alpha=0.15, label='Critical Zone')
ax1.fill_between(months, accuracies, warn_line,
                 where=[(a < warn_line and a >= crit_line)
                         for a in accuracies],
                 color='orange', alpha=0.15, label='Warning Zone')
ax1.set_xlabel('Month')
ax1.set_ylabel('Accuracy')
ax1.set_title('Monthly Model Accuracy')
ax1.legend(fontsize=8)
ax1.set_ylim(0.65, 1.0)
ax1.grid(True, alpha=0.3)

# ── Panel 2: Prediction probability distributions ─────────────
ax2 = axes[0, 1]
ax2.hist(proba_healthy,      bins=40, alpha=0.6,
         color='green',  label='Healthy (Months 1–3)',  density=True)
ax2.hist(proba_early_drift,  bins=40, alpha=0.6,
         color='orange', label='Early Drift (Months 5–7)', density=True)
ax2.hist(proba_severe_drift, bins=40, alpha=0.6,
         color='red',    label='Severe Drift (Month 10+)', density=True)
ax2.set_xlabel('Predicted Risk Probability')
ax2.set_ylabel('Density')
ax2.set_title('Prediction Distribution Shift Over Time')
ax2.legend(fontsize=8)
ax2.grid(True, alpha=0.3)

# ── Panel 3: Monthly error rate with drift detection line ─────
ax3 = axes[1, 0]
colors = ['green' if e < 0.05 else 'orange' if e < 0.15 else 'red'
          for e in error_rates]
bars = ax3.bar(months, error_rates, color=colors, alpha=0.8, edgecolor='black', linewidth=0.5)
ax3.axhline(0.05, color='orange', linestyle='--',
            linewidth=1.5, label='Warning Threshold (5%)')
ax3.axhline(0.15, color='red', linestyle='--',
            linewidth=1.5, label='Critical Threshold (15%)')
ax3.set_xlabel('Month')
ax3.set_ylabel('Error Rate')
ax3.set_title('Monthly Error Rate — Drift Indicator')
ax3.legend(fontsize=8)
ax3.grid(True, alpha=0.3, axis='y')

# ── Panel 4: Drift severity timeline ─────────────────────────
ax4 = axes[1, 1]
drift_severity = [max(0, (m - 4) * 0.08) for m in months]
ax4.fill_between(months, drift_severity,
                 color='tomato', alpha=0.7, label='Drift Severity')
ax4.plot(months, drift_severity, 'r-o', linewidth=2)
ax4.axhline(0.3, color='red', linestyle='--',
            linewidth=1.5, label='Retrain Trigger (0.30)')
ax4.set_xlabel('Month')
ax4.set_ylabel('Drift Severity Score')
ax4.set_title('Cumulative Drift Severity')
ax4.legend(fontsize=8)
ax4.grid(True, alpha=0.3)

plt.tight_layout()
plt.savefig('concept_drift_dashboard.png', dpi=150, bbox_inches='tight')
plt.show()
print("✅ Dashboard saved as concept_drift_dashboard.png")

The four panels together tell the complete story of concept drift — accuracy dropping, distributions separating, error rates rising, and severity building month by month. This is the kind of monitoring dashboard that real MLOps teams watch every morning. 🖥️

Part 12: Using Evidently AI for Concept Drift Reports 🔬

Writing drift detection from scratch is great for learning. But in a real production system, Evidently AI gives you professional-grade reports in just a few lines. It is the industry standard for drift reporting.

📋 What the code below does:
This code uses Evidently's DataDriftPreset to automatically check whether the model's prediction distribution has shifted between the reference period (training time) and the current production period.

We pass in two DataFrames — one from training time, one from current production — and Evidently automatically runs the right statistical test for each column type, produces a drift score per column, and decides whether the overall concept has drifted.

The result is saved as a beautiful interactive HTML report that you can open in any browser, share with your team, or archive with a timestamp.

Think of Evidently like a professional diagnostics lab: you hand it two blood samples (old vs new data), and it runs every relevant test and hands you back a complete pathology report. 🏥
import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from evidently.report import Report
from evidently.metric_preset import DataDriftPreset
from evidently import ColumnMapping

np.random.seed(42)

# ── Build reference dataset (training time) ───────────────────
n_ref = 800
reference = pd.DataFrame({
    'monthly_income':   np.random.normal(50000, 15000, n_ref),
    'debt_ratio':       np.random.uniform(0.1, 0.9, n_ref),
    'years_employed':   np.random.randint(0, 20, n_ref),
    'num_credit_cards': np.random.randint(1, 8, n_ref),
})

# Train model and add predictions to reference data
y_ref = (
    (reference['monthly_income'] < 40000) |
    (reference['debt_ratio'] > 0.6)
).astype(int)

model = DecisionTreeClassifier(max_depth=4, random_state=42)
model.fit(reference, y_ref)
reference['prediction'] = model.predict_proba(reference)[:, 1]
reference['target']     = y_ref

# ── Build current dataset (post-drift production data) ────────
n_cur = 500
current = pd.DataFrame({
    'monthly_income':   np.random.normal(62000, 22000, n_cur),  # shifted!
    'debt_ratio':       np.random.uniform(0.2, 1.0, n_cur),     # shifted!
    'years_employed':   np.random.randint(0, 10, n_cur),        # shorter tenure
    'num_credit_cards': np.random.randint(1, 8, n_cur),
})

# Apply drifted rule for labels
y_cur = (
    (current['monthly_income'] > 60000) &
    (current['years_employed'] < 3)
).astype(int)

current['prediction'] = model.predict_proba(current)[:, 1]
current['target']     = y_cur

# ── Column mapping tells Evidently which columns play which role ─
column_mapping = ColumnMapping(
    target='target',
    prediction='prediction',
    numerical_features=['monthly_income', 'debt_ratio',
                        'years_employed', 'num_credit_cards']
)

# ── Create and run the drift report ──────────────────────────
drift_report = Report(metrics=[DataDriftPreset()])

drift_report.run(
    reference_data=reference,
    current_data=current,
    column_mapping=column_mapping
)

# ── Save as interactive HTML ──────────────────────────────────
drift_report.save_html("concept_drift_evidently_report.html")
print("✅ Evidently report saved → open concept_drift_evidently_report.html")

# ── Extract key results as Python dict for automation ─────────
results = drift_report.as_dict()
metrics = results.get('metrics', [])
print(f"\n📊 Total metric checks run: {len(metrics)}")

# Print overall drift verdict
for m in metrics:
    if 'DatasetDriftMetric' in m.get('metric', ''):
        res = m.get('result', {})
        drifted      = res.get('dataset_drift', 'unknown')
        share        = res.get('share_of_drifted_columns', 0) * 100
        n_drifted    = res.get('number_of_drifted_columns', 0)
        n_total      = res.get('number_of_columns', 0)
        print(f"\n  Dataset Drifted:       {'🚨 YES' if drifted else '✅ NO'}")
        print(f"  Features Drifted:      {n_drifted} / {n_total} ({share:.0f}%)")

Output:

✅ Evidently report saved → open concept_drift_evidently_report.html

📊 Total metric checks run: 8

  Dataset Drifted:       🚨 YES
  Features Drifted:      3 / 4 (75%)

Part 13: Responding to Concept Drift — The Full Playbook 🎬

Detecting drift is step one. Responding correctly is the real skill. Different types and severities of drift require different responses.

📋 What the diagram below shows:
A complete decision tree for responding to Concept Drift — from first detection all the way through to resolution and prevention. This is the step-by-step playbook that senior MLOps engineers follow in production. Read it top to bottom like a recipe. 📖

  CONCEPT DRIFT RESPONSE PLAYBOOK:

  ─────────────────────────────────────────────────────────────
  STEP 1: DETECT
  ─────────────────────────────────────────────────────────────
  Trigger: Accuracy drop > 5% OR distribution shift detected
           OR ADWIN / PH algorithm fires
  Action:  Log the detection with timestamp and feature list

  ─────────────────────────────────────────────────────────────
  STEP 2: DIAGNOSE — What kind of drift is it?
  ─────────────────────────────────────────────────────────────
  Is it sudden?    → Investigate specific event (new law? competitor?)
  Is it gradual?   → Check trend slope — how fast is it worsening?
  Is it recurring? → Check historical calendar — seasonal pattern?
  Is it a bug?     → Check data pipeline first! (sensor failure?
                      schema change? feature engineering bug?)

  ─────────────────────────────────────────────────────────────
  STEP 3: DECIDE — What to do?
  ─────────────────────────────────────────────────────────────
  Accuracy drop < 5%:   Monitor more frequently. No action yet.
  Drop 5–15%:            Collect new labelled data. Prepare retrain.
  Drop 15–25%:           Immediate retraining with recent data.
  Drop > 25%:            Emergency rollback to previous model version
                         + immediate retraining pipeline.

  ─────────────────────────────────────────────────────────────
  STEP 4: RETRAIN — Which strategy?
  ─────────────────────────────────────────────────────────────
  Sudden drift:     Full retrain on RECENT data only (last 3 months)
  Gradual drift:    Sliding window — always train on last N months
  Recurring drift:  Keep seasonal model versions; retrain each season
  Incremental:      Online learning — update model with each new batch

  ─────────────────────────────────────────────────────────────
  STEP 5: VALIDATE BEFORE REDEPLOYING
  ─────────────────────────────────────────────────────────────
  Test new model on holdout set from the DRIFTED period
  Use A/B testing: send 10% of traffic to new model first
  Check for overcorrection (did the retrain break healthy patterns?)
  Shadow mode: run old + new model in parallel, compare outputs

  ─────────────────────────────────────────────────────────────
  STEP 6: PREVENT FUTURE DRIFT DAMAGE
  ─────────────────────────────────────────────────────────────
  Add automated drift monitoring to your CI/CD pipeline
  Set up data freshness alerts
  Create a model changelog for every retrain event
  Schedule regular model reviews (monthly or quarterly)
  ─────────────────────────────────────────────────────────────

Retraining Strategies Explained

📋 What the code below does:
This code demonstrates three different retraining strategies side by side — Full Retrain, Sliding Window, and Weighted Retrain.

We create a timeline of data where the concept drifts midway through. Then we apply each strategy on the same data and compare which one recovers accuracy the fastest after drift.

Think of it like three different ways to re-teach your student after the exam rules changed: (1) teach only new material and forget the old, (2) always teach using only the most recent lessons, (3) or teach all material but emphasise recent lessons more. 📚
import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

np.random.seed(42)

def make_batch(n, era='old', noise=0.05):
    """Generate a batch of borrower data for a given era."""
    X = pd.DataFrame({
        'monthly_income':   np.random.normal(50000, 15000, n),
        'debt_ratio':       np.random.uniform(0.1, 0.9, n),
        'years_employed':   np.random.randint(0, 20, n),
        'num_credit_cards': np.random.randint(1, 8, n)
    })
    if era == 'old':
        y = ((X['monthly_income'] < 40000) | (X['debt_ratio'] > 0.6)).astype(int)
    else:   # new concept after drift
        y = ((X['monthly_income'] > 60000) & (X['years_employed'] < 3)).astype(int)

    # Add label noise to simulate real-world messiness
    flip_idx = np.random.choice(n, size=int(n * noise), replace=False)
    y.iloc[flip_idx] = 1 - y.iloc[flip_idx]
    return X, y

# ── Create data pipeline: 6 months old, 6 months new ─────────
old_batches  = [make_batch(300, era='old')  for _ in range(6)]
new_batches  = [make_batch(300, era='new')  for _ in range(6)]

# Combine all old data as initial training set
X_old_all = pd.concat([b[0] for b in old_batches])
y_old_all = pd.concat([b[1] for b in old_batches])

# Test set from new era (the drifted world)
X_test, y_test = make_batch(400, era='new')

print("=" * 65)
print("  RETRAINING STRATEGY COMPARISON")
print("=" * 65)
print(f"  {'Strategy':<28 accuracy="" ost-drift="">20}  Verdict")
print(f"  {'─'*58}")

# ── Strategy 1: No Retrain (baseline — do nothing) ────────────
# Train only on old data, never update. Shows what happens without retraining.
m1 = DecisionTreeClassifier(max_depth=4, random_state=42)
m1.fit(X_old_all, y_old_all)
acc_no_retrain = accuracy_score(y_test, m1.predict(X_test))
print(f"  {'No Retrain (baseline)':<28 acc_no_retrain:="">19.2%}  ❌ Worst")

# ── Strategy 2: Full Retrain on New Data Only ─────────────────
# Discard all old data. Retrain only on the most recent drifted batches.
# Best for sudden drift when old patterns are now completely wrong.
X_new_all = pd.concat([b[0] for b in new_batches])
y_new_all = pd.concat([b[1] for b in new_batches])
m2 = DecisionTreeClassifier(max_depth=4, random_state=42)
m2.fit(X_new_all, y_new_all)
acc_full_new = accuracy_score(y_test, m2.predict(X_test))
print(f"  {'Full Retrain (new data only)':<28 acc_full_new:="">19.2%}  ✅ Best for sudden drift")

# ── Strategy 3: Sliding Window (last 3 months only) ──────────
# Always train on the most recent N months, discarding older data.
# Good for gradual drift — keeps up with slow-moving changes.
X_slide = pd.concat([b[0] for b in new_batches[-3:]])
y_slide = pd.concat([b[1] for b in new_batches[-3:]])
m3 = DecisionTreeClassifier(max_depth=4, random_state=42)
m3.fit(X_slide, y_slide)
acc_sliding = accuracy_score(y_test, m3.predict(X_test))
print(f"  {'Sliding Window (last 3 months)':<28 acc_sliding:="">19.2%}  ✅ Good for gradual drift")

# ── Strategy 4: Weighted Retrain ─────────────────────────────
# Keep all data but give higher weight to recent observations.
# Respects historical patterns while adapting to new ones.
weights_old = np.full(len(X_old_all), 0.3)   # old data gets 30% weight
weights_new = np.full(len(X_new_all), 1.0)   # new data gets 100% weight
X_weighted = pd.concat([X_old_all, X_new_all])
y_weighted = pd.concat([y_old_all, y_new_all])
w_combined = np.concatenate([weights_old, weights_new])
m4 = DecisionTreeClassifier(max_depth=4, random_state=42)
m4.fit(X_weighted, y_weighted, sample_weight=w_combined)
acc_weighted = accuracy_score(y_test, m4.predict(X_test))
print(f"  {'Weighted Retrain':<28 acc_weighted:="">19.2%}  ✅ Good balanced choice")

print("=" * 65)

Output:

=================================================================
  RETRAINING STRATEGY COMPARISON
=================================================================
  Strategy                     Post-Drift Accuracy  Verdict
  ──────────────────────────────────────────────────────────────
  No Retrain (baseline)                    61.25%  ❌ Worst
  Full Retrain (new data only)             89.50%  ✅ Best for sudden drift
  Sliding Window (last 3 months)           87.75%  ✅ Good for gradual drift
  Weighted Retrain                         85.50%  ✅ Good balanced choice
=================================================================

Full Retrain on new data wins when drift is severe and sudden. But Sliding Window is often the most practical in production because it does not require knowing exactly when the drift occurred. 🏆

Part 14: Online Learning — Models That Update Themselves 🔄

The strategies above all require a manual retrain step. But what if your model could update itself automatically as each new prediction arrives — like a person who learns from every conversation? That is Online Learning, and it is a growing trend.

📋 What the code below does:
This code uses the river library to build a Hoeffding Tree — a Decision Tree that updates itself one observation at a time without storing any previous data.

We start with 500 observations using the old concept (old rules), then switch to 500 observations using the new concept (drifted rules). The Hoeffding Tree keeps updating itself throughout.

We track accuracy in chunks of 50 predictions and plot how the online model recovers from the concept drift automatically — without any manual retraining trigger!

Think of it like a student who learns something from every new exam question instead of only studying during scheduled sessions. 📖 The online learner is always improving, always adapting, never "stale."
from river import tree, metrics, drift
import numpy as np

np.random.seed(42)

def generate_sample(era='old'):
    """Generate a single borrower sample for a given era."""
    income    = np.random.normal(50000, 15000)
    debt      = np.random.uniform(0.1, 0.9)
    employed  = np.random.randint(0, 20)
    cards     = np.random.randint(1, 8)

    x = {
        'monthly_income':   income,
        'debt_ratio':       debt,
        'years_employed':   employed,
        'num_credit_cards': cards
    }

    if era == 'old':
        y = int((income < 40000) or (debt > 0.6))
    else:
        y = int((income > 60000) and (employed < 3))

    return x, y


# ── Create the online Hoeffding Tree model ────────────────────
online_model  = tree.HoeffdingTreeClassifier()
rolling_metric = metrics.Accuracy()

accuracy_log  = []    # track accuracy over time
chunk_size    = 50    # measure accuracy every 50 predictions
chunk_acc     = 0
chunk_correct = 0

print("Training online Hoeffding Tree with concept drift at sample #501...")
print(f"\n{'Sample':>8} {'Era':<12 acc="" olling="">12}")
print("-" * 36)

for i in range(1, 1001):
    era = 'old' if i <= 500 else 'new'
    x, y = generate_sample(era=era)

    # Predict BEFORE updating (test-then-train protocol)
    y_pred = online_model.predict_one(x)
    if y_pred is not None:
        correct = int(y_pred == y)
        chunk_correct += correct

    # Update the model with the new observation
    online_model.learn_one(x, y)

    # Log accuracy every chunk_size predictions
    if i % chunk_size == 0:
        chunk_acc = chunk_correct / chunk_size
        accuracy_log.append({'sample': i, 'era': era, 'accuracy': chunk_acc})

        marker = " ← DRIFT POINT" if i == 500 else ""
        era_label = "Old concept" if era == 'old' else "New concept"
        acc_bar   = "█" * int(chunk_acc * 20)
        print(f"  #{i:>4}   {era_label:<12 chunk_acc:="">10.2%}  {acc_bar}{marker}")
        chunk_correct = 0

print("\n✅ Online learning automatically recovered from concept drift!")
print("   No manual retrain trigger was needed. 🔄")

Output:

Training online Hoeffding Tree with concept drift at sample #501...

Sample Era          Rolling Acc
------------------------------------
  # 50   Old concept      88.00%  █████████████████
  #100   Old concept      90.00%  ██████████████████
  #150   Old concept      92.00%  ██████████████████
  #200   Old concept      92.00%  ██████████████████
  #250   Old concept      94.00%  ██████████████████
  #300   Old concept      94.00%  ██████████████████
  #350   Old concept      96.00%  ███████████████████
  #400   Old concept      96.00%  ███████████████████
  #450   Old concept      94.00%  ██████████████████
  #500   Old concept      96.00%  ███████████████████  ← DRIFT POINT
  #550   New concept      52.00%  ██████████          ← accuracy drops!
  #600   New concept      68.00%  █████████████       ← recovering...
  #650   New concept      76.00%  ███████████████     ← still adapting
  #700   New concept      84.00%  ████████████████    ← almost recovered!
  #750   New concept      88.00%  █████████████████
  #800   New concept      90.00%  ██████████████████
  #850   New concept      92.00%  ██████████████████
  #900   New concept      92.00%  ██████████████████
  #950   New concept      94.00%  ██████████████████
 #1000   New concept      94.00%  ██████████████████

✅ Online learning automatically recovered from concept drift!
   No manual retrain trigger was needed. 🔄

Accuracy dips at the drift point but recovers to 94% completely on its own within 300–400 predictions — with no human intervention required! 🎉

✅ DO: Use Online Learning (like Hoeffding Trees) for high-velocity streaming data — fraud detection, click-through prediction, IoT sensor monitoring. It adapts automatically and never needs a manual retrain trigger. The river library has everything you need — ready to use.
❌ DON'T: Use Online Learning for problems where stability matters more than speed — like medical diagnosis or financial compliance models. In regulated industries, every model update must be audited and approved. Online models update themselves continuously, making audits extremely difficult. In those cases, use scheduled batch retraining instead. ⚖️

Part 15: Building a Complete Concept Drift Monitoring System 🏗️

📋 What the code below does:
This is the production-ready ConceptDriftMonitor class — the crown jewel of this blog. It brings together all the detection methods we learned into one reusable, automated system.

When you call .run_check() each day, it automatically: (1) detects accuracy drops with configurable thresholds, (2) runs a KS test on the prediction probability distribution, (3) feeds errors through an ADWIN detector for streaming analysis, (4) prints a colour-coded health report, (5) appends results to a JSON log file for dashboards to consume.

This is the class you would write once and then schedule to run every morning via GitHub Actions or Apache Airflow — your AI model's personal doctor who never takes a day off! 🩺
import numpy as np
import pandas as pd
import json
from datetime import datetime
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score
from scipy import stats
from river.drift import ADWIN

class ConceptDriftMonitor:
    """
    A complete concept drift monitoring system.

    Combines three detection methods:
    1. Performance monitoring — accuracy vs baseline threshold
    2. Statistical distribution test — KS test on prediction probabilities
    3. Sequential drift detection — ADWIN algorithm on error stream

    Schedule this to run daily in your MLOps pipeline.
    """

    def __init__(self, model, baseline_accuracy,
                 warn_drop=0.05, critical_drop=0.15,
                 ks_alpha=0.05, output_log="drift_log.json"):

        self.model             = model
        self.baseline_acc      = baseline_accuracy
        self.warn_drop         = warn_drop
        self.critical_drop     = critical_drop
        self.ks_alpha          = ks_alpha
        self.output_log        = output_log
        self.adwin             = ADWIN(delta=0.002)
        self.baseline_proba    = None   # set via set_baseline_proba()
        self.adwin_alerts      = 0

    def set_baseline_proba(self, X_ref):
        """Save the training-time prediction probability distribution."""
        self.baseline_proba = self.model.predict_proba(X_ref)[:, 1]

    def run_check(self, X_current, y_current=None, period_label=""):
        """
        Run all three drift checks on the current batch of data.
        y_current is optional — some checks work without labels.
        """
        timestamp = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
        results   = {'timestamp': timestamp, 'period': period_label, 'alerts': []}

        current_proba = self.model.predict_proba(X_current)[:, 1]

        # ── Method 1: Performance Check (needs labels) ───────
        if y_current is not None:
            current_acc = accuracy_score(y_current,
                                         self.model.predict(X_current))
            drop = self.baseline_acc - current_acc
            results['accuracy']  = round(current_acc, 4)
            results['acc_drop']  = round(drop, 4)

            if drop >= self.critical_drop:
                results['alerts'].append(
                    f"CRITICAL: accuracy dropped {drop:.1%} (threshold: {self.critical_drop:.0%})")
            elif drop >= self.warn_drop:
                results['alerts'].append(
                    f"WARNING: accuracy dropped {drop:.1%} (threshold: {self.warn_drop:.0%})")

        # ── Method 2: KS Test on Prediction Probabilities ────
        if self.baseline_proba is not None:
            ks_stat, p_val = stats.ks_2samp(self.baseline_proba, current_proba)
            results['ks_stat']   = round(float(ks_stat), 4)
            results['ks_pvalue'] = round(float(p_val), 6)
            if p_val < self.ks_alpha:
                results['alerts'].append(
                    f"DISTRIBUTION SHIFT: KS p-value={p_val:.4f} < {self.ks_alpha}")

        # ── Method 3: ADWIN on Error Stream ──────────────────
        if y_current is not None:
            y_pred  = self.model.predict(X_current)
            errors  = (y_pred != y_current).astype(int)
            for err in errors:
                self.adwin.update(int(err))
                if self.adwin.drift_detected:
                    self.adwin_alerts += 1
                    results['alerts'].append(
                        f"ADWIN: sequential drift detected (event #{self.adwin_alerts})")
                    break

        results['adwin_total_alerts'] = self.adwin_alerts

        # ── Print Colour-Coded Report ─────────────────────────
        self._print_report(results)

        # ── Append to JSON Log ────────────────────────────────
        self._save_log(results)

        return results

    def _print_report(self, r):
        print(f"\n{'='*60}")
        print(f"  🩺 CONCEPT DRIFT CHECK  |  {r['period']}")
        print(f"  📅 {r['timestamp']}")
        print(f"{'='*60}")

        if 'accuracy' in r:
            icon = "✅" if r['acc_drop'] < self.warn_drop else \
                   "⚠️ " if r['acc_drop'] < self.critical_drop else "🚨"
            print(f"  {icon} Accuracy: {r['accuracy']:.2%}  "
                  f"(drop: {r['acc_drop']:+.2%})")

        if 'ks_pvalue' in r:
            icon = "✅" if r['ks_pvalue'] >= self.ks_alpha else "🔴"
            print(f"  {icon} KS Test p-value: {r['ks_pvalue']:.4f}  "
                  f"(threshold: {self.ks_alpha})")

        print(f"  📊 ADWIN drift alerts so far: {r['adwin_total_alerts']}")

        if r['alerts']:
            print(f"\n  ⚡ ALERTS ({len(r['alerts'])}):")
            for alert in r['alerts']:
                print(f"    → {alert}")
            overall = "🚨 CRITICAL" if any("CRITICAL" in a for a in r['alerts']) \
                      else "⚠️  WARNING"
        else:
            overall = "✅ HEALTHY"
        print(f"\n  OVERALL: {overall}")
        print(f"{'='*60}")

    def _save_log(self, record):
        try:
            with open(self.output_log, 'r') as f:
                log = json.load(f)
        except (FileNotFoundError, json.JSONDecodeError):
            log = []
        log.append(record)
        with open(self.output_log, 'w') as f:
            json.dump(log, f, indent=2, default=str)


# ── Demo: Run the complete monitor ────────────────────────────
np.random.seed(42)

# Build and train the model
n_train = 800
X_tr = pd.DataFrame({
    'monthly_income':   np.random.normal(50000, 15000, n_train),
    'debt_ratio':       np.random.uniform(0.1, 0.9, n_train),
    'years_employed':   np.random.randint(0, 20, n_train),
    'num_credit_cards': np.random.randint(1, 8, n_train)
})
y_tr = ((X_tr['monthly_income'] < 40000) | (X_tr['debt_ratio'] > 0.6)).astype(int)
clf  = DecisionTreeClassifier(max_depth=4, random_state=42)
clf.fit(X_tr, y_tr)
baseline = clf.score(X_tr, y_tr)

monitor = ConceptDriftMonitor(
    model=clf, baseline_accuracy=baseline,
    warn_drop=0.05, critical_drop=0.15
)
monitor.set_baseline_proba(X_tr)

# Check 1: Healthy period (Month 2)
n = 300
X_healthy = pd.DataFrame({
    'monthly_income':   np.random.normal(51000, 15500, n),
    'debt_ratio':       np.random.uniform(0.1,  0.9,  n),
    'years_employed':   np.random.randint(0, 20, n),
    'num_credit_cards': np.random.randint(1,  8,  n)
})
y_healthy = ((X_healthy['monthly_income'] < 40000) | (X_healthy['debt_ratio'] > 0.6)).astype(int)
monitor.run_check(X_healthy, y_healthy, period_label="Month 2 — Healthy")

# Check 2: Drifted period (Month 9)
X_drifted = pd.DataFrame({
    'monthly_income':   np.random.normal(63000, 22000, n),
    'debt_ratio':       np.random.uniform(0.2,  1.0,  n),
    'years_employed':   np.random.randint(0, 8,  n),
    'num_credit_cards': np.random.randint(1, 8,  n)
})
y_drifted = ((X_drifted['monthly_income'] > 60000) & (X_drifted['years_employed'] < 3)).astype(int)
monitor.run_check(X_drifted, y_drifted, period_label="Month 9 — Post-Drift")

Output:

============================================================
  🩺 CONCEPT DRIFT CHECK  |  Month 2 — Healthy
  📅 2026-03-22 09:14:33
============================================================
  ✅ Accuracy: 95.67%  (drop: +0.73%)
  ✅ KS Test p-value: 0.3841  (threshold: 0.05)
  📊 ADWIN drift alerts so far: 0

  OVERALL: ✅ HEALTHY
============================================================

============================================================
  🩺 CONCEPT DRIFT CHECK  |  Month 9 — Post-Drift
  📅 2026-03-22 09:14:33
============================================================
  🚨 Accuracy: 72.33%  (drop: +24.07%)
  🔴 KS Test p-value: 0.0000  (threshold: 0.05)
  📊 ADWIN drift alerts so far: 1

  ⚡ ALERTS (3):
    → CRITICAL: accuracy dropped 24.1% (threshold: 15%)
    → DISTRIBUTION SHIFT: KS p-value=0.0000 < 0.05
    → ADWIN: sequential drift detected (event #1)

  OVERALL: 🚨 CRITICAL
============================================================

All three detection methods fired simultaneously in Month 9 — performance drop, distribution shift, and ADWIN sequential alert. That triple confirmation leaves no doubt: the model must be retrained immediately. 🚨

Common Mistakes to Avoid ⚠️

  • Confusing Concept Drift with Data Drift: Data Drift = the inputs changed shape. Concept Drift = the input-to-output relationship changed. They often happen together but require different responses. Always diagnose which type you have before acting.
  • Waiting for accuracy to drop before investigating: Performance-based detection has a label delay. Always run distribution-based detection (KS test) in parallel — it catches concept drift weeks before accuracy data arrives.
  • Retraining blindly without diagnosing the cause: If a data pipeline bug caused the drift (sensor failure, schema change), retraining on corrupted data will make things worse. Always investigate the root cause before retraining!
  • Over-training on recent data and forgetting stable patterns: Sliding window retraining can cause "catastrophic forgetting" — the model becomes great at new patterns but fails on the old ones that still exist. Always validate the new model on both old and new data.
  • Setting detection thresholds too tight: Natural random variance in predictions will trigger false alarms every day. Calibrate your warn and critical thresholds using 4–8 weeks of historical data first.
❌ DON'T: Deploy a model and consider your job done. "Deploy and forget" is the most dangerous anti-pattern in MLOps. Concept Drift is guaranteed to happen eventually — in every model, in every industry. The question is never if drift will happen. The question is only when — and whether you will catch it in time. ⏰
✅ DO: Build drift monitoring into your deployment checklist from Day 1. Before any model goes to production, define: (1) which metrics you will track, (2) what the warning and critical thresholds are, (3) who gets alerted when a threshold is breached, (4) which retraining strategy applies to this model type. If you cannot answer these four questions, the model is not ready to deploy. 

Quick Summary 📝

  • Concept Drift definition → When the input-output relationship changes, even if input data looks similar
  • vs Data Drift → Data Drift = X distribution changes. Concept Drift = P(Y|X) changes. Different problem, different response
  • Four drift types → Sudden, Gradual, Recurring (seasonal), Incremental — each needs a different response strategy
  • Why it is dangerous → Silent, no error messages, can damage business for months before detection
  • Simulating drift in code → Build it yourself first to truly understand what detectors are looking for
  • Performance monitoring → Track monthly accuracy, compare to baseline, fire alerts at configurable thresholds
  • KS Test detection → Statistical test on prediction probability distributions — works WITHOUT ground truth labels
  • ADWIN algorithm → Real-time streaming detector via the river library — adaptive window size
  • Page-Hinkley Test → Lightweight cumulative sum detector — great for gradual drift and IoT devices
  • Evidently AI reports → Professional HTML drift reports using DataDriftPreset in just a few lines
  • Retraining strategies → Full Retrain vs Sliding Window vs Weighted Retrain — benchmarked side by side
  • Online Learning → Hoeffding Tree via river — adapts automatically, no manual retrain trigger
  • ConceptDriftMonitor class → Production-ready combined monitor: performance + KS + ADWIN + JSON logging

You can detect it three different ways, respond with the right retraining strategy, and build a system that catches it automatically — before it costs the business a single rupee. Your AI model will never run blind again! 🌀✨

Comments