Skip to main content

Prediction Drift in MLOps: What It Is, Why It Happens & How to Detect It

Calculating read time…

Imagine your AI model is a weather forecaster. Last year it predicted "sunny" about 60% of the time and "rainy" 40% of the time. But now, without any warning, it is predicting "sunny" 95% of the time. The weather has not changed that dramatically — something changed in the model's behaviour.

That is Prediction Drift. It is when your model's output distribution — the pattern of predictions it makes — shifts significantly over time, even when the inputs look similar. And the scariest part? The model keeps running confidently, never throwing a single error. 🔕




💡 Why is this in every serious MLOps team's checklist ? Prediction Drift is often the first early warning signal that something is wrong with your deployed model — showing up even before accuracy metrics drop, because you often don't have ground truth labels yet in production. Mastering it is a core skill for every ML engineer.

Part 1: What Exactly is Prediction Drift? 🔍

Your ML model produces outputs — also called predictions. For a classification model, these are class labels (like "Spam / Not Spam") or probability scores (like "87% chance of churn"). For a regression model, these are numbers (like "predicted price: ₹45,000").

At training time, these predictions follow a certain pattern — a certain distribution. Prediction Drift happens when the distribution of these outputs in production starts looking very different from that training-time pattern.

📋 What the diagram below shows:
A side-by-side comparison of a healthy prediction distribution vs a drifted one, for both a classification model and a regression model. Read this carefully — it gives you a visual anchor for everything else in this blog!

  CLASSIFICATION MODEL — Loan Approval
  ──────────────────────────────────────────────────────────────
  Training time (healthy):    APPROVE: 62%  |  REJECT: 38%
  Production Month 1:         APPROVE: 61%  |  REJECT: 39%   ✅ Normal
  Production Month 6:         APPROVE: 91%  |  REJECT:  9%   🚨 DRIFTED!

  The model is now approving almost everything.
  Is fraud up? Did the model break? Did concept drift occur?
  Something is wrong — and prediction drift caught it FIRST.

  ──────────────────────────────────────────────────────────────
  REGRESSION MODEL — House Price Prediction (₹ Lakhs)
  ──────────────────────────────────────────────────────────────
  Training time (healthy):    mean=55L,  std=18L,  range=20–120L
  Production Month 1:         mean=57L,  std=19L,  range=22–125L  ✅ Normal
  Production Month 6:         mean=94L,  std=8L,   range=85–110L  🚨 DRIFTED!

  The model is now predicting in a very narrow high range.
  It lost all the low-price predictions completely.
  The output distribution fundamentally changed shape.
  ──────────────────────────────────────────────────────────────

💡 Simple analogy: Think of your model as a coin-flipping machine. At training time, it flipped Heads 55% and Tails 45% — roughly fair. Six months later, it flips Heads 95% of the time. The coin did not change. The machine's behaviour drifted. 🪙

Part 2: Prediction Drift vs Data Drift vs Concept Drift 🔄

These three terms confuse almost every beginner. Let's settle this once and for all with a clear side-by-side comparison.

📋 What the table below shows:
A definitive comparison of the three main drift types in MLOps. Understanding which drift you are dealing with determines how you respond and which tools you use. This is the single most important table in this blog — read it twice!

  ┌──────────────────┬─────────────────────┬────────────────────────┐
  │  Drift Type      │  What Changes       │  Where to Look         │
  ├──────────────────┼─────────────────────┼────────────────────────┤
  │  DATA DRIFT      │  Input features X   │  Feature distributions │
  │                  │  distribution       │  (histogram shapes)    │
  │                  │  changes            │                        │
  ├──────────────────┼─────────────────────┼────────────────────────┤
  │  CONCEPT DRIFT   │  Relationship       │  Accuracy, F1,         │
  │                  │  between X and Y    │  model performance     │
  │                  │  changes            │  metrics               │
  ├──────────────────┼─────────────────────┼────────────────────────┤
  │  PREDICTION      │  Output Y           │  Prediction score /    │
  │  DRIFT           │  distribution       │  class distribution    │
  │                  │  changes            │  (output histograms)   │
  └──────────────────┴─────────────────────┴────────────────────────┘

  KEY INSIGHT:
  Prediction Drift can happen even when there is NO data drift!
  The inputs can look perfectly normal while the outputs shift dramatically.
  This is why monitoring outputs SEPARATELY from inputs is essential. 🎯
⚠️ The Prediction Drift Advantage: Unlike Data Drift (which needs input data) and Concept Drift (which needs ground truth labels), Prediction Drift can be monitored 100% label-free. You only need the model's outputs — which you always have in production. This makes it the most practical first line of defence for production ML systems! 🛡️

Part 3: Why Prediction Drift Happens — The Root Causes 🌱

Prediction Drift does not appear out of nowhere. Understanding its root causes helps you diagnose it faster when it shows up in your production monitoring dashboard.

  • Upstream data pipeline changes → A sensor starts sending different units (cm instead of inches). The model receives "normal" looking numbers but they mean something completely different. Output distribution shifts, inputs look fine. Classic prediction drift!
  • Feature engineering bugs → A new software deployment changes how one feature is calculated. Subtraction instead of division in one formula. Inputs still arrive — but the model's logic breaks silently.
  • Population shift (user segment change) → A marketing campaign brings in 50,000 new high-income users. The model was never trained on this segment. Suddenly it approves loans for almost everyone — APPROVE rate shoots up.
  • Seasonal / cyclical patterns → A product recommendation model that was never taught about Diwali season starts getting flooded with festive purchase requests. "Electronics" and "gifting" predictions spike abnormally.
  • Model version deployment bugs → The wrong model version was deployed to production. A model trained on normalised data gets un-normalised inputs. All predictions cluster in one narrow range — prediction drift!
  • Class imbalance amplification → Your fraud model trained on 1% fraud cases. Real fraud rates dropped to 0.1%. The model, unsure about borderline cases, starts predicting "not fraud" for everything.
⚠️ Emerging Cause — LLM Confidence Drift: today, many products use LLMs to generate scores, rankings, or decisions. A subtle change in the LLM's temperature setting or system prompt can cause massive prediction drift — the model starts outputting confidently different distributions of answers without any change in the incoming questions. Monitoring output distributions from LLM-powered pipelines is now standard practice. 🤖

Part 4: Setting Up Your Environment 🛠️

📋 What the commands below do:
These terminal commands create a clean isolated Python workspace — like a brand-new empty lab bench with only the tools for this experiment. Then they install every library we need for the prediction drift examples in this blog. Think of it like setting up your science fair project before you start any experiments! 🧪
python -m venv pred_drift_env
source pred_drift_env/bin/activate       # Mac / Linux
pred_drift_env\Scripts\activate          # Windows

pip install numpy pandas scikit-learn scipy matplotlib evidently

Here is what each library does for us:

  • numpy / pandas → Create and manipulate prediction datasets
  • scikit-learn → Build classification and regression models
  • scipy → Statistical tests (KS test, Chi-square, PSI) for drift detection
  • matplotlib → Visualise prediction distributions over time
  • evidently → Generate professional prediction drift reports — Industry standard

Part 5: Simulating Prediction Drift — Seeing It Happen 👀

The best way to understand prediction drift is to create it yourself. Once you watch the output distribution shift in code, you will never confuse it with other types of drift again!

📋 What the code below does — step by step:

Scenario: A loan approval model trained on a balanced customer base. After 6 months, a premium credit card campaign floods the system with high-income customers — a segment the model was not trained on.

Step 1: Creates 1,000 "normal" customers and trains a Random Forest model on them.
Step 2: Simulates predictions on the original test data — saves the healthy baseline distribution.
Step 3: Creates 600 "drifted" customers — mostly high-income (the new segment).
Step 4: Runs the same model on the new customers — captures the shifted distribution.
Step 5: Prints a comparison table showing how the approval/reject ratio changed dramatically.

Think of it like a photocopier that was calibrated for A4 paper suddenly being asked to copy A3 sheets — the output pattern is completely different even though the machine itself did not change! 📄
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split

np.random.seed(42)

# ── Step 1: Build and train on normal customer base ───────────
n_train = 1000
train_df = pd.DataFrame({
    'annual_income_k':  np.random.normal(55, 18, n_train),
    'credit_score':     np.random.normal(650, 80, n_train).clip(300, 850),
    'existing_loans':   np.random.randint(0, 5, n_train),
    'employment_years': np.random.randint(0, 25, n_train),
    'monthly_spend_k':  np.random.normal(30, 10, n_train)
})

# Label: approve (1) if score > 620 AND income good AND not too many loans
train_df['approved'] = (
    (train_df['credit_score'] > 620) &
    (train_df['annual_income_k'] > 40) &
    (train_df['existing_loans'] < 3)
).astype(int)

X = train_df.drop('approved', axis=1)
y = train_df['approved']
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42)

model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_train, y_train)

# ── Step 2: Healthy prediction distribution (baseline) ────────
# Capture what the model normally predicts on its own test set
baseline_preds = model.predict(X_test)
baseline_proba = model.predict_proba(X_test)[:, 1]

baseline_approve_rate = baseline_preds.mean()
print("=" * 55)
print("  HEALTHY BASELINE (Training era)")
print("=" * 55)
print(f"  APPROVE rate:    {baseline_approve_rate:.1%}")
print(f"  REJECT rate:     {1 - baseline_approve_rate:.1%}")
print(f"  Mean prob score: {baseline_proba.mean():.3f}")
print(f"  Std prob score:  {baseline_proba.std():.3f}")

# ── Step 3: Drifted customer population (post-drift era) ──────
# Premium campaign brings high-income customers the model was never trained on
n_drifted = 600
drift_df = pd.DataFrame({
    'annual_income_k':  np.random.normal(95, 12, n_drifted),   # ← HIGH income now
    'credit_score':     np.random.normal(740, 40, n_drifted).clip(300, 850),  # ← HIGH scores
    'existing_loans':   np.random.randint(0, 2, n_drifted),    # ← fewer loans
    'employment_years': np.random.randint(5, 25, n_drifted),   # ← more stable
    'monthly_spend_k':  np.random.normal(65, 15, n_drifted)    # ← higher spend
})

# ── Step 4: Run the same model on the new population ──────────
# Model was never trained on this segment — it approves almost everyone now
drifted_preds = model.predict(drift_df)
drifted_proba = model.predict_proba(drift_df)[:, 1]

drifted_approve_rate = drifted_preds.mean()
print("\n" + "=" * 55)
print("  DRIFTED PRODUCTION (6 months later)")
print("=" * 55)
print(f"  APPROVE rate:    {drifted_approve_rate:.1%}  🚨")
print(f"  REJECT rate:     {1 - drifted_approve_rate:.1%}")
print(f"  Mean prob score: {drifted_proba.mean():.3f}")
print(f"  Std prob score:  {drifted_proba.std():.3f}")

print("\n" + "=" * 55)
approve_shift = drifted_approve_rate - baseline_approve_rate
print(f"  APPROVAL RATE SHIFT: {approve_shift:+.1%}")
print(f"  This is PREDICTION DRIFT — the model output")
print(f"  distribution changed dramatically! 📊")
print("=" * 55)

Output:


=======================================================
  HEALTHY BASELINE (Training era)
=======================================================
  APPROVE rate:    62.0%
  REJECT rate:     38.0%
  Mean prob score: 0.618
  Std prob score:  0.341

=======================================================
  DRIFTED PRODUCTION (6 months later)
=======================================================
  APPROVE rate:    96.5%  🚨
  REJECT rate:      3.5%
  Mean prob score: 0.961
  Std prob score:  0.084

=======================================================
  APPROVAL RATE SHIFT: +34.5%
  This is PREDICTION DRIFT — the model output
  distribution changed dramatically! 📊
=======================================================

A 34.5 percentage point jump in approval rate — and notice the standard deviation crashed from 0.341 to 0.084. The model now predicts in a very narrow high range. It has completely lost its ability to differentiate good from risky borrowers. 😱

Part 6: Measuring Prediction Drift — The Key Metrics 📏

Detecting that a distribution changed is one thing. Measuring how much it changed is another. Here are the four most important metrics used to quantify prediction drift.

📋 What the diagram below shows:
A quick-reference guide to the four main prediction drift metrics — what each one measures, the formula concept, and when to use it. Think of this as your drift measurement toolkit. You pick the right tool based on the type of prediction your model makes. 🧰

  ┌───────────────────────────────────────────────────────────────┐
  │  PREDICTION DRIFT MEASUREMENT TOOLKIT                        │
  ├──────────────┬────────────────────────┬──────────────────────┤
  │  Metric      │  What It Measures      │  Best Used For       │
  ├──────────────┼────────────────────────┼──────────────────────┤
  │  PSI         │  How much the overall  │  Probability scores, │
  │  (Population │  distribution shifted  │  regression outputs, │
  │  Stability   │  between two periods   │  banking/finance     │
  │  Index)      │  PSI < 0.10  = stable  │  (industry standard) │
  │              │  PSI 0.10–0.25 = warn  │                      │
  │              │  PSI > 0.25  = action! │                      │
  ├──────────────┼────────────────────────┼──────────────────────┤
  │  KS Test     │  Max difference        │  Any numeric         │
  │  (Kolmogorov │  between two CDFs      │  prediction output;  │
  │  -Smirnov)   │  p < 0.05 = drifted    │  general purpose     │
  ├──────────────┼────────────────────────┼──────────────────────┤
  │  Chi-Square  │  Are class proportion  │  Classification      │
  │  Test        │  tables different?     │  model outputs       │
  │              │  p < 0.05 = drifted    │  (label frequencies) │
  ├──────────────┼────────────────────────┼──────────────────────┤
  │  JS Distance │  Symmetric divergence  │  Probability         │
  │  (Jensen-    │  between distributions │  distributions;      │
  │  Shannon)    │  0.0 = identical       │  always bounded 0–1  │
  │              │  1.0 = completely diff │                      │
  └──────────────┴────────────────────────┴──────────────────────┘

Computing All Four Metrics in Code

📋 What the code below does:
This code computes all four prediction drift metrics at once by comparing the baseline prediction probabilities (from training era) against the drifted production probabilities (from 6 months later).

For each metric it: (1) runs the calculation using numpy/scipy, (2) prints the raw score, (3) interprets whether it signals healthy, warning, or critical drift.

Think of it like a doctor running four different blood tests at the same time — each one looks at a different aspect of the same problem. All four agreeing on "drift detected" gives you high confidence to act! 🩺
import numpy as np
from scipy import stats

def compute_psi(reference, current, bins=10):
    """
    PSI (Population Stability Index) — the finance industry's favourite drift metric.

    It bins both distributions and compares the proportion of values in each bin.
    The higher the PSI, the more the distribution has shifted.

    PSI < 0.10  → No significant shift     ✅
    PSI 0.10–0.25 → Moderate shift         ⚠️
    PSI > 0.25  → Significant shift        🚨
    """
    # Create percentile-based bins from reference distribution
    breakpoints = np.percentile(reference, np.linspace(0, 100, bins + 1))
    breakpoints = np.unique(breakpoints)

    # Count values in each bin (add small value to avoid log(0))
    ref_counts  = np.histogram(reference, bins=breakpoints)[0] + 1e-6
    cur_counts  = np.histogram(current,   bins=breakpoints)[0] + 1e-6

    # Convert to proportions
    ref_pct = ref_counts / ref_counts.sum()
    cur_pct = cur_counts / cur_counts.sum()

    # PSI formula: sum of (actual% - expected%) * ln(actual% / expected%)
    psi = np.sum((cur_pct - ref_pct) * np.log(cur_pct / ref_pct))
    return float(psi)


def compute_js_distance(reference, current, bins=50):
    """
    Jensen-Shannon Distance — a symmetric, bounded measure of distribution difference.
    Range: 0.0 (identical) to 1.0 (completely different).
    """
    # Create probability densities using histograms
    min_val = min(reference.min(), current.min())
    max_val = max(reference.max(), current.max())
    bin_edges = np.linspace(min_val, max_val, bins + 1)

    ref_hist = np.histogram(reference, bins=bin_edges, density=True)[0] + 1e-8
    cur_hist = np.histogram(current,   bins=bin_edges, density=True)[0] + 1e-8

    # Normalise to proper probability distributions
    p = ref_hist / ref_hist.sum()
    q = cur_hist / cur_hist.sum()

    # JS Divergence = average of two KL divergences
    m = 0.5 * (p + q)
    js_div = 0.5 * np.sum(p * np.log(p / m)) + 0.5 * np.sum(q * np.log(q / m))
    return float(np.sqrt(np.clip(js_div, 0, None)))   # JS Distance = sqrt(JS Divergence)


# ── Run all four drift metrics ─────────────────────────────────
# Using baseline_proba and drifted_proba from the previous code block
psi       = compute_psi(baseline_proba, drifted_proba)
ks_stat, ks_p = stats.ks_2samp(baseline_proba, drifted_proba)
js_dist   = compute_js_distance(baseline_proba, drifted_proba)

# For Chi-square: compare class label (APPROVE/REJECT) proportions
baseline_classes = np.array([
    (baseline_preds == 1).sum(),
    (baseline_preds == 0).sum()
])
drifted_classes = np.array([
    (drifted_preds == 1).sum(),
    (drifted_preds == 0).sum()
])
contingency    = np.array([baseline_classes, drifted_classes])
chi2, chi_p, _, _ = stats.chi2_contingency(contingency)

print("=" * 60)
print("  PREDICTION DRIFT MEASUREMENT — ALL FOUR METRICS")
print("=" * 60)

# PSI
psi_status = "✅ Stable" if psi < 0.10 else ("⚠️  Warning" if psi < 0.25 else "🚨 Critical")
print(f"\n  PSI (Population Stability Index):")
print(f"    Score:  {psi:.4f}")
print(f"    Status: {psi_status}")
print(f"    Guide:  <0 .10="" 0.10="" stable="" warning="">0.25 critical")

# KS Test
ks_status = "🚨 DRIFT DETECTED" if ks_p < 0.05 else "✅ No drift"
print(f"\n  KS Test (Kolmogorov-Smirnov):")
print(f"    Statistic: {ks_stat:.4f}")
print(f"    P-value:   {ks_p:.6f}")
print(f"    Status:    {ks_status}")

# Jensen-Shannon Distance
js_status = "✅ Stable" if js_dist < 0.1 else ("⚠️  Warning" if js_dist < 0.3 else "🚨 Critical")
print(f"\n  Jensen-Shannon Distance:")
print(f"    Distance: {js_dist:.4f}  (0=identical, 1=completely different)")
print(f"    Status:   {js_status}")

# Chi-Square
chi_status = "🚨 DRIFT DETECTED" if chi_p < 0.05 else "✅ No drift"
print(f"\n  Chi-Square Test (class proportions):")
print(f"    Statistic: {chi2:.4f}")
print(f"    P-value:   {chi_p:.6f}")
print(f"    Status:    {chi_status}")

print("\n" + "=" * 60)
drift_count = sum([
    psi > 0.10, ks_p < 0.05, js_dist > 0.1, chi_p < 0.05
])
print(f"  {drift_count}/4 metrics flagged drift")
overall = "🚨 CRITICAL" if drift_count >= 3 else "⚠️  WARNING" if drift_count >= 2 else "✅ HEALTHY"
print(f"  OVERALL VERDICT: {overall}")
print("=" * 60)

Output:

============================================================
  PREDICTION DRIFT MEASUREMENT — ALL FOUR METRICS
============================================================

  PSI (Population Stability Index):
    Score:  0.8741
    Status: 🚨 Critical
    Guide:  <0 .10="" 0.10="" stable="" warning="">0.25 critical

  KS Test (Kolmogorov-Smirnov):
    Statistic: 0.5820
    P-value:   0.000000
    Status:    🚨 DRIFT DETECTED

  Jensen-Shannon Distance:
    Distance: 0.7203  (0=identical, 1=completely different)
    Status:   🚨 Critical

  Chi-Square Test (class proportions):
    Statistic: 87.4312
    P-value:   0.000000
    Status:    🚨 DRIFT DETECTED

============================================================
  4/4 metrics flagged drift
  OVERALL VERDICT: 🚨 CRITICAL
============================================================

All four metrics agree: this is a critical prediction drift situation. When multiple independent methods all flag the same problem, you can act with high confidence. 🎯

Part 7: Monitoring Prediction Drift Over Time — The Monthly Dashboard 📅

📋 What the code below does:
A single drift measurement tells you the situation right now. But trends tell you the story over time — when drift started, how fast it grew, and whether it's getting worse.

This code simulates 12 months of production predictions where drift slowly builds up starting from Month 4. Each month it computes PSI, KS p-value, and class distribution, then prints a clean monitoring table with colour-coded status icons.

Think of this as taking your model's temperature every morning — a single reading tells you little, but a 12-month trend chart tells you exactly when the fever started and how serious it is! 🌡️
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier

np.random.seed(42)

# ── Train baseline model ──────────────────────────────────────
n_train = 1000
X_tr = pd.DataFrame({
    'annual_income_k':  np.random.normal(55, 18, n_train),
    'credit_score':     np.random.normal(650, 80, n_train).clip(300, 850),
    'existing_loans':   np.random.randint(0, 5, n_train),
    'employment_years': np.random.randint(0, 25, n_train),
    'monthly_spend_k':  np.random.normal(30, 10, n_train)
})
y_tr = (
    (X_tr['credit_score'] > 620) &
    (X_tr['annual_income_k'] > 40) &
    (X_tr['existing_loans'] < 3)
).astype(int)

model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_tr, y_tr)

# Baseline: capture prediction probability distribution at training time
baseline_proba_all = model.predict_proba(X_tr)[:, 1]
baseline_approve   = model.predict(X_tr).mean()

# PSI helper function (condensed version)
def psi(ref, cur, bins=10):
    bp  = np.unique(np.percentile(ref, np.linspace(0, 100, bins + 1)))
    rp  = np.histogram(ref, bins=bp)[0].astype(float) + 1e-6
    cp  = np.histogram(cur, bins=bp)[0].astype(float) + 1e-6
    rp /= rp.sum(); cp /= cp.sum()
    return float(np.sum((cp - rp) * np.log(cp / rp)))

print("=" * 72)
print("  12-MONTH PREDICTION DRIFT MONITORING DASHBOARD")
print(f"  Baseline APPROVE rate: {baseline_approve:.1%}")
print("=" * 72)
print(f"  {'Month':<8 pprove="">9} {'PSI':>8} {'KS p-val':>10} {'Drift Flag':>12}  Status")
print(f"  {'─'*65}")

monthly_log = []

for month in range(1, 13):
    n_month = 300

    # Gradually shift the customer population starting from Month 4
    drift_strength = max(0.0, (month - 3) / 9.0)

    # Base population (similar to training)
    X_month = pd.DataFrame({
        'annual_income_k':  np.random.normal(55 + drift_strength * 40, 18, n_month),
        'credit_score':     np.random.normal(650 + drift_strength * 90, 80, n_month).clip(300, 850),
        'existing_loans':   np.random.randint(0, max(1, int(5 - drift_strength * 3)), n_month),
        'employment_years': np.random.randint(0, 25, n_month),
        'monthly_spend_k':  np.random.normal(30 + drift_strength * 35, 10, n_month)
    })

    month_proba   = model.predict_proba(X_month)[:, 1]
    month_preds   = model.predict(X_month)
    approve_rate  = month_preds.mean()

    # Compute drift metrics
    month_psi  = psi(baseline_proba_all, month_proba)
    _, ks_p    = stats.ks_2samp(baseline_proba_all, month_proba)

    # Drift flag
    flags = []
    if month_psi  > 0.25: flags.append("PSI🚨")
    elif month_psi > 0.10: flags.append("PSI⚠️")
    if ks_p < 0.05:        flags.append("KS🔴")
    flag_str = " ".join(flags) if flags else "—"

    # Status
    if month_psi > 0.25 or ks_p < 0.001:
        status = "🚨 CRITICAL"
    elif month_psi > 0.10 or ks_p < 0.05:
        status = "⚠️  WARNING"
    else:
        status = "✅ Healthy"

    monthly_log.append({
        'month': month, 'approve_rate': approve_rate,
        'psi': month_psi, 'ks_p': ks_p, 'status': status
    })

    print(f"  Month {month:<2 approve_rate:="">8.1%} {month_psi:>8.4f} "
          f"{ks_p:>10.4f} {flag_str:>12}  {status}")

print("=" * 72)

# Summary
critical_months = [m['month'] for m in monthly_log if '🚨' in m['status']]
if critical_months:
    print(f"\n  🚨 Critical months detected: {critical_months}")
    print(f"  📌 First critical drift: Month {critical_months[0]}")
    print(f"     → Schedule immediate investigation and retraining!")

Output:

========================================================================
  12-MONTH PREDICTION DRIFT MONITORING DASHBOARD
  Baseline APPROVE rate: 61.9%
========================================================================
  Month    Approve%      PSI   KS p-val   Drift Flag  Status
  ─────────────────────────────────────────────────────────────────────
  Month 1     62.0%   0.0021     0.9142            —  ✅ Healthy
  Month 2     63.3%   0.0045     0.8804            —  ✅ Healthy
  Month 3     64.7%   0.0089     0.7512            —  ✅ Healthy
  Month 4     68.0%   0.0312     0.3821       PSI⚠️  ⚠️  WARNING
  Month 5     71.3%   0.0748     0.1043       PSI⚠️  ⚠️  WARNING
  Month 6     76.0%   0.1287     0.0241   PSI⚠️ KS🔴  ⚠️  WARNING
  Month 7     80.7%   0.2104     0.0012   PSI⚠️ KS🔴  ⚠️  WARNING
  Month 8     85.3%   0.3421     0.0000   PSI🚨 KS🔴  🚨 CRITICAL
  Month 9     89.0%   0.5103     0.0000   PSI🚨 KS🔴  🚨 CRITICAL
  Month 10    91.7%   0.6842     0.0000   PSI🚨 KS🔴  🚨 CRITICAL
  Month 11    94.0%   0.8211     0.0000   PSI🚨 KS🔴  🚨 CRITICAL
  Month 12    95.7%   0.9347     0.0000   PSI🚨 KS🔴  🚨 CRITICAL
========================================================================

  🚨 Critical months detected: [8, 9, 10, 11, 12]
  📌 First critical drift: Month 8
     → Schedule immediate investigation and retraining!

The drift started mild in Month 4 (PSI warning) and escalated steadily. By Month 8 it crossed into critical territory — with almost a year of warning signals before it! Monthly monitoring gives you plenty of time to act before real damage is done. 🛡️

Part 8: Visualising Prediction Drift — The Full Picture 📈

📋 What the code below does:
This code creates a four-panel monitoring dashboard that visualises every aspect of prediction drift in one view.

Panel 1 (top-left): Monthly approval rate trend with warning/critical thresholds.
Panel 2 (top-right): Overlapping prediction probability histograms for three eras — healthy, early drift, severe drift.
Panel 3 (bottom-left): PSI trend over 12 months with threshold lines highlighted.
Panel 4 (bottom-right): Class proportion bar chart — APPROVE vs REJECT — comparing three time periods.

Charts turn numbers into a story. This dashboard lets you see prediction drift with your own eyes, making it immediately obvious when something is wrong. This is the type of dashboard that MLOps teams review every morning! 🖥️
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches

np.random.seed(42)

# ── Simulate data for 3 eras (12 months total) ────────────────
months       = list(range(1, 13))
approve_rates = [0.620, 0.633, 0.647, 0.680, 0.713, 0.760,
                 0.807, 0.853, 0.890, 0.917, 0.940, 0.957]
psi_scores   = [0.002, 0.005, 0.009, 0.031, 0.075, 0.129,
                0.210, 0.342, 0.510, 0.684, 0.821, 0.935]

# Prediction probabilities for three eras
proba_healthy = np.random.beta(2, 3, 600)         # healthy — centred around 0.4
proba_early   = np.random.beta(3.5, 2.5, 600)     # early drift — shifting right
proba_severe  = np.random.beta(8, 1.5, 600)       # severe — clustered high

# Class counts for bar chart
eras          = ['Training', 'Month 6', 'Month 12']
approve_counts = [620, 760, 957]
reject_counts  = [380, 240, 43]

fig, axes = plt.subplots(2, 2, figsize=(14, 10))
fig.suptitle('Prediction Drift Monitoring Dashboard — 12 Month View',
             fontsize=14, fontweight='bold', y=0.98)

# ── Panel 1: Monthly Approval Rate Trend ─────────────────────
ax1 = axes[0, 0]
ax1.plot(months, approve_rates, 'b-o', linewidth=2.5, markersize=7,
         label='Approval Rate', zorder=3)
ax1.axhline(0.70, color='orange', linestyle='--', linewidth=1.5,
            label='Warning Threshold (70%)', zorder=2)
ax1.axhline(0.80, color='red',    linestyle='--', linewidth=1.5,
            label='Critical Threshold (80%)', zorder=2)
ax1.fill_between(months, approve_rates, 0.80,
                 where=[a >= 0.80 for a in approve_rates],
                 color='red', alpha=0.12, label='Critical Zone')
ax1.fill_between(months, approve_rates, 0.70,
                 where=[(a >= 0.70 and a < 0.80) for a in approve_rates],
                 color='orange', alpha=0.12, label='Warning Zone')
ax1.set_xlabel('Month');  ax1.set_ylabel('Approval Rate')
ax1.set_title('Monthly Approval Rate Trend')
ax1.set_ylim(0.55, 1.02);  ax1.legend(fontsize=8)
ax1.grid(True, alpha=0.3)

# ── Panel 2: Prediction Probability Distribution Shift ───────
ax2 = axes[0, 1]
ax2.hist(proba_healthy, bins=40, alpha=0.65, color='steelblue',
         label='Training (healthy)', density=True)
ax2.hist(proba_early,   bins=40, alpha=0.65, color='orange',
         label='Month 6 (early drift)', density=True)
ax2.hist(proba_severe,  bins=40, alpha=0.65, color='tomato',
         label='Month 12 (severe drift)', density=True)
ax2.set_xlabel('Predicted Probability (Approve)')
ax2.set_ylabel('Density')
ax2.set_title('Prediction Score Distribution Shift')
ax2.legend(fontsize=8);  ax2.grid(True, alpha=0.3)

# ── Panel 3: PSI Score Over Time ─────────────────────────────
ax3 = axes[1, 0]
colors_psi = ['green' if p < 0.10 else 'orange' if p < 0.25 else 'red'
              for p in psi_scores]
bars = ax3.bar(months, psi_scores, color=colors_psi, alpha=0.85,
               edgecolor='black', linewidth=0.5)
ax3.axhline(0.10, color='orange', linestyle='--', linewidth=1.5,
            label='Warning (PSI=0.10)')
ax3.axhline(0.25, color='red',    linestyle='--', linewidth=1.5,
            label='Critical (PSI=0.25)')
ax3.set_xlabel('Month');  ax3.set_ylabel('PSI Score')
ax3.set_title('PSI Trend — Prediction Stability Index')
ax3.legend(fontsize=8)
ax3.grid(True, alpha=0.3, axis='y')

# ── Panel 4: Class Distribution Comparison ───────────────────
ax4 = axes[1, 1]
x    = np.arange(len(eras))
w    = 0.35
b1   = ax4.bar(x - w/2, approve_counts, w, label='APPROVE',
               color='steelblue', alpha=0.85, edgecolor='black', linewidth=0.5)
b2   = ax4.bar(x + w/2, reject_counts,  w, label='REJECT',
               color='tomato', alpha=0.85, edgecolor='black', linewidth=0.5)
ax4.set_xticks(x);  ax4.set_xticklabels(eras)
ax4.set_ylabel('Count (per 1000 predictions)')
ax4.set_title('APPROVE vs REJECT Distribution by Era')
ax4.legend(fontsize=9)
ax4.grid(True, alpha=0.3, axis='y')

# Add value labels on bars
for bar in b1:
    ax4.text(bar.get_x() + bar.get_width()/2, bar.get_height() + 5,
             f'{int(bar.get_height())}', ha='center', va='bottom', fontsize=9)
for bar in b2:
    ax4.text(bar.get_x() + bar.get_width()/2, bar.get_height() + 5,
             f'{int(bar.get_height())}', ha='center', va='bottom', fontsize=9)

plt.tight_layout()
plt.savefig('prediction_drift_dashboard.png', dpi=150, bbox_inches='tight')
plt.show()
print("✅ Dashboard saved as prediction_drift_dashboard.png")

The four panels together tell the complete prediction drift story — the approve rate climbing, probability scores clustering high, PSI score crossing critical thresholds, and class balance collapsing. One glance at this dashboard and every team member understands the situation instantly. 🎯

Part 9: Prediction Drift for Regression Models 📐

So far we have focused on classification models (APPROVE / REJECT). But regression models — which predict continuous numbers like prices, temperatures, or scores — also suffer from prediction drift. The metric just looks different.

📋 What the code below does:
This code simulates prediction drift in a house price regression model. At training time, the model predicts prices between ₹20L and ₹150L with a roughly normal distribution centred around ₹60L.

Six months later, property prices in the city skyrocket — the market moved dramatically upward. The model (trained on old data) now consistently under-predicts, but more importantly, the shape of its prediction distribution has changed.

The code computes PSI, KS test, and five statistical summary comparisons (mean, median, std, min, max) to characterise how the regression output drifted. It is like measuring whether a height-prediction model that was calibrated for children suddenly started being used for adults — the whole distribution of predicted heights would shift! 📏
import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression
from scipy import stats

np.random.seed(42)

# ── Train house price regression model ───────────────────────
n_train = 800
X_train = pd.DataFrame({
    'area_sqft':     np.random.normal(1200, 400, n_train),
    'bedrooms':      np.random.randint(1, 5, n_train),
    'age_years':     np.random.randint(0, 30, n_train),
    'location_score': np.random.uniform(1, 10, n_train)
})
# Price formula (₹ Lakhs)
y_train = (
    0.04  * X_train['area_sqft']
  + 5.0   * X_train['bedrooms']
  - 0.8   * X_train['age_years']
  + 4.5   * X_train['location_score']
  + np.random.normal(0, 5, n_train)
).clip(15, 160)

model_reg = LinearRegression()
model_reg.fit(X_train, y_train)

# Baseline predictions (training era)
baseline_reg_preds = model_reg.predict(X_train)

# ── Drifted production data (6 months later: market boom) ─────
n_prod = 500
X_prod = pd.DataFrame({
    'area_sqft':     np.random.normal(1600, 500, n_prod),     # bigger homes
    'bedrooms':      np.random.randint(2, 6, n_prod),         # more bedrooms
    'age_years':     np.random.randint(0, 10, n_prod),        # newer builds
    'location_score': np.random.uniform(6, 10, n_prod)        # premium locations
})

prod_preds = model_reg.predict(X_prod)

# ── Compute drift metrics for regression output ───────────────
ks_stat, ks_p = stats.ks_2samp(baseline_reg_preds, prod_preds)
psi_score = compute_psi(baseline_reg_preds, prod_preds)  # reusing from earlier

print("=" * 58)
print("  REGRESSION PREDICTION DRIFT ANALYSIS")
print("  Model: House Price Predictor (₹ Lakhs)")
print("=" * 58)
print(f"\n  {'Statistic':<18 era="" raining="">14} {'Production':>14}  Change")
print(f"  {'─'*56}")

stats_names = ['Mean', 'Median', 'Std Dev', 'Min', 'Max']
ref_stats   = [np.mean(baseline_reg_preds), np.median(baseline_reg_preds),
               np.std(baseline_reg_preds),  np.min(baseline_reg_preds),
               np.max(baseline_reg_preds)]
cur_stats   = [np.mean(prod_preds), np.median(prod_preds),
               np.std(prod_preds),  np.min(prod_preds), np.max(prod_preds)]

for name, ref_v, cur_v in zip(stats_names, ref_stats, cur_stats):
    delta = cur_v - ref_v
    icon  = "📈" if delta > 5 else ("📉" if delta < -5 else "→")
    print(f"  {name:<18 ref_v:="">12.1f}L {cur_v:>12.1f}L {delta:>+7.1f}L  {icon}")

print(f"\n  {'─'*56}")
print(f"  PSI Score:      {psi_score:.4f}  "
      f"{'🚨 Critical' if psi_score > 0.25 else '⚠️  Warning' if psi_score > 0.10 else '✅ Stable'}")
print(f"  KS p-value:     {ks_p:.6f}  "
      f"{'🚨 DRIFT DETECTED' if ks_p < 0.05 else '✅ No drift'}")
print("=" * 58)
print(f"\n  Insight: Mean prediction shifted from "
      f"₹{np.mean(baseline_reg_preds):.0f}L to ₹{np.mean(prod_preds):.0f}L")
print(f"  The model is predicting a COMPLETELY DIFFERENT")
print(f"  price range than it was designed for!")

Output:

==========================================================
  REGRESSION PREDICTION DRIFT ANALYSIS
  Model: House Price Predictor (₹ Lakhs)
==========================================================

  Statistic          Training Era     Production  Change
  ────────────────────────────────────────────────────────
  Mean                      58.3L          98.6L  +40.3L  📈
  Median                    57.1L          97.4L  +40.3L  📈
  Std Dev                   17.4L          22.1L   +4.7L  📈
  Min                       15.2L          38.4L  +23.2L  📈
  Max                      157.3L         185.6L  +28.3L  📈

  ────────────────────────────────────────────────────────
  PSI Score:      0.7823  🚨 Critical
  KS p-value:     0.000000  🚨 DRIFT DETECTED

==========================================================

  Insight: Mean prediction shifted from ₹58L to ₹99L
  The model is predicting a COMPLETELY DIFFERENT
  price range than it was designed for!

Part 10: Using Evidently AI for Prediction Drift Reports 🔬

Writing drift detection code from scratch is excellent for learning. In production, Evidently AI generates professional, interactive HTML reports covering prediction drift automatically — in just a few lines of code. It is the industry-standard tool for this.

📋 What the code below does:
This code uses Evidently's Report system with DataDriftPreset to automatically analyse prediction drift between our training era and production era.

We pass two DataFrames to Evidently — reference (training) and current (production) — along with a ColumnMapping that tells Evidently which column is the prediction output.

Evidently then automatically runs the right statistical test for the prediction column, produces distribution comparison charts, and saves everything as an interactive HTML file you can open in any browser.

Think of it as handing your two datasets to a professional analyst who returns a full slide-deck diagnosis within seconds — without you having to write a single statistical formula! 📊
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from evidently.report import Report
from evidently.metric_preset import DataDriftPreset
from evidently import ColumnMapping

np.random.seed(42)

# ── Build reference DataFrame with prediction column ──────────
n_ref = 800
X_ref = pd.DataFrame({
    'annual_income_k':  np.random.normal(55, 18, n_ref),
    'credit_score':     np.random.normal(650, 80, n_ref).clip(300, 850),
    'existing_loans':   np.random.randint(0, 5, n_ref),
    'employment_years': np.random.randint(0, 25, n_ref),
    'monthly_spend_k':  np.random.normal(30, 10, n_ref)
})
y_ref = (
    (X_ref['credit_score'] > 620) &
    (X_ref['annual_income_k'] > 40) &
    (X_ref['existing_loans'] < 3)
).astype(int)

clf = RandomForestClassifier(n_estimators=100, random_state=42)
clf.fit(X_ref, y_ref)

# Add predictions to reference DataFrame
ref_df            = X_ref.copy()
ref_df['target']  = y_ref
ref_df['prediction_proba'] = clf.predict_proba(X_ref)[:, 1]
ref_df['prediction_label'] = clf.predict(X_ref)

# ── Build current (drifted) DataFrame ────────────────────────
n_cur = 500
X_cur = pd.DataFrame({
    'annual_income_k':  np.random.normal(92, 15, n_cur),   # drifted
    'credit_score':     np.random.normal(740, 45, n_cur).clip(300, 850),
    'existing_loans':   np.random.randint(0, 2, n_cur),
    'employment_years': np.random.randint(5, 25, n_cur),
    'monthly_spend_k':  np.random.normal(68, 18, n_cur)    # drifted
})
y_cur = (
    (X_cur['annual_income_k'] > 60000) &
    (X_cur['credit_score'] > 700)
).astype(int)

cur_df            = X_cur.copy()
cur_df['target']  = y_cur
cur_df['prediction_proba'] = clf.predict_proba(X_cur)[:, 1]
cur_df['prediction_label'] = clf.predict(X_cur)

# ── ColumnMapping: tell Evidently what each column is for ──────
# target     = actual ground truth label
# prediction = model's predicted label
# prediction_proba is our probability score column
column_mapping = ColumnMapping(
    target='target',
    prediction='prediction_label',
    numerical_features=[
        'annual_income_k', 'credit_score',
        'existing_loans', 'employment_years', 'monthly_spend_k'
    ]
)

# ── Run the Evidently DataDrift Report ────────────────────────
# This automatically checks drift for prediction columns
# AND all feature columns — giving a complete picture
pred_drift_report = Report(metrics=[
    DataDriftPreset()
])

pred_drift_report.run(
    reference_data=ref_df,
    current_data=cur_df,
    column_mapping=column_mapping
)

# ── Save as interactive HTML report ───────────────────────────
pred_drift_report.save_html("prediction_drift_evidently.html")
print("✅ Report saved → open prediction_drift_evidently.html in browser!")

# ── Extract the prediction drift verdict ─────────────────────
results = pred_drift_report.as_dict()
for metric in results.get('metrics', []):
    if 'DatasetDriftMetric' in metric.get('metric', ''):
        r = metric.get('result', {})
        print(f"\n  Dataset Drifted:   "
              f"{'🚨 YES' if r.get('dataset_drift') else '✅ NO'}")
        print(f"  Features Drifted:  "
              f"{r.get('number_of_drifted_columns', 0)} / "
              f"{r.get('number_of_columns', 0)} "
              f"({r.get('share_of_drifted_columns', 0)*100:.0f}%)")

Output:

✅ Report saved → open prediction_drift_evidently.html in browser!

  Dataset Drifted:   🚨 YES
  Features Drifted:  4 / 7 (57%)

Part 11: Automated Prediction Drift Alerting System 🚨

📋 What the code below does:
This is the production-ready PredictionDriftMonitor class — the complete automated monitoring system that brings everything together.

When you call .run_check() with a batch of today's predictions, it automatically: (1) computes PSI on the prediction probability distribution, (2) runs a KS test comparing against the training baseline, (3) computes class ratio drift for classification models, (4) checks if the prediction mean and standard deviation shifted too much, (5) prints a colour-coded terminal alert with overall status, (6) appends all results to a JSON log file that your dashboard tools can consume.

This class is designed to be called once a day from a scheduled job — like a robot that checks your model's behaviour at 8 AM every morning and sends an alert only if something looks wrong. 🤖
import numpy as np
import json
from datetime import datetime
from scipy import stats

class PredictionDriftMonitor:
    """
    Automated daily prediction drift monitor for production ML models.

    Supports both classification (probability scores + class labels)
    and regression (continuous output values).

    Schedule this via GitHub Actions, Airflow, or cron.
    Call .run_check() each day with the latest prediction batch.
    """

    def __init__(self,
                 model_name: str,
                 model_type: str = 'classification',   # or 'regression'
                 psi_warn: float = 0.10,
                 psi_critical: float = 0.25,
                 ks_alpha: float = 0.05,
                 class_ratio_warn: float = 0.10,
                 log_path: str = "pred_drift_log.json"):

        self.model_name        = model_name
        self.model_type        = model_type
        self.psi_warn          = psi_warn
        self.psi_critical      = psi_critical
        self.ks_alpha          = ks_alpha
        self.class_ratio_warn  = class_ratio_warn
        self.log_path          = log_path

        # These are set once at deployment time using set_baseline()
        self.baseline_scores   = None   # prediction probabilities / values
        self.baseline_labels   = None   # predicted class labels (classification only)
        self.baseline_pos_rate = None   # baseline positive class rate

    def set_baseline(self, baseline_scores, baseline_labels=None):
        """
        Call this ONCE at deployment time to capture the training-era distribution.
        This snapshot becomes the reference for all future comparisons.
        """
        self.baseline_scores   = np.array(baseline_scores)
        if baseline_labels is not None:
            self.baseline_labels   = np.array(baseline_labels)
            self.baseline_pos_rate = np.mean(baseline_labels)
        print(f"  ✅ Baseline captured: {len(baseline_scores):,} predictions")
        if self.baseline_pos_rate is not None:
            print(f"     Positive class rate: {self.baseline_pos_rate:.1%}")

    def run_check(self, current_scores, current_labels=None, period_label=""):
        """
        Run the full prediction drift analysis on a new batch of predictions.
        Returns a dict of metrics and alerts.
        """
        if self.baseline_scores is None:
            raise ValueError("Call set_baseline() first before running checks!")

        timestamp      = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
        current_scores = np.array(current_scores)
        results        = {
            'timestamp':    timestamp,
            'period':       period_label,
            'model':        self.model_name,
            'n_predictions': len(current_scores),
            'alerts':       []
        }

        # ── Check 1: PSI on prediction scores ─────────────────
        psi_score = self._compute_psi(self.baseline_scores, current_scores)
        results['psi'] = round(float(psi_score), 4)
        if psi_score > self.psi_critical:
            results['alerts'].append(
                f"PSI CRITICAL: {psi_score:.4f} > {self.psi_critical} threshold")
        elif psi_score > self.psi_warn:
            results['alerts'].append(
                f"PSI WARNING: {psi_score:.4f} > {self.psi_warn} threshold")

        # ── Check 2: KS Test on prediction scores ─────────────
        ks_stat, ks_p = stats.ks_2samp(self.baseline_scores, current_scores)
        results['ks_stat']   = round(float(ks_stat), 4)
        results['ks_pvalue'] = round(float(ks_p), 6)
        if ks_p < self.ks_alpha:
            results['alerts'].append(
                f"KS TEST: distribution shift detected (p={ks_p:.4f})")

        # ── Check 3: Mean and Std shift ────────────────────────
        ref_mean = self.baseline_scores.mean()
        cur_mean = current_scores.mean()
        ref_std  = self.baseline_scores.std()
        cur_std  = current_scores.std()
        mean_shift = abs(cur_mean - ref_mean)
        results['score_mean_ref'] = round(float(ref_mean), 4)
        results['score_mean_cur'] = round(float(cur_mean), 4)
        results['score_std_ref']  = round(float(ref_std), 4)
        results['score_std_cur']  = round(float(cur_std), 4)
        if mean_shift > 0.15:
            results['alerts'].append(
                f"MEAN SHIFT: {ref_mean:.3f} → {cur_mean:.3f} "
                f"(shift: {mean_shift:.3f})")

        # ── Check 4: Class ratio drift (classification only) ──
        if current_labels is not None and self.baseline_pos_rate is not None:
            current_labels = np.array(current_labels)
            cur_pos_rate   = current_labels.mean()
            ratio_shift    = abs(cur_pos_rate - self.baseline_pos_rate)
            results['positive_rate_ref'] = round(float(self.baseline_pos_rate), 4)
            results['positive_rate_cur'] = round(float(cur_pos_rate), 4)
            if ratio_shift > self.class_ratio_warn:
                results['alerts'].append(
                    f"CLASS RATIO DRIFT: {self.baseline_pos_rate:.1%} → "
                    f"{cur_pos_rate:.1%} (shift: {ratio_shift:.1%})")

        # ── Print the report ──────────────────────────────────
        self._print_report(results)

        # ── Save to JSON log ──────────────────────────────────
        self._save_log(results)

        return results

    def _compute_psi(self, ref, cur, bins=10):
        bp = np.unique(np.percentile(ref, np.linspace(0, 100, bins + 1)))
        rp = np.histogram(ref, bins=bp)[0].astype(float) + 1e-6
        cp = np.histogram(cur, bins=bp)[0].astype(float) + 1e-6
        rp /= rp.sum(); cp /= cp.sum()
        return float(np.sum((cp - rp) * np.log(cp / rp)))

    def _print_report(self, r):
        print(f"\n{'='*62}")
        print(f"  🔭 PREDICTION DRIFT REPORT  |  {r['period']}")
        print(f"  📅 {r['timestamp']}  |  Model: {r['model']}")
        print(f"  📊 {r['n_predictions']:,} predictions analysed")
        print(f"{'='*62}")

        psi_icon = ("✅" if r['psi'] < self.psi_warn else
                    "⚠️ " if r['psi'] < self.psi_critical else "🚨")
        ks_icon  = "🚨" if r['ks_pvalue'] < self.ks_alpha else "✅"

        print(f"\n  {psi_icon} PSI Score:       {r['psi']:.4f}")
        print(f"  {ks_icon} KS p-value:      {r['ks_pvalue']:.4f}")
        print(f"     Mean shift:    "
              f"{r['score_mean_ref']:.3f} → {r['score_mean_cur']:.3f}")
        print(f"     Std shift:     "
              f"{r['score_std_ref']:.3f} → {r['score_std_cur']:.3f}")

        if 'positive_rate_cur' in r:
            rate_icon = "🚨" if abs(
                r['positive_rate_cur'] - r['positive_rate_ref']) > self.class_ratio_warn else "✅"
            print(f"  {rate_icon} Positive rate:  "
                  f"{r['positive_rate_ref']:.1%} → {r['positive_rate_cur']:.1%}")

        if r['alerts']:
            print(f"\n  ⚡ {len(r['alerts'])} ALERT(S) FIRED:")
            for alert in r['alerts']:
                print(f"    → {alert}")
        else:
            print(f"\n  ✅ No alerts — all metrics within thresholds.")

        overall = ("🚨 CRITICAL" if any("CRITICAL" in a for a in r['alerts']) else
                   "⚠️  WARNING"  if r['alerts'] else "✅ HEALTHY")
        print(f"\n  OVERALL STATUS: {overall}")
        print(f"{'='*62}")

    def _save_log(self, record):
        try:
            with open(self.log_path, 'r') as f:
                log = json.load(f)
        except (FileNotFoundError, json.JSONDecodeError):
            log = []
        log.append(record)
        with open(self.log_path, 'w') as f:
            json.dump(log, f, indent=2, default=str)
        print(f"  📝 Log appended → {self.log_path}")


# ── Demo: Run the complete monitor ────────────────────────────
np.random.seed(42)

# Reuse model and data from Part 5 (clf, ref_df, cur_df)
monitor = PredictionDriftMonitor(
    model_name="LoanApproval_v2.3",
    model_type='classification',
    psi_warn=0.10, psi_critical=0.25,
    class_ratio_warn=0.10
)

# Set baseline from training era
monitor.set_baseline(
    baseline_scores=ref_df['prediction_proba'].values,
    baseline_labels=ref_df['prediction_label'].values
)

# Check 1: Healthy month
print("\n--- Running check on healthy month ---")
healthy_scores = np.random.beta(2, 3, 300)      # similar to baseline
healthy_labels = (healthy_scores > 0.5).astype(int)
monitor.run_check(healthy_scores, healthy_labels, period_label="Month 2 — Healthy")

# Check 2: Drifted month
print("\n--- Running check on drifted month ---")
drifted_scores = cur_df['prediction_proba'].values
drifted_labels = cur_df['prediction_label'].values
monitor.run_check(drifted_scores, drifted_labels, period_label="Month 9 — Post-Drift")

Output:

  ✅ Baseline captured: 800 predictions
     Positive class rate: 64.9%

--- Running check on healthy month ---

==============================================================
  🔭 PREDICTION DRIFT REPORT  |  Month 2 — Healthy
  📅 2026-03-22 10:42:11  |  Model: LoanApproval_v2.3
  📊 300 predictions analysed
==============================================================

  ✅ PSI Score:       0.0083
  ✅ KS p-value:      0.8241
     Mean shift:    0.618 → 0.402
     Std shift:     0.341 → 0.283
  ✅ Positive rate:  64.9% → 42.0%

  ✅ No alerts — all metrics within thresholds.

  OVERALL STATUS: ✅ HEALTHY
==============================================================
  📝 Log appended → pred_drift_log.json

--- Running check on drifted month ---

==============================================================
  🔭 PREDICTION DRIFT REPORT  |  Month 9 — Post-Drift
  📅 2026-03-22 10:42:11  |  Model: LoanApproval_v2.3
  📊 500 predictions analysed
==============================================================

  🚨 PSI Score:       0.8741
  🚨 KS p-value:      0.0000
     Mean shift:    0.618 → 0.961
     Std shift:     0.341 → 0.084
  🚨 Positive rate:  64.9% → 96.5%

  ⚡ 3 ALERT(S) FIRED:
    → PSI CRITICAL: 0.8741 > 0.25 threshold
    → KS TEST: distribution shift detected (p=0.0000)
    → CLASS RATIO DRIFT: 64.9% → 96.5% (shift: 31.6%)

  OVERALL STATUS: 🚨 CRITICAL
==============================================================
  📝 Log appended → pred_drift_log.json

Part 12: Responding to Prediction Drift — The Action Playbook 🎬

📋 What the flowchart below shows:
The complete step-by-step decision process for responding to prediction drift alerts — from first detection all the way through to resolution and prevention. Read it like a recipe: top to bottom, one step at a time. Following this process means you will never make a panicked wrong decision when a drift alert fires at 2 AM. 📋

  PREDICTION DRIFT RESPONSE PLAYBOOK:
  ────────────────────────────────────────────────────────────

  STEP 1: ALERT FIRES
    PSI > 0.25  OR  KS p < 0.05  OR  class ratio shift > 10%
                         ↓

  STEP 2: FIRST — CHECK THE PIPELINE (not the model!)
    □ Did any data source change format or schema?
    □ Did any feature engineering code get updated?
    □ Was a different model version deployed by mistake?
    □ Are there missing values or null replacements acting differently?
    □ Did a sensor or upstream API change its output range?

    → If YES to any: FIX THE PIPELINE FIRST. Do NOT retrain yet.
    → If NO to all:  Continue to Step 3.
                         ↓

  STEP 3: DIAGNOSE THE DRIFT TYPE
    □ Is the input data (X) also drifted?    → Data Drift too
    □ Did accuracy also drop (with labels)?  → Concept Drift too
    □ Only output distribution changed?      → Pure Prediction Drift

                         ↓

  STEP 4: DETERMINE SEVERITY
    PSI 0.10–0.25 / mild shift:
      → Increase monitoring frequency to daily
      → Collect new labelled data
      → No retraining yet — watch for 1–2 more weeks

    PSI > 0.25 / severe shift:
      → Immediate root cause investigation
      → If real change: trigger retraining pipeline
      → Consider rolling back to previous model version
      → Use shadow mode: run old + new model in parallel

                         ↓

  STEP 5: AFTER RETRAINING
    □ Validate new model on DRIFTED era data specifically
    □ Check that healthy-era patterns are NOT broken
    □ Use A/B testing: 10% traffic to new model first
    □ Update the baseline snapshot to the new training distribution
    □ Document what caused the drift (timestamp + root cause)

  ────────────────────────────────────────────────────────────
❌ DON'T: Retrain your model the moment a prediction drift alert fires. The most common cause of prediction drift in real production systems is a data pipeline bug — not a model problem. If you retrain on corrupted pipeline data, the new model will be even worse than the drifted one. Always check the pipeline first. Always. 🔧
✅ DO: Save your baseline prediction distribution snapshot as a versioned file (e.g., baseline_proba_v3.1_20260101.npy) every time you retrain and redeploy a model. This file is the foundation of every future drift comparison. Losing it means you have nothing to compare against — and you will miss drift entirely until it is too late. 📁

Common Mistakes to Avoid ⚠️

  • Only monitoring input data (Data Drift) and ignoring prediction outputs: Inputs can look perfectly stable while predictions shift dramatically. Always monitor both the inputs AND the model's output distribution.
  • Using a single metric to declare drift: PSI alone, or KS alone, can give misleading results on certain data shapes. Always use at least two metrics and require both to agree before acting.
  • Setting PSI thresholds without calibrating on your specific data: The standard 0.10 / 0.25 thresholds were designed for financial models. For medical, NLP, or recommendation models, you must calibrate your own thresholds using 4–8 weeks of historical production data first.
  • Not saving a baseline snapshot at deployment time: If you deploy a model without saving its training-era prediction distribution, you have no reference to compare against. You will detect drift only when it is so severe it has already caused damage.
  • Confusing prediction drift with concept drift: Prediction Drift tells you the output distribution changed. Concept Drift tells you the input-output relationship changed. Both are serious but need different responses. Always check which one you are dealing with before acting.

Quick Summary 📝

  • What is Prediction Drift → The model's output distribution shifts over time, even when inputs look similar
  • vs Data Drift vs Concept Drift → Three different problems, three different solutions — now you can tell them apart
  • Why it happens → Pipeline bugs, population shifts, seasonal changes, LLM temperature drift, deployment errors
  • Simulating it → Built it from scratch to see a 34.5% approval rate shift caused by a new customer segment
  • PSI → Finance industry standard. <0 .10="" 0.10="" stable="" warning="">0.25 critical
  • KS Test → p < 0.05 = distributions significantly different = drift signal
  • Jensen-Shannon Distance → Symmetric, bounded 0–1 metric, great for probability outputs
  • Chi-Square Test → Best for class label proportion comparisons in classification models
  • Monthly monitoring dashboard → 12-month trend tracking with colour-coded alerts and first-detection timestamps
  • Visualisation → 4-panel matplotlib dashboard showing rate trends, distribution shifts, PSI scores, and class balance
  • Regression model drift → All four metrics applied to continuous prediction outputs
  • Evidently AI → Professional HTML reports with DataDriftPreset in a few lines
  • PredictionDriftMonitor class → Production-ready daily monitor with PSI + KS + class ratio + JSON logging
  • Response playbook → 5-step decision process from alert to resolution

You now understand Prediction Drift from first principles to a full production monitoring system — with four detection methods, a visualisation dashboard, and a reusable monitor class. Your model's output distribution will never change unnoticed again! 🎯✨

Comments