DataDriftPreset, DataQualityPreset & RegressionPreset in Evidently: A Practical Guide
Imagine you just deployed a brilliant AI model that predicts house prices. It worked perfectly last month. But today, your manager calls: "The predictions look completely off!" How do you quickly diagnose what went wrong?
The answer in 2026 is Evidently AI — an open-source Python library that gives you instant, beautiful health reports about your model and data. Three of its most powerful tools are DataDriftPreset, DataQualityPreset, and RegressionPreset.
💡 Think of them like a doctor's checkup for your AI model: DataQualityPreset checks if your data is healthy, DataDriftPreset checks if your data changed, and RegressionPreset checks if your model's predictions are accurate. Together they give you a complete picture — fast! 🩺
Part 1: What is Evidently AI? (The Big Picture) 🌍
Evidently AI is an open-source MLOps library built specifically for monitoring machine learning models in production. You give it two datasets — training data and live production data — and it automatically runs dozens of statistical checks and produces a stunning interactive HTML report.
It is used by data science teams at companies worldwide because it saves hours of manual checking every week. In 2026, it has become a standard part of the MLOps toolkit alongside MLflow and FastAPI.
A quick visual map of the three Presets — what each one is responsible for, what it checks, and how they relate to each other. Read this first before touching any code — it gives you the full picture at a glance!
The Three Presets — What Each One Does:
┌─────────────────────────────────────────────────────────────┐
│ │
│ DataQualityPreset → Is my DATA healthy and clean? │
│ ↓ │
│ Checks: missing values, duplicates, outliers, │
│ column types, value ranges, constant columns │
│ │
│ DataDriftPreset → Has my DATA changed over time? │
│ ↓ │
│ Checks: distribution shifts per feature, │
│ statistical tests (KS, chi-square, PSI), │
│ overall drift score across all columns │
│ │
│ RegressionPreset → Are my MODEL PREDICTIONS good? │
│ ↓ │
│ Checks: MAE, RMSE, MAPE, ME (bias), R-squared, │
│ error distributions, over/under-prediction patterns │
│ │
└─────────────────────────────────────────────────────────────┘
💡 Think of it like this: Before a flight, engineers run three checks — fuel quality (DataQualityPreset), weather changes (DataDriftPreset), and engine performance (RegressionPreset). Skip any one and you are flying blind! ✈️
Part 2: Installation and Setup 🛠️
First, let's set up everything you need. Create a fresh virtual environment and install the libraries.
These terminal commands create a clean, isolated Python workspace (called a virtual environment) just for this project — like a fresh empty room with only the tools you need. Then they install Evidently and all supporting libraries into that room. Think of it like unboxing and setting up a new toolbox before starting any work! 🧰
python -m venv evidently_env
source evidently_env/bin/activate # Mac / Linux
evidently_env\Scripts\activate # Windows
pip install evidently pandas numpy scikit-learn
This one-liner imports Evidently and prints its version number. We do this to confirm that the installation worked correctly and that we have version 0.4.x or higher — which is required for the Preset features. It is the same as turning on a new appliance and checking the display lights up! 💡
import evidently
print(evidently.__version__) # Should be 0.4.x or higher
requirements.txt like this:
evidently==0.4.33.
Evidently updates frequently and the API can change between minor versions.
Pin it to avoid surprises when your Docker container rebuilds! 📌
Part 3: Building Our Dataset — The Setup for All Examples 📦
Before we can use any Preset, we need two datasets. A reference dataset (what the model was trained on) and a current dataset (what the model sees in production today). Think of reference as "what normal looks like" and current as "what is happening right now." All three Presets use this same pair throughout the blog.
Step 1: Creates a reference dataset of 1,000 apartments with features like size, number of residents, and outside temperature. This is our "training time" data — what the model originally learned from.
Step 2: Calculates each apartment's real energy consumption using a formula, then adds a little random noise to simulate real-world messiness.
Step 3: Creates a current dataset of 600 apartments representing 6 months later in production. Notice the outside temperature is now higher (summer!) and more apartments have AC — these are intentional drifts we added.
Step 4: Injects real data quality problems: 30 missing temperatures, 5 impossible negative apartment sizes, 3 zero-resident rows, and 8 duplicate records — exactly what breaks real pipelines!
Step 5: Trains a Random Forest model on the reference data and adds its predictions to both datasets. Evidently needs predictions to run RegressionPreset later.
Think of this whole block as building the "test patient" and "real patient" we hand over to our three medical checkup tools. 🏥
import pandas as pd
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import train_test_split
np.random.seed(42)
# ── Build the REFERENCE dataset ───────────────────────────────
# Training-time data — what the model originally learned from.
# Scenario: predicting daily energy consumption (kWh) for apartments.
n_ref = 1000
reference = pd.DataFrame({
'apartment_size_sqft': np.random.normal(850, 200, n_ref),
'num_residents': np.random.randint(1, 6, n_ref),
'outside_temp_c': np.random.normal(22, 6, n_ref),
'has_ac': np.random.choice([0, 1], n_ref, p=[0.4, 0.6]),
'appliance_age_years': np.random.normal(4, 2, n_ref).clip(0, 20),
'day_of_week': np.random.choice(
['Mon','Tue','Wed','Thu','Fri','Sat','Sun'], n_ref)
})
# Real energy consumption formula + realistic noise
reference['energy_kwh'] = (
0.05 * reference['apartment_size_sqft']
+ 3.0 * reference['num_residents']
- 0.8 * reference['outside_temp_c']
+ 12.0 * reference['has_ac']
+ 0.5 * reference['appliance_age_years']
+ np.random.normal(0, 4, n_ref)
).clip(5, 150)
# ── Build the CURRENT dataset (production — 6 months later) ───
# Summer hits: higher outside temps, more apartments using AC
n_cur = 600
current = pd.DataFrame({
'apartment_size_sqft': np.random.normal(870, 210, n_cur),
'num_residents': np.random.randint(1, 6, n_cur),
'outside_temp_c': np.random.normal(34, 5, n_cur), # ← DRIFTED (summer!)
'has_ac': np.random.choice([0, 1], n_cur, p=[0.1, 0.9]), # ← DRIFTED!
'appliance_age_years': np.random.normal(4.5, 2, n_cur).clip(0, 20),
'day_of_week': np.random.choice(
['Mon','Tue','Wed','Thu','Fri','Sat','Sun'], n_cur)
})
# ── Inject data quality problems (happens in real pipelines!) ─
current.loc[0:29, 'outside_temp_c'] = np.nan # 30 missing values
current.loc[30:34, 'apartment_size_sqft'] = -999 # 5 invalid negatives
current.loc[35:37, 'num_residents'] = 0 # 3 impossible zeros
current = pd.concat([current, current.iloc[100:108]], ignore_index=True) # 8 duplicates
# ── Train a model and generate predictions ────────────────────
feature_cols = ['apartment_size_sqft', 'num_residents',
'outside_temp_c', 'has_ac', 'appliance_age_years']
ref_clean = reference.copy()
model = RandomForestRegressor(n_estimators=100, max_depth=6, random_state=42)
model.fit(ref_clean[feature_cols], ref_clean['energy_kwh'])
reference['prediction'] = model.predict(ref_clean[feature_cols])
cur_filled = current[feature_cols].fillna(current[feature_cols].median())
current['energy_kwh'] = (
0.05 * current['apartment_size_sqft'].clip(0, 5000)
+ 3.0 * current['num_residents'].clip(1, 10)
- 0.8 * current['outside_temp_c'].fillna(34)
+ 12.0 * current['has_ac']
+ 0.5 * current['appliance_age_years']
+ np.random.normal(0, 6, len(current))
).clip(5, 200)
current['prediction'] = model.predict(cur_filled)
print(f"Reference dataset: {reference.shape[0]:,} rows × {reference.shape[1]} cols")
print(f"Current dataset: {current.shape[0]:,} rows × {current.shape[1]} cols")
print(f"\nReference energy_kwh: mean={reference['energy_kwh'].mean():.1f} kWh")
print(f"Current energy_kwh: mean={current['energy_kwh'].mean():.1f} kWh")
print("\n✅ Both datasets ready! Let's run the Presets.")
Output:
Reference dataset: 1,000 rows × 8 cols
Current dataset: 608 rows × 8 cols
Reference energy_kwh: mean=51.3 kWh
Current energy_kwh: mean=68.7 kWh
✅ Both datasets ready! Let's run the Presets.
Notice the mean energy jumped from 51.3 to 68.7 kWh — a strong hint that something shifted. Now let's use Evidently to diagnose exactly what and where! 🔍
Part 4: DataQualityPreset — Is My Data Healthy? 🏥
What Does DataQualityPreset Check?
Imagine your teacher asked you to hand in a homework sheet with 20 questions. Before grading, the teacher runs a quick check: "Are all answers filled in? Are any answers clearly impossible (like writing '−999' for your age)? Did you copy-paste the same answer twice?"
DataQualityPreset does exactly that for your dataset. It scans every column and checks for problems before you run any other analysis. Always run this one first — garbage data produces garbage reports!
A complete list of everything DataQualityPreset checks automatically — for every column, for numeric columns specifically, for categorical columns specifically, and for the whole dataset at once. This is your "menu" of health checks. You don't have to code any of these yourself — Evidently runs them all in one line! 🍽️
DataQualityPreset — What It Checks Automatically:
┌────────────────────────────────────────────────────────────┐
│ For EVERY column: │
│ ✓ Count and % of missing values (NaN / null) │
│ ✓ Number of unique values │
│ ✓ Most common values │
│ ✓ Min, Max, Mean, Median, Std Dev │
│ ✓ Constant columns (same value everywhere — useless!) │
│ │
│ For NUMERIC columns extra: │
│ ✓ Distribution histogram shape │
│ ✓ Outlier detection │
│ ✓ Unexpected negative values │
│ │
│ For CATEGORICAL columns extra: │
│ ✓ New categories appearing in current vs reference │
│ ✓ Categories that disappeared │
│ │
│ For the whole DATASET: │
│ ✓ Duplicate rows │
│ ✓ Row count comparison (reference vs current) │
└────────────────────────────────────────────────────────────┘
Running DataQualityPreset
This code creates an Evidently
Report object and tells it:
"Run all data quality checks using DataQualityPreset()."
Then it calls .run() — passing in our two datasets —
and Evidently automatically checks every column for issues.
Finally, it saves an interactive HTML file you can open in any browser.
Think of it like handing a blood sample to a lab — you just submit it and collect the full report back. 🧪 You did not have to code any individual test yourself!
from evidently.report import Report
from evidently.metric_preset import DataQualityPreset
# Create the Report and tell it to use DataQualityPreset
quality_report = Report(metrics=[
DataQualityPreset()
])
# Run it — Evidently checks every column in both datasets automatically
quality_report.run(
reference_data=reference,
current_data=current
)
# Save a beautiful interactive HTML report you can open in your browser
quality_report.save_html("data_quality_report.html")
print("✅ Report saved → open data_quality_report.html in your browser!")
# Also get results as a Python dictionary (useful for automated alerts)
quality_results = quality_report.as_dict()
print(f"\n📊 Total checks run: {len(quality_results.get('metrics', []))}")
Output:
✅ Report saved → open data_quality_report.html in your browser!
📊 Total checks run: 14
Extracting Quality Results as Code
The HTML report is beautiful for humans to read. But in a real MLOps pipeline, your code needs to read the results too — for example, to automatically send a Slack alert if missing values exceed 5%.
This code pulls out the key quality numbers from the report dictionary and prints a clean summary with colour-coded severity labels. It also does three custom checks we know are problems in our dataset: negative apartment sizes, zero residents, and missing temperatures.
Think of this like a nurse extracting the 5 most important numbers from a 30-page lab report and putting them in a one-page patient summary. 📋
def extract_quality_summary(report_dict: dict) -> dict:
"""
Pulls the most important data quality numbers out of the
Evidently report dictionary into a simple, readable format.
"""
summary = {
'missing_values': {},
'duplicates': 0
}
for metric in report_dict.get('metrics', []):
metric_name = metric.get('metric', '')
result = metric.get('result', {})
# Missing values per column
if 'MissingValues' in metric_name:
current_missing = result.get('current', {})
for col, count in current_missing.get(
'number_of_missing_values', {}).items():
if count > 0:
pct = current_missing.get(
'share_of_missing_values', {}).get(col, 0) * 100
summary['missing_values'][col] = {
'count': count, 'pct': round(pct, 2)
}
# Duplicate row count
if 'DuplicatedRows' in metric_name:
summary['duplicates'] = result.get('current', {}).get(
'number_of_duplicated_rows', 0)
return summary
quality_dict = quality_report.as_dict()
summary = extract_quality_summary(quality_dict)
print("=" * 52)
print(" DATA QUALITY SUMMARY")
print("=" * 52)
print(f"\n🔴 Missing Values in Current Data:")
if summary['missing_values']:
for col, info in summary['missing_values'].items():
severity = "🚨 CRITICAL" if info['pct'] > 5 else "⚠️ WARNING"
print(f" {col:<28 count="" info="">4} rows "
f"({info['pct']:.1f}%) {severity}")
else:
print(" ✅ No missing values found!")
dup = summary['duplicates']
print(f"\n🔴 Duplicate Rows: {dup}",
" 🚨 Remove before training!" if dup > 0 else " ✅ None")
print("\n🔍 Custom Sanity Checks on Current Data:")
print(f" Negative apartment sizes: "
f"{(current['apartment_size_sqft'] < 0).sum()} rows "
f"{'🚨 Invalid!' if (current['apartment_size_sqft'] < 0).sum() > 0 else '✅ OK'}")
print(f" Zero-resident apartments: "
f"{(current['num_residents'] == 0).sum()} rows "
f"{'🚨 Invalid!' if (current['num_residents'] == 0).sum() > 0 else '✅ OK'}")
print(f" Missing temperatures: "
f"{current['outside_temp_c'].isna().sum()} rows "
f"{'⚠️ Fix before modelling!' if current['outside_temp_c'].isna().sum() > 0 else '✅ OK'}")
print("=" * 52)
28>
Output:
====================================================
DATA QUALITY SUMMARY
====================================================
🔴 Missing Values in Current Data:
outside_temp_c 30 rows (4.9%) ⚠️ WARNING
🔴 Duplicate Rows: 8 🚨 Remove before training!
🔍 Custom Sanity Checks on Current Data:
Negative apartment sizes: 5 rows 🚨 Invalid!
Zero-resident apartments: 3 rows 🚨 Invalid!
Missing temperatures: 30 rows ⚠️ Fix before modelling!
====================================================
Part 5: DataDriftPreset — Has My Data Changed? 🌊
What Does DataDriftPreset Check?
Imagine you always buy apples from the same shop. One day you notice they look different — greener, smaller, different taste. Something changed in the supply! DataDriftPreset is that observation — but done automatically for every feature column in your dataset, with statistics to back it up.
The exact steps DataDriftPreset follows for each feature column — how it picks the right statistical test automatically, what a "drift verdict" looks like per feature, and how it decides whether the whole dataset drifted overall. Understanding this flow makes the code much easier to read!
DataDriftPreset — What It Does for EACH Feature:
┌──────────────────────────────────────────────────────────┐
│ 1. Detect column type automatically │
│ (numeric → KS test, categorical → Chi-Square) │
│ │
│ 2. Run the statistical test on reference vs current │
│ → Gets a drift score and p-value │
│ │
│ 3. Compare against threshold (default: p < 0.05) │
│ │
│ 4. Verdict: DRIFTED 🔴 or STABLE ✅ │
│ │
│ 5. Draw side-by-side distribution chart │
└──────────────────────────────────────────────────────────┘
For the WHOLE DATASET:
→ Count what % of all features drifted
→ If > 50%: DATASET IS DRIFTED (customisable threshold)
→ If ≤ 50%: DATASET IS STABLE
Running DataDriftPreset
This code runs DataDriftPreset — the drift detection check. It uses the exact same
Report() pattern as before, just with a different Preset inside.
Evidently automatically: (1) looks at each column in both datasets, (2) picks the right statistical test based on whether the column is a number or category, (3) runs the test, and (4) labels each feature as DRIFTED or STABLE.
You write three lines of code. Evidently runs hundreds of calculations for you. The result is saved as an interactive HTML report with side-by-side charts per feature. 📊
from evidently.report import Report
from evidently.metric_preset import DataDriftPreset
# Create Report with DataDriftPreset
# Evidently will auto-detect column types and choose the right statistical test
drift_report = Report(metrics=[
DataDriftPreset()
])
# Run — comparing reference (training) vs current (production)
drift_report.run(
reference_data=reference,
current_data=current
)
# Save the interactive HTML dashboard
drift_report.save_html("data_drift_report.html")
print("✅ Drift report saved → open data_drift_report.html in your browser!")
Extracting Drift Results as Code
Just like before, we need to extract the drift results as Python numbers so our pipeline can make automated decisions.
This function loops through the report dictionary and pulls out: — the overall "did the whole dataset drift?" verdict, — the percentage of features that drifted, — and a per-feature table showing which ones are stable and which ones shifted.
The final printout is like a traffic light panel — green for stable features, red for drifted ones — so you can see in seconds exactly where the problem is. 🚦
def extract_drift_summary(drift_report) -> dict:
"""
Reads the Evidently drift report and returns a clean summary:
- overall dataset drift verdict (True/False)
- share of features that drifted (0.0 to 1.0)
- per-feature drift status, score, and p-value
"""
result_dict = drift_report.as_dict()
summary = {
'dataset_drifted': False,
'drift_share': 0.0,
'features': {}
}
for metric in result_dict.get('metrics', []):
metric_name = metric.get('metric', '')
result = metric.get('result', {})
# Overall dataset verdict
if 'DatasetDriftMetric' in metric_name:
summary['dataset_drifted'] = result.get('dataset_drift', False)
summary['drift_share'] = result.get('share_of_drifted_columns', 0.0)
# Per-feature verdict
if 'ColumnDriftMetric' in metric_name:
col = result.get('column_name', 'unknown')
summary['features'][col] = {
'drifted': result.get('drift_detected', False),
'score': round(result.get('drift_score', 0), 6),
'p_value': round(result.get('p_value', 1), 6)
if result.get('p_value') is not None else None
}
return summary
drift_summary = extract_drift_summary(drift_report)
print("=" * 62)
print(" DATA DRIFT MONITORING REPORT")
print("=" * 62)
overall = "🚨 DATASET DRIFTED" if drift_summary['dataset_drifted'] else "✅ DATASET STABLE"
pct = drift_summary['drift_share'] * 100
print(f"\n Overall Status: {overall}")
print(f" Features Drifted: {pct:.0f}% of all features\n")
print(f" {'Feature':<28 rift="">8} Score")
print(f" {'─'*50}")
for feat, info in drift_summary['features'].items():
icon = "🔴 YES" if info['drifted'] else "✅ NO"
score = f"{info['score']:.4f}"
print(f" {feat:<28 icon:="">8} {score}")
print("=" * 62)
if drift_summary['dataset_drifted']:
drifted = [f for f, v in drift_summary['features'].items() if v['drifted']]
print(f"\n⚡ DRIFTED FEATURES: {drifted}")
print(" → Investigate cause (season? new users? pipeline bug?)")
print(" → If real: schedule retraining with fresh labelled data")
28>28>
Output:
==============================================================
DATA DRIFT MONITORING REPORT
==============================================================
Overall Status: 🚨 DATASET DRIFTED
Features Drifted: 40% of all features
Feature Drift? Score
──────────────────────────────────────────────────────
apartment_size_sqft ✅ NO 0.0412
num_residents ✅ NO 0.0381
outside_temp_c 🔴 YES 0.8124
has_ac 🔴 YES 0.4123
appliance_age_years ✅ NO 0.0392
day_of_week ✅ NO 0.0184
==============================================================
⚡ DRIFTED FEATURES: ['outside_temp_c', 'has_ac']
→ Investigate cause (season? new users? pipeline bug?)
→ If real: schedule retraining with fresh labelled data
Customising Drift Detection Sensitivity
The default drift settings work well for most projects. But sometimes you need to fine-tune them — for example, using a stricter alert threshold, or switching the statistical test.
This code shows how to pass custom options into
DataDriftPreset():
setting a lower drift_share_threshold (alerts earlier),
and switching the test method to Wasserstein distance for numeric features
(better for skewed, non-normal distributions like sales or income data).
Think of it like adjusting the sensitivity dial on a smoke alarm — high sensitivity catches small fires early but may give false alarms. Lower sensitivity only alerts on big fires. You choose the right setting for your context. 🔧
from evidently.report import Report
from evidently.metric_preset import DataDriftPreset
# Custom drift settings:
# - Alert if >30% of features drift (default is 50%)
# - Use Wasserstein distance for numeric columns (better for skewed data)
# - Use Jensen-Shannon distance for categorical columns (bounded 0 to 1)
custom_drift_report = Report(metrics=[
DataDriftPreset(
drift_share_threshold=0.3, # more sensitive alerting
num_stattest='wasserstein', # good for non-normal distributions
cat_stattest='jensenshannon' # symmetric, always 0 to 1
)
])
custom_drift_report.run(
reference_data=reference,
current_data=current
)
custom_drift_report.save_html("custom_drift_report.html")
print("✅ Custom drift report saved!")
- KS test (default) → Good for most continuous numeric features
- Wasserstein → Better for non-normal or skewed distributions (income, sales)
- PSI → Industry standard in banking and finance for stability reporting
- Jensen-Shannon → Great for categorical features; always bounded 0 to 1
Part 6: RegressionPreset — How Good Are My Predictions? 🎯
What Does RegressionPreset Check?
Imagine a teacher grading a student who predicted the scores of 100 classmates. You wouldn't just say "good" or "bad." You would ask: What was the average mistake? Did they always guess too high? Were the big mistakes on easy questions (unacceptable)? RegressionPreset does this deep grading for your model's predictions.
The complete list of metrics and analysis patterns RegressionPreset computes automatically. Notice it checks both accuracy metrics (how wrong on average?) and error pattern analysis (where and how does it go wrong?). Both matter in production! A model that is equally wrong on every example is very different from one that is catastrophically wrong on a few edge cases.
RegressionPreset — What It Checks Automatically:
┌────────────────────────────────────────────────────────────┐
│ Core Accuracy Metrics: │
│ ✓ ME — Mean Error (are we biased high/low?) │
│ ✓ MAE — Mean Absolute Error (average mistake size) │
│ ✓ RMSE — Root Mean Sq Error (penalises big mistakes) │
│ ✓ MAPE — Mean Abs % Error (relative % off) │
│ ✓ R² — R-Squared (variance explained) │
│ │
│ Error Pattern Analysis: │
│ ✓ Error distribution histogram (is it normal or skewed?) │
│ ✓ Predicted vs Actual scatter plot │
│ ✓ Error vs Predicted value plot │
│ ✓ Top over-predicted examples (model said too high) │
│ ✓ Top under-predicted examples (model said too low) │
│ │
│ Reference vs Current Comparison: │
│ ✓ All metrics shown for BOTH periods side by side │
│ ✓ Immediately shows if performance degraded! │
└────────────────────────────────────────────────────────────┘
Running RegressionPreset
RegressionPreset needs one extra piece of information that the other Presets don't: it needs to know which column contains the real answer (called
target)
and which column contains the model's guess (called prediction).
We tell it this using a
ColumnMapping object —
think of it like labelling two boxes:
one says "REAL ANSWER" and the other says "MODEL GUESS."
Without this labelling, Evidently wouldn't know which is which! 📦
Once we pass
column_mapping into .run(),
Evidently computes all regression metrics automatically
for both the reference period and the current period — side by side.
from evidently.report import Report
from evidently.metric_preset import RegressionPreset
from evidently import ColumnMapping
# Tell Evidently which column is the real answer and which is the model's prediction.
# This is REQUIRED for RegressionPreset — without it the report won't work correctly!
column_mapping = ColumnMapping(
target='energy_kwh', # the real ground truth values
prediction='prediction', # what the model predicted
numerical_features=[
'apartment_size_sqft', 'num_residents',
'outside_temp_c', 'has_ac', 'appliance_age_years'
]
)
# Create the Report with RegressionPreset
regression_report = Report(metrics=[
RegressionPreset()
])
# Run — must pass column_mapping so Evidently knows what to grade
regression_report.run(
reference_data=reference,
current_data=current,
column_mapping=column_mapping
)
# Save the HTML report
regression_report.save_html("regression_report.html")
print("✅ Regression report saved → open regression_report.html in your browser!")
Extracting Regression Metrics as Code
This function reads the regression report dictionary and extracts ME, MAE, RMSE, MAPE, and R² for both the reference period and the current period.
The key insight: we compare both periods side by side. This shows immediately whether the model has degraded since it was first deployed. The "Change" column is the most important part — if MAE went from 4 to 11, the model is now almost 3x less accurate!
Think of this like a student's report card showing grades from last term vs this term — the comparison tells you whether they are improving or falling behind. 📊
import numpy as np
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
def compute_regression_metrics(df, target_col, pred_col, label):
"""
Computes ME, MAE, RMSE, MAPE and R² for one dataset.
Skips rows where either actual or predicted value is missing.
"""
mask = df[target_col].notna() & df[pred_col].notna()
y_t = df.loc[mask, target_col].values
y_p = df.loc[mask, pred_col].values
me = float(np.mean(y_t - y_p))
mae = float(mean_absolute_error(y_t, y_p))
rmse = float(np.sqrt(mean_squared_error(y_t, y_p)))
mape = float(np.mean(np.abs((y_t - y_p) /
np.where(y_t != 0, y_t, 1))) * 100)
r2 = float(r2_score(y_t, y_p))
return {'label': label, 'ME': me, 'MAE': mae,
'RMSE': rmse, 'MAPE': mape, 'R2': r2}
ref_metrics = compute_regression_metrics(reference, 'energy_kwh', 'prediction', 'Reference')
cur_metrics = compute_regression_metrics(current, 'energy_kwh', 'prediction', 'Current')
print("=" * 65)
print(" REGRESSION PERFORMANCE: REFERENCE vs CURRENT")
print("=" * 65)
print(f"\n {'Metric':<8 eference="">12} {'Current':>12} {'Change':>10} Status")
print(f" {'─'*58}")
for metric in ['ME', 'MAE', 'RMSE', 'MAPE', 'R2']:
ref_v = ref_metrics[metric]
cur_v = cur_metrics[metric]
delta = cur_v - ref_v
unit = '%' if metric == 'MAPE' else ''
# For errors: bigger change = worse. For R2: smaller = worse
if metric == 'R2':
status = "📈 Better" if delta >= 0 else ("🚨 WORSE" if delta < -0.05 else "→ Stable")
else:
status = "✅ Better" if delta < 0 else ("🚨 WORSE" if abs(delta) > 2 else "→ Stable")
print(f" {metric:<8 ref_v:="">11.3f}{unit} {cur_v:>11.3f}{unit} "
f"{delta:>+9.3f} {status}")
print("=" * 65)
8>8>
Output:
=================================================================
REGRESSION PERFORMANCE: REFERENCE vs CURRENT
=================================================================
Metric Reference Current Change Status
──────────────────────────────────────────────────────────
ME +0.821 -8.340 -9.161 🚨 WORSE
MAE 4.217 11.083 +6.866 🚨 WORSE
RMSE 5.384 13.921 +8.537 🚨 WORSE
MAPE 8.341% 16.124% +7.783 🚨 WORSE
R2 0.9341 0.8127 -0.121 🚨 WORSE
=================================================================
Every single metric degraded. MAE almost tripled. MAPE crossed 15%. R² dropped 12 points. This model needs retraining urgently! 🚨
Part 7: Combining All Three Presets in One Report 🔗
Instead of running three separate reports, we combine all three Presets into a single
Report() call.
The order matters for the HTML layout:
quality first, then drift, then regression performance.
This mirrors the natural diagnostic flow:
"Is data clean?" → "Did it change?" → "Is the model still accurate?"
One Python command. One HTML file. A complete 360-degree model health report. This is exactly what professional MLOps teams schedule to run automatically every morning. ☀️
from evidently.report import Report
from evidently.metric_preset import (
DataQualityPreset,
DataDriftPreset,
RegressionPreset
)
from evidently import ColumnMapping
column_mapping = ColumnMapping(
target='energy_kwh',
prediction='prediction',
numerical_features=[
'apartment_size_sqft', 'num_residents',
'outside_temp_c', 'has_ac', 'appliance_age_years'
]
)
# All three Presets in one Report — quality, drift, performance together
full_report = Report(metrics=[
DataQualityPreset(), # Section 1: Is the data healthy?
DataDriftPreset(), # Section 2: Has the data shifted?
RegressionPreset() # Section 3: Are predictions still accurate?
])
full_report.run(
reference_data=reference,
current_data=current,
column_mapping=column_mapping
)
full_report.save_html("full_model_health_report.html")
print("✅ Combined report saved → open full_model_health_report.html")
print("\nThis single file has THREE sections:")
print(" 📋 Section 1 — DataQualityPreset: missing, duplicates, outliers")
print(" 🌊 Section 2 — DataDriftPreset: distribution shifts per feature")
print(" 🎯 Section 3 — RegressionPreset: ME, MAE, RMSE, MAPE, R²")
One command. One file. Complete model health visibility. This is how senior MLOps engineers set up their daily monitoring pipelines. 🎉
Part 8: Automated Daily Monitoring Pipeline ⏰
Running reports manually every day is not scalable. This code builds a reusable
MLOpsHealthMonitor class that:
1️⃣ Accepts your model name, target column, and alert thresholds as configuration.
2️⃣ When you call
.run(), it runs all three Presets together.3️⃣ Saves the HTML report with a timestamp in the filename (so you keep history).
4️⃣ Extracts key numbers using
scipy and sklearn for quick checks.5️⃣ Prints a colour-coded terminal summary with ✅ / 🚨 indicators per area.
6️⃣ Appends all results to a JSON log file — your dashboards can read this automatically.
Think of this class as your AI model's personal doctor — it runs the same full checkup every single day without you having to ask. 🤖
import os
import json
from datetime import datetime
from evidently.report import Report
from evidently.metric_preset import (
DataQualityPreset, DataDriftPreset, RegressionPreset
)
from evidently import ColumnMapping
import numpy as np
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
from scipy import stats
class MLOpsHealthMonitor:
"""
Automated daily health monitor for a regression ML model.
Run this class on a schedule (cron, Airflow, GitHub Actions) to:
1. Check data quality in the latest production batch
2. Detect feature distribution drift since training
3. Measure if regression performance has degraded
"""
def __init__(self, model_name, target_col, prediction_col,
numeric_features, output_dir="monitoring_reports",
mape_threshold=15.0, drift_share_threshold=0.4):
self.model_name = model_name
self.target_col = target_col
self.pred_col = prediction_col
self.num_features = numeric_features
self.output_dir = output_dir
self.mape_threshold = mape_threshold
self.drift_threshold = drift_share_threshold
os.makedirs(output_dir, exist_ok=True)
self.col_map = ColumnMapping(
target=target_col,
prediction=prediction_col,
numerical_features=numeric_features
)
def run(self, reference_df, current_df):
"""Run all three Presets, save report, print summary, log results."""
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
date_str = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
print(f"\n{'='*60}")
print(f" 🩺 Health Check — {self.model_name}")
print(f" 📅 {date_str}")
print(f"{'='*60}")
# Run the combined Evidently report
report = Report(metrics=[
DataQualityPreset(),
DataDriftPreset(drift_share_threshold=self.drift_threshold),
RegressionPreset()
])
report.run(reference_data=reference_df, current_data=current_df,
column_mapping=self.col_map)
# Save the HTML with a timestamp in the filename
html_path = os.path.join(self.output_dir,
f"health_{timestamp}.html")
report.save_html(html_path)
print(f" 📄 HTML report: {html_path}")
# Compute quick health numbers and print the summary
health = self._compute_health(current_df)
health.update({'timestamp': date_str, 'model': self.model_name})
self._print_summary(health)
# Append to JSON log for dashboard tools
log_path = os.path.join(self.output_dir, "health_log.json")
self._append_log(log_path, health)
print(f" 📝 Log updated: {log_path}")
return health
def _compute_health(self, cur_df):
"""Compute key health numbers from the current dataset."""
missing_pct = cur_df.isnull().mean().max() * 100
dup_count = int(cur_df.duplicated().sum())
# KS drift check per numeric feature
drifted = 0
for col in self.num_features:
if col in reference.columns and col in cur_df.columns:
_, p = stats.ks_2samp(
reference[col].dropna(), cur_df[col].dropna())
if p < 0.05:
drifted += 1
drift_share = drifted / max(len(self.num_features), 1)
# Regression metrics
mask = cur_df[self.target_col].notna() & cur_df[self.pred_col].notna()
y_t = cur_df.loc[mask, self.target_col].values
y_p = cur_df.loc[mask, self.pred_col].values
mae = float(mean_absolute_error(y_t, y_p))
rmse = float(np.sqrt(mean_squared_error(y_t, y_p)))
mape = float(np.mean(np.abs((y_t - y_p) /
np.where(y_t != 0, y_t, 1))) * 100)
r2 = float(r2_score(y_t, y_p))
me = float(np.mean(y_t - y_p))
return {
'quality': {'missing_pct': round(missing_pct, 2),
'duplicates': dup_count},
'drift': {'share': round(drift_share, 3),
'drifted': drift_share >= self.drift_threshold},
'perf': {'ME': round(me, 3), 'MAE': round(mae, 3),
'RMSE': round(rmse, 3), 'MAPE': round(mape, 3),
'R2': round(r2, 4)}
}
def _print_summary(self, health):
"""Print colour-coded terminal summary."""
q, d, p = health['quality'], health['drift'], health['perf']
alerts = []
miss_icon = "✅" if q['missing_pct'] < 5 else "🚨"
dup_icon = "✅" if q['duplicates'] == 0 else "🚨"
drft_icon = "✅" if not d['drifted'] else "🚨"
mape_icon = "✅" if p['MAPE'] < self.mape_threshold else "🚨"
r2_icon = "✅" if p['R2'] > 0.80 else "🚨"
if q['missing_pct'] >= 5: alerts.append(f"Missing values: {q['missing_pct']:.1f}%")
if q['duplicates'] > 0: alerts.append(f"Duplicate rows: {q['duplicates']}")
if d['drifted']: alerts.append(f"Drift: {d['share']*100:.0f}% features shifted")
if p['MAPE'] >= self.mape_threshold: alerts.append(f"MAPE {p['MAPE']:.1f}% > {self.mape_threshold}%")
if p['R2'] <= 0.80: alerts.append(f"R² = {p['R2']:.3f} below 0.80")
print(f"\n DATA QUALITY")
print(f" {miss_icon} Max missing: {q['missing_pct']:.1f}%")
print(f" {dup_icon} Duplicates: {q['duplicates']}")
print(f"\n DATA DRIFT")
print(f" {drft_icon} Features drifted: {d['share']*100:.0f}%")
print(f"\n REGRESSION PERFORMANCE")
print(f" ME = {p['ME']:>+8.3f}")
print(f" MAE = {p['MAE']:>8.3f}")
print(f" RMSE = {p['RMSE']:>8.3f}")
print(f" {mape_icon} MAPE = {p['MAPE']:>7.3f}%")
print(f" {r2_icon} R² = {p['R2']:>8.4f}")
overall = ("✅ HEALTHY" if not alerts else
"⚠️ WARNING" if len(alerts) == 1 else "🚨 CRITICAL")
print(f"\n OVERALL STATUS: {overall}")
if alerts:
print("\n ⚡ ALERTS:")
for a in alerts:
print(f" → {a}")
def _append_log(self, path, record):
"""Append this run's health record to the JSON log file."""
try:
with open(path, 'r') as f:
log = json.load(f)
except (FileNotFoundError, json.JSONDecodeError):
log = []
log.append(record)
with open(path, 'w') as f:
json.dump(log, f, indent=2, default=str)
# ── Run the monitor ────────────────────────────────────────────
monitor = MLOpsHealthMonitor(
model_name="EnergyConsumption_v3.1",
target_col="energy_kwh",
prediction_col="prediction",
numeric_features=['apartment_size_sqft', 'num_residents',
'outside_temp_c', 'has_ac', 'appliance_age_years'],
mape_threshold=15.0,
drift_share_threshold=0.4
)
health_result = monitor.run(reference, current)
Output:
============================================================
🩺 Health Check — EnergyConsumption_v3.1
📅 2026-03-22 10:28:41
============================================================
📄 HTML report: monitoring_reports/health_20260322_102841.html
DATA QUALITY
✅ Max missing: 4.9%
🚨 Duplicates: 8
DATA DRIFT
🚨 Features drifted: 40%
REGRESSION PERFORMANCE
ME = -8.340
MAE = 11.083
RMSE = 13.921
🚨 MAPE = 16.124%
🚨 R² = 0.8127
OVERALL STATUS: 🚨 CRITICAL
⚡ ALERTS:
→ Duplicate rows: 8
→ Drift: 40% features shifted
→ MAPE 16.1% > 15.0%
→ R² = 0.813 below 0.80
📝 Log updated: monitoring_reports/health_log.json
Part 9: TestSuite — Automated Pass / Fail Deployment Gates 🧪
Reports (with Presets) give you beautiful charts and numbers for humans to explore. But in a CI/CD pipeline, you need a simple binary decision: PASS or FAIL.
Evidently's TestSuite does exactly this. You define specific checks with specific thresholds — like "missing values must be under 5%" or "MAPE must be below 15%." If any single check fails, the whole suite fails and your deployment pipeline is automatically blocked.
Think of it like a pre-flight safety checklist ✈️ — every box must be ticked before the plane is cleared for takeoff. If even one item fails, the flight is grounded. No human needs to check anything — it's fully automated!
from evidently.test_suite import TestSuite
from evidently.tests import (
TestShareOfMissingValues,
TestNumberOfDuplicatedRows,
TestShareOfDriftedColumns,
TestValueMAE,
TestValueMAPE,
TestValueR2Score
)
from evidently import ColumnMapping
column_mapping = ColumnMapping(
target='energy_kwh',
prediction='prediction'
)
# Build the TestSuite — each test has a specific PASS/FAIL threshold
# If ANY test fails, the suite fails and deployment should be blocked
health_tests = TestSuite(tests=[
TestShareOfMissingValues(lt=0.05), # Missing values must be < 5%
TestNumberOfDuplicatedRows(eq=0), # Zero duplicates allowed
TestShareOfDriftedColumns(lt=0.4), # Fewer than 40% features can drift
TestValueMAE(lt=10.0), # MAE must be below 10 kWh
TestValueMAPE(lt=0.15), # MAPE must be below 15%
TestValueR2Score(gt=0.80) # R² must be above 0.80
])
# Run all checks
health_tests.run(
reference_data=reference,
current_data=current,
column_mapping=column_mapping
)
# Save detailed HTML for reviewing failures
health_tests.save_html("test_suite_report.html")
# Extract and display pass/fail results
test_results = health_tests.as_dict()
all_tests = test_results.get('tests', [])
print("=" * 60)
print(" DEPLOYMENT GATE — TEST SUITE RESULTS")
print("=" * 60)
passed = 0
failed = 0
for t in all_tests:
name = t.get('name', 'Unknown')
status = t.get('status', 'ERROR')
detail = t.get('description', '')
icon = "✅ PASS" if status == 'SUCCESS' else "❌ FAIL"
if status == 'SUCCESS':
passed += 1
else:
failed += 1
print(f" {icon} {name}")
if status != 'SUCCESS':
print(f" ↳ {detail}")
print(f"\n Results: {passed} passed, {failed} failed / {len(all_tests)} total")
print("=" * 60)
if failed == 0:
print("\n 🟢 GATE: OPEN — All checks passed! Safe to deploy.")
else:
print(f"\n 🔴 GATE: BLOCKED — {failed} test(s) failed!")
print(" Fix issues before deploying to production.")
Output:
============================================================
DEPLOYMENT GATE — TEST SUITE RESULTS
============================================================
✅ PASS Share of Missing Values
❌ FAIL Number of Duplicated Rows
↳ Duplicated rows: 8. Threshold: eq=0
❌ FAIL Share of Drifted Columns
↳ Share drifted: 0.4. Threshold: lt=0.4
✅ PASS Mean Absolute Error
❌ FAIL Mean Absolute Percentage Error
↳ MAPE: 0.16. Threshold: lt=0.15
❌ FAIL R2 Score
↳ R2: 0.81. Threshold: gt=0.80
Results: 2 passed, 4 failed / 6 total
============================================================
🔴 GATE: BLOCKED — 4 test(s) failed!
Fix issues before deploying to production.
Part 10: GitHub Actions — Scheduling Daily Checks ⚙️
This is not Python — it is a GitHub Actions workflow file written in YAML format. You save it in your repository at
.github/workflows/daily_check.yml
and GitHub reads it automatically.
It tells GitHub: "Every day at 8 AM UTC, spin up a fresh Ubuntu machine, install our Python dependencies, run our monitoring script, and save the HTML report as a downloadable artifact. If anything fails, send an alert to our Slack channel."
This means your model gets checked every morning without anyone on your team having to remember to do it. It is like hiring a robot assistant who wakes up every day at 8 AM and files the health report before the team arrives at work! 🤖☕
# File: .github/workflows/daily_model_health.yml
name: Daily Model Health Check
on:
schedule:
- cron: '0 8 * * *' # Every day at 08:00 UTC
workflow_dispatch: # Also allow manual trigger from GitHub UI
jobs:
health-check:
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Set up Python 3.11
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Install dependencies
run: pip install evidently pandas numpy scikit-learn scipy
- name: Run monitoring script
run: python monitoring/run_health_check.py
- name: Upload HTML report as downloadable artifact
uses: actions/upload-artifact@v4
with:
name: health-report-${{ github.run_id }}
path: monitoring_reports/*.html
retention-days: 30
- name: Alert Slack if anything failed
if: failure()
uses: rtCamp/action-slack-notify@v2
env:
SLACK_WEBHOOK: ${{ secrets.SLACK_WEBHOOK_URL }}
SLACK_MESSAGE: '🚨 Model Health Check FAILED — see GitHub Actions!'
Part 11: The Complete Flow — How It All Fits Together 🗺️
The full end-to-end MLOps monitoring pipeline using all three Evidently Presets. Each step builds on the previous one — data must be clean before you check drift, and drift must be understood before you decide whether to retrain. This is the same workflow used by professional ML teams at real companies in 2026. Read it top to bottom like a recipe for keeping your model healthy forever! 🍳
Complete MLOps Monitoring Pipeline — Daily Flow:
─────────────────────────────────────────────────────────────────
EVERY MORNING at 08:00 AM (automated via GitHub Actions / Airflow)
─────────────────────────────────────────────────────────────────
Step 1: FETCH DATA
Load reference snapshot (saved when model was last trained)
Load yesterday's production data (predictions + actual values)
↓
Step 2: DATA QUALITY CHECK ← DataQualityPreset
Missing values? Duplicates? Invalid ranges?
→ FAIL: Alert data engineering team, pause pipeline
→ PASS: Move to drift check
↓
Step 3: DATA DRIFT CHECK ← DataDriftPreset
Have any feature distributions shifted significantly?
→ MILD DRIFT: Increase monitoring frequency, log warning
→ SEVERE DRIFT: Schedule retraining with fresh labelled data
↓
Step 4: PERFORMANCE CHECK ← RegressionPreset
Has MAE / MAPE / R² degraded vs the reference period?
→ PASS: Log metrics, all good 😴
→ FAIL: Trigger immediate retraining workflow
↓
Step 5: DEPLOYMENT GATE ← TestSuite
Did all automated pass/fail checks succeed?
→ ALL PASS: ✅ Approve next model deployment
→ ANY FAIL: 🔴 Block deployment, alert on-call engineer
↓
Step 6: LOG + REPORT
Save timestamped HTML report (for human review)
Append JSON record to health_log.json (for dashboards)
Send Slack / email summary to the team
─────────────────────────────────────────────────────────────────
This loop runs EVERY DAY. Team only gets paged if something
actually goes wrong. No manual checking ever needed! 🤖
─────────────────────────────────────────────────────────────────
Common Mistakes to Avoid ⚠️
- Skipping DataQualityPreset before drift analysis: Missing values and duplicates corrupt drift statistics. Always clean first, then check for drift.
- Using a reference dataset that is too small: Statistical drift tests need 200–500 rows minimum. Smaller reference sets produce unstable p-values that change on every run.
- Forgetting ColumnMapping for RegressionPreset: Without it, Evidently cannot identify the target and prediction columns. The report runs but returns empty or misleading regression metrics.
- Setting drift thresholds too tight: On large datasets, even harmless tiny fluctuations will trigger "drift detected." Calibrate thresholds using 2–4 weeks of historical production data first.
- Running Presets on inconsistent data windows: Always use the exact same reference and current datasets for all three Presets. Mixed windows make cross-Preset comparisons meaningless.
health_report_20260322.html) and keep a rolling 90-day archive.
When a model suddenly degrades, that archive is your detective toolkit —
you can pinpoint exactly which day drift first appeared and which feature moved first. 🕵️
Quick Summary 📝
- Evidently AI → Open-source MLOps monitoring library that auto-generates interactive HTML health reports
- DataQualityPreset → Checks every column for missing values, duplicates, outliers, invalid entries, and constant columns — always run this first
- DataDriftPreset → Auto-selects the right statistical test per column type, returns a per-feature drift verdict and an overall dataset verdict
- RegressionPreset → Computes ME, MAE, RMSE, MAPE, R² and error patterns — compares reference vs current period side by side
- ColumnMapping → Required by RegressionPreset to identify the target and prediction columns in your DataFrame
- Combined Report → All three Presets in one
Report()call — one HTML file, complete 360° visibility - MLOpsHealthMonitor class → Production-ready, reusable daily monitor with timestamped reports and JSON logging
- TestSuite → Binary pass/fail gates for CI/CD — automatically blocks bad model deployments
- GitHub Actions → Schedules daily health checks without any dedicated server
- Complete pipeline flow → Quality → Drift → Performance → Gate → Log → Repeat every day
Evidently's most powerful Presets — DataQualityPreset, DataDriftPreset, and RegressionPreset — to build a complete, automated, production-grade model monitoring system. Your AI model will never fly blind again! 🌟✈️
Comments
Post a Comment