Skip to main content

Experiment Tracking in MLOps

Calculating read time…

Imagine you are a scientist trying to bake the perfect chocolate cake 🍰.
You try 50 different recipes — different amounts of sugar, butter, baking time, temperature.
But you forget to write anything down.
Three weeks later you cannot remember which recipe made the best cake!

That is exactly what happens when data scientists train machine learning models without experiment tracking.
They run hundreds of training experiments, change hyperparameters, swap datasets — and forget what worked.




Experiment Tracking is the lab notebook for your ML models.
It records every detail of every training run automatically — so you can always go back, compare, reproduce, and improve.

What Is Experiment Tracking?

When you train a machine learning model, you make dozens of decisions:

  • What data did I use? Which version?
  • What hyperparameters did I set? Learning rate? Batch size? Number of layers?
  • How good was the final model? Accuracy? Loss? F1 score?
  • Which model file should go to production?

Experiment tracking tools record all of this automatically — every single time you run training.
No manual spreadsheets. No "I think I used 0.001 learning rate… maybe 0.01?"
Everything is stored, timestamped, searchable, and comparable.

💡 Think of it like: Google Drive for your ML experiments.
Every training run gets its own folder with all the evidence — code, data version, settings, results, and the actual model file.
You can pull up any run from 6 months ago and reproduce it exactly.

Why Experiment Tracking Matters — The Problem Without It

Without tracking, here is what a typical ML team's life looks like:

🧒 What this shows: Imagine keeping score in a game but writing nothing down. By Thursday, nobody remembers who was winning on Monday!
Monday:    Run experiment with lr=0.01     → accuracy 78%  (not saved)
Tuesday:   Run experiment with lr=0.001    → accuracy 84%  (forgot settings)
Wednesday: Run experiment with lr=0.0001   → accuracy 81%  (can't reproduce)
Thursday:  "Which experiment had 84% again?? I need to deploy THAT model!"
Friday:    😭 Nobody knows. Start over.

With experiment tracking, the same week looks like:

🧒 What this shows: Same game, but now someone wrote every score in a notebook. On Thursday, you just flip to Tuesday's page — answer found instantly!
Monday:    Run #001  lr=0.01    → 78%    ← tracked automatically ✅
Tuesday:   Run #002  lr=0.001   → 84%    ← tracked automatically ✅
Wednesday: Run #003  lr=0.0001  → 81%    ← tracked automatically ✅
Thursday:  "Run #002 had 84%. Here is the model file, the dataset hash,
            the exact code commit, and all hyperparameters."
Friday:    🚀 Deployed to production in 10 minutes.

That is the power of experiment tracking. It turns chaotic research into reproducible engineering.

The 4 Things Every Tracking Tool Records

Every experiment tracking tool — no matter which one you pick — records these four things:

  • Parameters → The settings you chose before training starts. Learning rate, batch size, number of epochs, model architecture, optimizer type.
  • Metrics → The numbers that measure how good your model is. Training loss, validation accuracy, F1 score, AUC-ROC — recorded at every epoch.
  • Artifacts → The files your training produces. The saved model file, plots, confusion matrix images, sample predictions, trained weights.
  • Metadata → Everything else about the context. Which Git commit ran, which dataset version was used, which Python version, which GPU, start time, end time, duration.

Together these four things mean you can completely reconstruct any past experiment — even years later. That is what reproducibility means in MLOps.

The Toolchain — What Everyone Is Using

The experiment tracking landscape has settled into three clear choices depending on your team size and budget:

  • MLflow → The open-source de facto standard. 20,000+ GitHub stars, 14 million monthly downloads. Free to self-host. Used by Netflix, Shopify, Databricks, Zillow. Best if you want zero vendor lock-in.
  • Weights & Biases (W&B) → The developer-favourite. Beautiful UI, easiest setup, best visualisations. Used by OpenAI, NVIDIA, Stability AI, Microsoft. Best for research teams and individuals.
  • Neptune.ai → The enterprise-grade metadata database. Best for large teams running thousands of experiments with strict governance requirements.

Other valid choices in the ecosystem:

  • ClearML → Fully open-source, includes pipeline management and experiment tracking in one tool
  • Comet ML → Strong visualisation and LLM experiment support, great for GenAI teams
  • DVC → Git-native tracking for data and models, now part of lakeFS ecosystem
  • AWS SageMaker Experiments → Best choice if your entire stack is already on AWS
  • Google Vertex AI Experiments → Best choice if your stack is on Google Cloud

For this post we go deep on MLflow — the most widely used open-source option — and also show W&B for comparison.

Part 1 — MLflow: The Open-Source Standard

MLflow has four components that work together. Think of them as four departments in the same office:

  • MLflow Tracking → The lab notebook. Records parameters, metrics, and artifacts from every run.
  • MLflow Projects → The reproducibility department. Packages your code so anyone can rerun it exactly.
  • MLflow Models → The packaging department. Wraps your model in a standard format that works with any serving framework.
  • MLflow Model Registry → The promotion department. Manages model versions and their lifecycle — Staging → Production → Archived.

Step 1: Install MLflow

🧒 What this code does: This is like downloading an app on your phone.
Line 1 installs MLflow onto your computer using pip (the Python app store).
Line 2 checks the version to make sure it installed correctly.
Line 3 opens a website on your computer at localhost:5000 where you can see all your experiments in a nice table — like opening your lab notebook in a browser.
# Install MLflow (Python 3.8+ required)
pip install mlflow

# Verify the installation
mlflow --version
# Output: mlflow, version 2.20.0

# Launch the local tracking server (opens in your browser)
mlflow ui
# Visit http://localhost:5000 to see the MLflow dashboard

Step 2: Your First Tracked Experiment

🧒 What this code does: Think of this like doing a science experiment and filling in your lab sheet at the same time — automatically!

Step by step what happens:
① We make some fake weather data (like temperature readings) to practise with.
② We press "record" by calling mlflow.start_run() — everything after this is captured.
③ We write down our settings (alpha, max_iter) — like writing "I used 2 cups of sugar" in a recipe.
④ We train the model — this is the actual learning part.
⑤ We measure how good it is (MAE — lower is better).
⑥ We save the result and the model file.
⑦ The with block closes — MLflow saves everything to your dashboard automatically. Go to localhost:5000 and see it!
import mlflow
import mlflow.sklearn
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split
import numpy as np

# --- Fake weather data for the example ---
X = np.random.randn(1000, 10)   # 1000 samples, 10 features
y = np.random.randn(1000)        # temperature values

X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2)

# --- The hyperparameters we want to test ---
alpha = 0.5          # regularisation strength
max_iter = 1000      # maximum iterations

# ✨ Everything between start_run() and end is automatically tracked ✨
with mlflow.start_run(run_name="ridge-alpha-0.5"):

    # Log the hyperparameters we chose
    mlflow.log_param("alpha", alpha)
    mlflow.log_param("max_iter", max_iter)
    mlflow.log_param("model_type", "Ridge")
    mlflow.log_param("dataset_version", "v2.3")

    # Train the model
    model = Ridge(alpha=alpha, max_iter=max_iter)
    model.fit(X_train, y_train)

    # Evaluate
    predictions = model.predict(X_val)
    mae = mean_absolute_error(y_val, predictions)

    # Log the result metrics
    mlflow.log_metric("mae", mae)
    mlflow.log_metric("train_samples", len(X_train))
    mlflow.log_metric("val_samples", len(X_val))

    # Save the trained model as an artifact
    mlflow.sklearn.log_model(model, "weather-model")

    print(f"Run complete! MAE = {mae:.4f}")
    print(f"View at: http://localhost:5000")

What happens automatically:

  • MLflow creates a unique Run ID for this experiment
  • All parameters are stored — alpha, max_iter, model_type, dataset_version
  • The MAE metric is stored and plotted over time
  • The trained model is saved and downloadable from the UI
  • The Git commit hash is captured automatically
  • Start time, end time, and duration are recorded

Open http://localhost:5000 in your browser and you will see your experiment with all these details in a clean table. 🎯

Step 3: Autologging — Track Everything with One Line

MLflow can automatically capture every parameter and metric from popular libraries — with just one line of code:

🧒 What this code does: Remember how you had to write down every ingredient yourself in the last example?
mlflow.sklearn.autolog() is like having a robot helper who follows you around the kitchen and writes everything down FOR you.
You don't write a single log_param or log_metric line — the robot does it all the moment training finishes. Magic! ✨
import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import train_test_split

# Enable autologging — captures EVERYTHING automatically
mlflow.sklearn.autolog()

X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2)

with mlflow.start_run(run_name="random-forest-autolog"):
    model = RandomForestRegressor(n_estimators=100, max_depth=5)
    model.fit(X_train, y_train)
    # That's it! MLflow captured all params, metrics, and the model.

Autologging works with sklearn, PyTorch, TensorFlow/Keras, XGBoost, LightGBM, Hugging Face Transformers, and more.
For deep learning models it logs the loss and metrics at every single epoch automatically.

Step 4: Logging Metrics at Every Epoch (Deep Learning)

For neural networks you want to track how loss and accuracy improve during training:

🧒 What this code does: Imagine you are running a 50-lap race 🏃 and someone calls out your speed after every single lap.
That is what this code does — after every training epoch (one lap of learning), it calls out the train_loss and val_mae scores and writes them into your notebook.
At the end, MLflow draws a smooth graph showing how your model got smarter and smarter over 50 laps.
You can then compare that graph against other training runs to see which one learned the fastest!
import mlflow
import torch

mlflow.pytorch.autolog()  # one line — logs everything!

with mlflow.start_run(run_name="weather-lstm-v1"):

    mlflow.log_params({
        "epochs": 50,
        "learning_rate": 0.001,
        "batch_size": 64,
        "hidden_size": 128,
        "num_layers": 2,
        "dropout": 0.2,
        "optimizer": "Adam",
        "dataset": "global_weather_2024_v3"
    })

    model = WeatherLSTM(hidden_size=128, num_layers=2, dropout=0.2)
    optimizer = torch.optim.Adam(model.parameters(), lr=0.001)

    for epoch in range(50):
        train_loss = train_one_epoch(model, optimizer)
        val_loss, val_mae = evaluate(model)

        # Log metrics at each epoch — MLflow plots these automatically
        mlflow.log_metrics({
            "train_loss": train_loss,
            "val_loss": val_loss,
            "val_mae": val_mae
        }, step=epoch)

    # Save the trained model
    mlflow.pytorch.log_model(model, "weather-lstm")

In the MLflow UI you will see smooth line charts showing how train_loss and val_loss decreased over 50 epochs — for every single run. You can overlay multiple runs to compare them visually. 📈

Step 5: Comparing Multiple Runs — Grid Search

Now let's run many experiments — trying different hyperparameters — and let MLflow track all of them at once:

🧒 What this code does: Imagine trying every combination of pizza toppings to find the tastiest pizza 🍕.
You try small/medium/large sizes with thin/thick crusts and 3 different cheeses = 27 combinations total.
Instead of tasting one pizza per day (27 days!), this code bakes all 27 automatically, measures how delicious each one is (val_mae score), and writes everything in a neat table.
At the end you just look at the table, sort by tastiest, and you know which pizza wins — in minutes instead of a month! 🏆
import mlflow
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_absolute_error

mlflow.set_experiment("weather-model-hyperparameter-search")

# Try all combinations — 27 runs total
learning_rates = [0.001, 0.01, 0.1]
n_estimators_list = [50, 100, 200]
max_depths = [3, 5, 10]

for lr in learning_rates:
    for n_est in n_estimators_list:
        for depth in max_depths:
            with mlflow.start_run(
                run_name=f"rf-lr{lr}-n{n_est}-d{depth}"
            ):
                mlflow.log_params({
                    "learning_rate": lr,
                    "n_estimators": n_est,
                    "max_depth": depth
                })

                model = RandomForestRegressor(
                    n_estimators=n_est,
                    max_depth=depth
                )
                model.fit(X_train, y_train)
                mae = mean_absolute_error(y_val, model.predict(X_val))

                mlflow.log_metric("val_mae", mae)
                mlflow.sklearn.log_model(model, "model")

Expected MLflow UI output — a sortable comparison table:

Run Name                     lr      n_est  depth   val_mae
─────────────────────────────────────────────────────────────
rf-lr0.01-n200-d10          0.01    200    10      0.1823  ← BEST ✅
rf-lr0.01-n100-d10          0.01    100    10      0.1891
rf-lr0.001-n200-d10         0.001   200    10      0.1934
rf-lr0.1-n50-d3             0.1     50     3       0.2891  ← WORST

Sort by val_mae ascending and your best model is instantly visible. Click on it to see every detail.
This is why teams run 50 experiments in one afternoon instead of one per day. 🏎️

Part 2 — The MLflow Model Registry

Finding the best experiment is step one. Promoting it to production safely is step two.
That is what the Model Registry handles.

Think of it like a promotion process at work 👔:
Junior → Senior → Manager → VP.
A model goes through: None → Staging → Production → Archived.

🧒 What this code does: You finished baking 27 pizzas and found the tastiest one (Run #002). Now you want to put it on the restaurant menu officially 🍕🍽️.

Step 1 puts the pizza on a "trying it out" menu (Staging) — only internal staff test it.
Step 2 puts it on the main menu for all customers (champion = Production).
Step 3 is how any chef in the kitchen can grab the exact same recipe later with one command.
No more emailing model files around — everyone loads from the same registered source!
import mlflow
from mlflow.tracking import MlflowClient

client = MlflowClient()

# Step 1: Register a model from a run into the registry
model_uri = "runs:/abc123def456/model"    # run ID from your best experiment
registered_model = mlflow.register_model(
    model_uri=model_uri,
    name="weather-prediction-model"
)
print(f"Registered as version {registered_model.version}")
# Output: Registered as version 1

# Step 2: Promote version 1 to Staging (internal testing)
client.set_registered_model_alias(
    name="weather-prediction-model",
    alias="staging",
    version="1"
)

# Step 3: After testing passes — promote to Production!
client.set_registered_model_alias(
    name="weather-prediction-model",
    alias="champion",       # 'champion' is the 2026 standard alias for production
    version="1"
)

# Step 4: Load the production model anywhere in your codebase
champion_model = mlflow.pyfunc.load_model(
    "models:/weather-prediction-model@champion"
)

predictions = champion_model.predict(new_weather_data)
print(predictions)

Querying the Registry for the Best Run Automatically

🧒 What this code does: Instead of a human reading the table and picking the winner, this code does it automatically.
It searches ALL your experiments, finds the one with the lowest val_mae score (lowest error = best model), and then registers it straight into the Model Registry without any human clicking.
Think of it like a robot judge at a science fair that reads every project, picks the winner, and hands them the trophy — instantly! 🤖🏆
from mlflow.tracking import MlflowClient

client = MlflowClient()

# Find the single best run across ALL experiments — by lowest val_mae
best_run = client.search_runs(
    experiment_ids=["1"],
    filter_string="metrics.val_mae < 0.20",
    order_by=["metrics.val_mae ASC"],
    max_results=1
)[0]

print(f"Best Run ID: {best_run.info.run_id}")
print(f"Best MAE:    {best_run.data.metrics['val_mae']:.4f}")
print(f"Parameters:  {best_run.data.params}")

# Automatically register the best run in one step
mlflow.register_model(
    model_uri=f"runs:/{best_run.info.run_id}/model",
    name="weather-prediction-model"
)

Part 3 — Weights & Biases: The Developer-Favourite

W&B is loved by researchers for its beautiful real-time dashboards and zero-configuration setup.
Used by OpenAI, NVIDIA, Stability AI, Microsoft, and most of the top AI research labs worldwide.

Step 1: Install and Set Up W&B

🧒 What this code does: Two quick steps — install the W&B library (like downloading an app), then log in.
Logging in gives W&B permission to save your experiments on their website at wandb.ai.
After this, every training run you do will appear on a beautiful online dashboard — like your own personal science fair display board that updates in real time while the model trains! 🎨
# Install
pip install wandb

# Log in (creates free account at wandb.ai)
wandb login
# Paste your API key from wandb.ai/authorize

Step 2: Track an Experiment with W&B

🧒 What this code does: This is very similar to MLflow but with a nicer live dashboard.

① wandb.init() — presses the record button and opens a new page on wandb.ai for this experiment.
② The config dictionary — writes down all your settings before training starts (like filling in the top of a test paper: name, date, settings).
③ wandb.log() inside the loop — after every single training lap, posts the latest scores to your live dashboard. You can watch the graphs update in real time from your phone! 📱
④ wandb.Artifact — saves the trained model file so you can download it later from any computer.
⑤ run.finish() — presses stop on the recording. The link in the output takes you straight to your results page.
import wandb
import torch

# Initialise a new experiment run
run = wandb.init(
    project="weather-ai",          # group experiments by project
    name="lstm-v1-batch64",        # human-readable run name
    config={                       # log all hyperparameters upfront
        "learning_rate": 0.001,
        "batch_size": 64,
        "epochs": 50,
        "hidden_size": 128,
        "dropout": 0.2,
        "dataset": "global_weather_2024_v3",
        "architecture": "LSTM"
    }
)

config = wandb.config

model = WeatherLSTM(
    hidden_size=config.hidden_size,
    dropout=config.dropout
)
optimizer = torch.optim.Adam(model.parameters(), lr=config.learning_rate)

for epoch in range(config.epochs):
    train_loss = train_one_epoch(model, optimizer)
    val_loss, val_mae = evaluate(model)

    # Log metrics — they appear in real-time in your browser
    wandb.log({
        "train_loss": train_loss,
        "val_loss": val_loss,
        "val_mae": val_mae,
        "epoch": epoch
    })

# Save the model as a W&B Artifact
artifact = wandb.Artifact("weather-lstm", type="model")
torch.save(model.state_dict(), "model.pt")
artifact.add_file("model.pt")
run.log_artifact(artifact)

run.finish()
print(f"View your run at: {run.url}")

W&B Hyperparameter Sweeps — Find the Best Config Automatically

W&B Sweeps automatically search for the best hyperparameter combination using Bayesian optimisation.
Instead of guessing which learning rate to try next, the sweep learns from previous runs and suggests the most promising settings.

🧒 What this code does: Normal hyperparameter search is like trying flavours of ice cream in random order — you might eat 19 bad flavours before finding the best one.
Bayesian optimisation is smarter — after eating a few flavours it starts to figure out which type you like (sweet vs sour, fruity vs creamy) and starts suggesting flavours more likely to be your favourite.

① The sweep_config dictionary defines the "menu" of options to try — ranges for learning_rate, batch_size, hidden_size, and dropout.
② wandb.sweep() creates a manager that runs 20 experiments, learning from each one to suggest smarter settings for the next.
③ wandb.agent() runs the 20 experiments one by one. At the end, W&B draws a chart showing which combination won! 🏆
import wandb

# Define the search space
sweep_config = {
    "method": "bayes",           # Bayesian optimisation — smarter than random!
    "metric": {
        "name": "val_mae",
        "goal": "minimize"       # find settings that minimise MAE
    },
    "parameters": {
        "learning_rate": {
            "distribution": "log_uniform_values",
            "min": 1e-4,
            "max": 1e-1
        },
        "batch_size": {
            "values": [32, 64, 128, 256]
        },
        "hidden_size": {
            "values": [64, 128, 256, 512]
        },
        "dropout": {
            "min": 0.1,
            "max": 0.5
        }
    }
}

# Create the sweep
sweep_id = wandb.sweep(
    sweep=sweep_config,
    project="weather-ai"
)

def train():
    with wandb.init() as run:
        config = run.config
        # ... your training code using config values ...
        wandb.log({"val_mae": val_mae})

# Launch 20 sweep runs — Bayesian will find the best combo
wandb.agent(sweep_id, function=train, count=20)

Part 4 — Data and Dataset Tracking

Tracking your model is not enough.
You must also track your data.
A model is only as good as the data it trained on — if your data changes, your model behaves differently.

Log Dataset Version with MLflow

🧒 What this code does: Imagine you baked the best cake ever but forgot which bag of flour you used. Two months later, you grab a different flour and the cake tastes different — and you have no idea why!
mlflow.log_input() is like stapling the flour bag label to your recipe card.
It records exactly which dataset file (and which version of it) was used to train this model — so you can always go back and use the exact same data again.
import mlflow

with mlflow.start_run():
    # Log dataset as an input — tracks which data version trained this model
    dataset = mlflow.data.from_numpy(
        X_train,
        targets=y_train,
        name="weather_training_data",
        source="s3://my-bucket/weather/train_v2.3.parquet"
    )
    mlflow.log_input(dataset, context="training")

    # Also tag the run with useful data details
    mlflow.set_tags({
        "dataset.version": "v2.3",
        "dataset.size": len(X_train),
        "dataset.date_range": "2020-01-01 to 2024-12-31",
        "git.commit": "abc123def456"
    })

Track Datasets with DVC (Data Version Control)

DVC is the Git for your data files. It stores tiny pointer files in Git while the actual large data files live in cloud storage.

🧒 What this code does: Your data file might be 50 GB — you cannot put that in Git (Git is for code, not giant files!).
DVC is clever: it puts a tiny sticky note in Git that says "the real file is in S3 at this address with this checksum".

① dvc init — sets up DVC in your project (like git init but for data).
② dvc remote add — tells DVC where to actually store the big files (your S3 bucket).
③ dvc add — creates the tiny sticky note (.dvc file) for your big data file and adds the real file to .gitignore.
④ git commit — saves the sticky note into Git.
⑤ dvc push — uploads the actual big file to S3.
Now anyone can run git clone + dvc pull to get the exact same data. Perfect reproducibility! 🎯
# Install DVC
pip install dvc[s3]        # for AWS S3 storage

# Initialise DVC in your project
dvc init

# Tell DVC where to store your large files
dvc remote add -d myremote s3://my-bucket/dvc-store

# Track your training dataset
dvc add data/weather_train_v2.3.parquet

# This creates a tiny .dvc pointer file — commit it to Git!
git add data/weather_train_v2.3.parquet.dvc .gitignore
git commit -m "data: add weather training data v2.3"

# Push the actual big data file to S3
dvc push

# Anyone on the team can now get the exact same data:
git clone your-repo
dvc pull

Part 5 — Automated Model Promotion in a Full MLOps Pipeline

In production, experiment tracking does not run manually — it is embedded in automated training pipelines.
Here is how it all connects:

New Data Arrives (daily/weekly)
        ↓
Data Validation + DVC versioning
        ↓
Training Pipeline triggered (GitHub Actions / Kubeflow)
        ↓
MLflow autolog() runs throughout training
        ↓
All params, metrics, artifacts logged to MLflow Tracking Server
        ↓
Automated model evaluation:
    → Is val_mae better than the current production model?
    → Does it pass fairness and bias checks?
        ↓
If YES → Register in Model Registry → deploy to production
If NO  → Stays in Staging → alert ML team to investigate

Automated Model Promotion Script

🧒 What this code does: This is the robot referee 🤖🏆 that decides if the new model is good enough to go live.

Imagine the current production model is like the reigning champion at a sports tournament.
Every new training run is a challenger who wants to take the title.

① The function first checks the score of the current champion (the model already in production).
② It then checks the score of the new challenger (the model we just trained).
③ If the challenger's MAE is lower (better!) than the champion's — it wins! The function promotes it to "champion" automatically.
④ If the challenger is worse — the current champion keeps the title. Nothing changes in production.
⑤ At the very bottom, we call this function at the end of every training run — so this check happens automatically every single time.
import mlflow
from mlflow.tracking import MlflowClient

client = MlflowClient()

def promote_if_better(new_run_id, metric="val_mae"):
    """
    Automatically promote a new model to champion
    if it beats the current production model.
    """

    # Step 1: Get the current production (champion) model metrics
    try:
        champion = client.get_model_version_by_alias(
            name="weather-prediction-model",
            alias="champion"
        )
        champion_run = client.get_run(champion.run_id)
        champion_mae = champion_run.data.metrics[metric]
        print(f"Current champion MAE: {champion_mae:.4f}")
    except Exception:
        champion_mae = float("inf")   # no champion yet — promote anything
        print("No champion yet — first model promoted automatically")

    # Step 2: Get the challenger's metric
    new_run = client.get_run(new_run_id)
    challenger_mae = new_run.data.metrics[metric]
    print(f"Challenger MAE: {challenger_mae:.4f}")

    # Step 3: Compare and decide
    if challenger_mae < champion_mae:
        new_version = mlflow.register_model(
            model_uri=f"runs:/{new_run_id}/model",
            name="weather-prediction-model"
        )
        client.set_registered_model_alias(
            name="weather-prediction-model",
            alias="champion",
            version=new_version.version
        )
        print(f"✅ Version {new_version.version} promoted to champion!")
        return True
    else:
        print(f"❌ Challenger did not beat champion. Current champion stays.")
        return False


# Step 5: Call this automatically at the end of every training run
promote_if_better(new_run_id="abc123def456")

Part 6 — GenAI and LLM Experiment Tracking

Many teams are training and fine-tuning Large Language Models (LLMs) — not just traditional ML models.
Experiment tracking for LLMs has some unique requirements:

  • Prompt versioning → Track which prompt template produced which output
  • Evaluation scores → Track BLEU, ROUGE, human eval scores, toxicity scores
  • Token costs → Track how many tokens each run consumed (and the $ cost)
  • Response samples → Log actual model responses as artifacts for manual review
🧒 What this code does: This is the same as our regular MLflow tracking — but now for a giant AI language model (like a mini version of ChatGPT).

The special LLM things we track are:
🔢 bleu_score and rouge_l — special scores that measure how close the model's answers are to correct human answers (like a spell-checker score for AI writing).
💰 total_tokens_trained — LLMs process millions of words at a time, which costs real money. Tracking tokens = tracking your cloud bill.
📝 sample_prediction.json — we save an actual example question and the model's real answer, so a human can review it later and say "yes that sounds right" or "no that's wrong".
📁 lora_adapters — LoRA is a fast fine-tuning method. Instead of saving the whole giant model (100GB!), we save only the small changes (a few MB). This gets saved as an artifact.
import mlflow

# MLflow 2.20+ has native LLM tracking support
with mlflow.start_run(run_name="llm-finetune-weather-v3"):

    mlflow.log_params({
        "base_model": "meta-llama/Llama-3-8B",
        "fine_tune_method": "LoRA",
        "lora_rank": 16,
        "lora_alpha": 32,
        "learning_rate": 2e-4,
        "batch_size": 8,
        "gradient_accumulation": 4,
        "epochs": 3,
        "prompt_template_version": "v2.1",
        "training_dataset": "weather-qa-pairs-v3.jsonl"
    })

    # Track LLM-specific metrics
    mlflow.log_metrics({
        "train_loss": 0.423,
        "eval_loss": 0.391,
        "bleu_score": 0.782,        # how accurate the text output is
        "rouge_l": 0.841,           # how complete the answers are
        "perplexity": 12.4,         # how "surprised" the model is by the test data
        "tokens_per_second": 847,   # training speed
        "total_tokens_trained": 15_400_000  # cloud cost indicator
    })

    # Log an actual sample prediction for human review
    sample_outputs = {
        "question": "What will the temperature be in Mumbai tomorrow?",
        "expected": "32°C with 80% humidity",
        "model_output": "Tomorrow in Mumbai: 31°C, high humidity expected"
    }
    with open("sample_prediction.json", "w") as f:
        import json
        json.dump(sample_outputs, f, indent=2)
    mlflow.log_artifact("sample_prediction.json")

    # Log the fine-tuned adapter weights (small file — just the changes)
    mlflow.log_artifact("lora_adapters/")

Part 7 — Setting Up a Production MLflow Server

Running MLflow locally is fine for learning. For a real team, you need a shared server that everyone connects to.

🧒 What this code does: Right now your MLflow saves experiments only on your own laptop. If your laptop dies — everything is gone! 😱
This docker-compose.yaml sets up a proper shared MLflow server for your whole team, with two parts working together:

🖥️ mlflow-server — the MLflow website that everyone visits (like a shared Google Drive for experiments).
🗄️ postgres — a real database that permanently stores all your experiment data (so it never gets lost even if the server restarts).
☁️ S3 bucket — stores the actual model files and artifacts in the cloud (because model files can be gigabytes).

The second code block shows how every data scientist on the team points their laptop at this shared server — so they all log to the same place.
# docker-compose.yaml for production MLflow

services:
  mlflow-server:
    image: ghcr.io/mlflow/mlflow:v2.20.0
    ports:
      - "5000:5000"
    environment:
      MLFLOW_BACKEND_STORE_URI: "postgresql://mlflow:${DB_PASSWORD}@postgres:5432/mlflowdb"
      MLFLOW_DEFAULT_ARTIFACT_ROOT: "s3://my-company-mlflow-artifacts/"
      AWS_ACCESS_KEY_ID: "${AWS_ACCESS_KEY_ID}"
      AWS_SECRET_ACCESS_KEY: "${AWS_SECRET_ACCESS_KEY}"
    command: >
      mlflow server
      --backend-store-uri postgresql://mlflow:${DB_PASSWORD}@postgres:5432/mlflowdb
      --default-artifact-root s3://my-company-mlflow-artifacts/
      --host 0.0.0.0
      --port 5000
      --workers 4
    depends_on:
      - postgres

  postgres:
    image: postgres:16-alpine
    environment:
      POSTGRES_USER: mlflow
      POSTGRES_PASSWORD: "${DB_PASSWORD}"
      POSTGRES_DB: mlflowdb
    volumes:
      - pgdata:/var/lib/postgresql/data

volumes:
  pgdata:
🧒 What this code does: Two quick steps to activate the team server.
① docker-compose up -d — starts the server in the background (like pressing "on" on a router).
② export MLFLOW_TRACKING_URI — tells your laptop "when you log experiments, send them to THIS address instead of your own computer".
Now when any teammate runs python train.py, their experiments show up on the shared team dashboard automatically — no extra code needed! 🎉
# Start the production server
docker-compose up -d

# Every data scientist on the team runs this once:
export MLFLOW_TRACKING_URI=http://mlflow.my-company.com:5000

# Now all training runs automatically log to the shared server
python train.py

Every data scientist on the team now logs to the same shared tracking server. You can see each other's experiments, compare runs across team members, and the Model Registry becomes a company-wide asset.

Choosing the Right Tool — Decision Guide 

Ask yourself these questions to pick the right experiment tracking tool:

  • Want open-source, zero vendor lock-in, self-hosted? → Use MLflow. Industry standard, free forever. ✅
  • Want the most beautiful UI and easiest setup for research? → Use Weights & Biases. Used by OpenAI. ✅
  • Running 10,000+ experiments with strict governance? → Use Neptune.ai. ✅
  • Want full pipeline + tracking in one open-source tool? → Use ClearML. ✅
  • Already fully on AWS? → Use SageMaker Experiments. Zero extra setup. ✅
  • Already fully on Google Cloud? → Use Vertex AI Experiments. ✅
  • Fine-tuning LLMs and tracking prompts? → MLflow 2.20+ or W&B — both have LLM-native support. ✅

Quick Summary 📝

What we learned

  • What experiment tracking is → The lab notebook for ML models. Records params, metrics, artifacts, and metadata from every training run.
  • Why it matters → Without it, you cannot reproduce results, compare experiments, or safely promote models to production.
  • The four things tracked → Parameters, Metrics, Artifacts, and Metadata (code commit, data version, environment).
  • MLflow → Open-source standard. Tracking, Projects, Models, and Model Registry. 14M monthly downloads.
  • Autologging → One line of code captures everything for sklearn, PyTorch, TensorFlow, XGBoost, and Hugging Face.
  • Model Registry → Version and promote models from Staging to champion. Full lineage and rollback.
  • W&B → Beautiful UI, best for research teams, Bayesian hyperparameter sweeps built in.
  • Data tracking → DVC for dataset versioning. Log dataset inputs with mlflow.log_input().
  • GenAI tracking → Track prompts, BLEU/ROUGE scores, token costs, and response samples for LLM fine-tuning.
  • Production setup → MLflow server + PostgreSQL + S3 = team-wide shared tracking server.

Happy tracking! 📊✨

Comments