Imagine you are a scientist trying to bake the perfect chocolate cake 🍰.
You try 50 different recipes — different amounts of sugar, butter, baking time, temperature.
But you forget to write anything down.
Three weeks later you cannot remember which recipe made the best cake!
That is exactly what happens when data scientists train machine learning models without experiment tracking.
They run hundreds of training experiments, change hyperparameters, swap datasets — and forget what worked.
Experiment Tracking is the lab notebook for your ML models.
It records every detail of every training run automatically — so you can always go back, compare, reproduce, and improve.
What Is Experiment Tracking?
When you train a machine learning model, you make dozens of decisions:
- What data did I use? Which version?
- What hyperparameters did I set? Learning rate? Batch size? Number of layers?
- How good was the final model? Accuracy? Loss? F1 score?
- Which model file should go to production?
Experiment tracking tools record all of this automatically — every single time you run training.
No manual spreadsheets. No "I think I used 0.001 learning rate… maybe 0.01?"
Everything is stored, timestamped, searchable, and comparable.
💡 Think of it like: Google Drive for your ML experiments.
Every training run gets its own folder with all the evidence — code, data version, settings, results, and the actual model file.
You can pull up any run from 6 months ago and reproduce it exactly.
Why Experiment Tracking Matters — The Problem Without It
Without tracking, here is what a typical ML team's life looks like:
Monday: Run experiment with lr=0.01 → accuracy 78% (not saved)
Tuesday: Run experiment with lr=0.001 → accuracy 84% (forgot settings)
Wednesday: Run experiment with lr=0.0001 → accuracy 81% (can't reproduce)
Thursday: "Which experiment had 84% again?? I need to deploy THAT model!"
Friday: 😭 Nobody knows. Start over.
With experiment tracking, the same week looks like:
Monday: Run #001 lr=0.01 → 78% ← tracked automatically ✅
Tuesday: Run #002 lr=0.001 → 84% ← tracked automatically ✅
Wednesday: Run #003 lr=0.0001 → 81% ← tracked automatically ✅
Thursday: "Run #002 had 84%. Here is the model file, the dataset hash,
the exact code commit, and all hyperparameters."
Friday: 🚀 Deployed to production in 10 minutes.
That is the power of experiment tracking. It turns chaotic research into reproducible engineering.
The 4 Things Every Tracking Tool Records
Every experiment tracking tool — no matter which one you pick — records these four things:
- Parameters → The settings you chose before training starts. Learning rate, batch size, number of epochs, model architecture, optimizer type.
- Metrics → The numbers that measure how good your model is. Training loss, validation accuracy, F1 score, AUC-ROC — recorded at every epoch.
- Artifacts → The files your training produces. The saved model file, plots, confusion matrix images, sample predictions, trained weights.
- Metadata → Everything else about the context. Which Git commit ran, which dataset version was used, which Python version, which GPU, start time, end time, duration.
Together these four things mean you can completely reconstruct any past experiment — even years later. That is what reproducibility means in MLOps.
The Toolchain — What Everyone Is Using
The experiment tracking landscape has settled into three clear choices depending on your team size and budget:
- MLflow → The open-source de facto standard. 20,000+ GitHub stars, 14 million monthly downloads. Free to self-host. Used by Netflix, Shopify, Databricks, Zillow. Best if you want zero vendor lock-in.
- Weights & Biases (W&B) → The developer-favourite. Beautiful UI, easiest setup, best visualisations. Used by OpenAI, NVIDIA, Stability AI, Microsoft. Best for research teams and individuals.
- Neptune.ai → The enterprise-grade metadata database. Best for large teams running thousands of experiments with strict governance requirements.
Other valid choices in the ecosystem:
- ClearML → Fully open-source, includes pipeline management and experiment tracking in one tool
- Comet ML → Strong visualisation and LLM experiment support, great for GenAI teams
- DVC → Git-native tracking for data and models, now part of lakeFS ecosystem
- AWS SageMaker Experiments → Best choice if your entire stack is already on AWS
- Google Vertex AI Experiments → Best choice if your stack is on Google Cloud
For this post we go deep on MLflow — the most widely used open-source option — and also show W&B for comparison.
Part 1 — MLflow: The Open-Source Standard
MLflow has four components that work together. Think of them as four departments in the same office:
- MLflow Tracking → The lab notebook. Records parameters, metrics, and artifacts from every run.
- MLflow Projects → The reproducibility department. Packages your code so anyone can rerun it exactly.
- MLflow Models → The packaging department. Wraps your model in a standard format that works with any serving framework.
- MLflow Model Registry → The promotion department. Manages model versions and their lifecycle — Staging → Production → Archived.
Step 1: Install MLflow
Line 1 installs MLflow onto your computer using pip (the Python app store).
Line 2 checks the version to make sure it installed correctly.
Line 3 opens a website on your computer at
localhost:5000 where you can see all your experiments in a nice table — like opening your lab notebook in a browser.
# Install MLflow (Python 3.8+ required)
pip install mlflow
# Verify the installation
mlflow --version
# Output: mlflow, version 2.20.0
# Launch the local tracking server (opens in your browser)
mlflow ui
# Visit http://localhost:5000 to see the MLflow dashboard
Step 2: Your First Tracked Experiment
Step by step what happens:
① We make some fake weather data (like temperature readings) to practise with.
② We press "record" by calling
mlflow.start_run() — everything after this is captured.③ We write down our settings (alpha, max_iter) — like writing "I used 2 cups of sugar" in a recipe.
④ We train the model — this is the actual learning part.
⑤ We measure how good it is (MAE — lower is better).
⑥ We save the result and the model file.
⑦ The
with block closes — MLflow saves everything to your dashboard automatically. Go to localhost:5000 and see it!
import mlflow
import mlflow.sklearn
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split
import numpy as np
# --- Fake weather data for the example ---
X = np.random.randn(1000, 10) # 1000 samples, 10 features
y = np.random.randn(1000) # temperature values
X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2)
# --- The hyperparameters we want to test ---
alpha = 0.5 # regularisation strength
max_iter = 1000 # maximum iterations
# ✨ Everything between start_run() and end is automatically tracked ✨
with mlflow.start_run(run_name="ridge-alpha-0.5"):
# Log the hyperparameters we chose
mlflow.log_param("alpha", alpha)
mlflow.log_param("max_iter", max_iter)
mlflow.log_param("model_type", "Ridge")
mlflow.log_param("dataset_version", "v2.3")
# Train the model
model = Ridge(alpha=alpha, max_iter=max_iter)
model.fit(X_train, y_train)
# Evaluate
predictions = model.predict(X_val)
mae = mean_absolute_error(y_val, predictions)
# Log the result metrics
mlflow.log_metric("mae", mae)
mlflow.log_metric("train_samples", len(X_train))
mlflow.log_metric("val_samples", len(X_val))
# Save the trained model as an artifact
mlflow.sklearn.log_model(model, "weather-model")
print(f"Run complete! MAE = {mae:.4f}")
print(f"View at: http://localhost:5000")
What happens automatically:
- MLflow creates a unique Run ID for this experiment
- All parameters are stored — alpha, max_iter, model_type, dataset_version
- The MAE metric is stored and plotted over time
- The trained model is saved and downloadable from the UI
- The Git commit hash is captured automatically
- Start time, end time, and duration are recorded
Open http://localhost:5000 in your browser and you will see your experiment with all these details in a clean table. 🎯
Step 3: Autologging — Track Everything with One Line
MLflow can automatically capture every parameter and metric from popular libraries — with just one line of code:
mlflow.sklearn.autolog() is like having a robot helper who follows you around the kitchen and writes everything down FOR you.You don't write a single
log_param or log_metric line — the robot does it all the moment training finishes. Magic! ✨
import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import train_test_split
# Enable autologging — captures EVERYTHING automatically
mlflow.sklearn.autolog()
X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2)
with mlflow.start_run(run_name="random-forest-autolog"):
model = RandomForestRegressor(n_estimators=100, max_depth=5)
model.fit(X_train, y_train)
# That's it! MLflow captured all params, metrics, and the model.
Autologging works with sklearn, PyTorch, TensorFlow/Keras, XGBoost, LightGBM, Hugging Face Transformers, and more.
For deep learning models it logs the loss and metrics at every single epoch automatically.
Step 4: Logging Metrics at Every Epoch (Deep Learning)
For neural networks you want to track how loss and accuracy improve during training:
That is what this code does — after every training epoch (one lap of learning), it calls out the train_loss and val_mae scores and writes them into your notebook.
At the end, MLflow draws a smooth graph showing how your model got smarter and smarter over 50 laps.
You can then compare that graph against other training runs to see which one learned the fastest!
import mlflow
import torch
mlflow.pytorch.autolog() # one line — logs everything!
with mlflow.start_run(run_name="weather-lstm-v1"):
mlflow.log_params({
"epochs": 50,
"learning_rate": 0.001,
"batch_size": 64,
"hidden_size": 128,
"num_layers": 2,
"dropout": 0.2,
"optimizer": "Adam",
"dataset": "global_weather_2024_v3"
})
model = WeatherLSTM(hidden_size=128, num_layers=2, dropout=0.2)
optimizer = torch.optim.Adam(model.parameters(), lr=0.001)
for epoch in range(50):
train_loss = train_one_epoch(model, optimizer)
val_loss, val_mae = evaluate(model)
# Log metrics at each epoch — MLflow plots these automatically
mlflow.log_metrics({
"train_loss": train_loss,
"val_loss": val_loss,
"val_mae": val_mae
}, step=epoch)
# Save the trained model
mlflow.pytorch.log_model(model, "weather-lstm")
In the MLflow UI you will see smooth line charts showing how train_loss and val_loss decreased over 50 epochs — for every single run. You can overlay multiple runs to compare them visually. 📈
Step 5: Comparing Multiple Runs — Grid Search
Now let's run many experiments — trying different hyperparameters — and let MLflow track all of them at once:
You try small/medium/large sizes with thin/thick crusts and 3 different cheeses = 27 combinations total.
Instead of tasting one pizza per day (27 days!), this code bakes all 27 automatically, measures how delicious each one is (val_mae score), and writes everything in a neat table.
At the end you just look at the table, sort by tastiest, and you know which pizza wins — in minutes instead of a month! 🏆
import mlflow
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_absolute_error
mlflow.set_experiment("weather-model-hyperparameter-search")
# Try all combinations — 27 runs total
learning_rates = [0.001, 0.01, 0.1]
n_estimators_list = [50, 100, 200]
max_depths = [3, 5, 10]
for lr in learning_rates:
for n_est in n_estimators_list:
for depth in max_depths:
with mlflow.start_run(
run_name=f"rf-lr{lr}-n{n_est}-d{depth}"
):
mlflow.log_params({
"learning_rate": lr,
"n_estimators": n_est,
"max_depth": depth
})
model = RandomForestRegressor(
n_estimators=n_est,
max_depth=depth
)
model.fit(X_train, y_train)
mae = mean_absolute_error(y_val, model.predict(X_val))
mlflow.log_metric("val_mae", mae)
mlflow.sklearn.log_model(model, "model")
Expected MLflow UI output — a sortable comparison table:
Run Name lr n_est depth val_mae
─────────────────────────────────────────────────────────────
rf-lr0.01-n200-d10 0.01 200 10 0.1823 ← BEST ✅
rf-lr0.01-n100-d10 0.01 100 10 0.1891
rf-lr0.001-n200-d10 0.001 200 10 0.1934
rf-lr0.1-n50-d3 0.1 50 3 0.2891 ← WORST
Sort by val_mae ascending and your best model is instantly visible. Click on it to see every detail.
This is why teams run 50 experiments in one afternoon instead of one per day. 🏎️
Part 2 — The MLflow Model Registry
Finding the best experiment is step one. Promoting it to production safely is step two.
That is what the Model Registry handles.
Think of it like a promotion process at work 👔:
Junior → Senior → Manager → VP.
A model goes through: None → Staging → Production → Archived.
Step 1 puts the pizza on a "trying it out" menu (Staging) — only internal staff test it.
Step 2 puts it on the main menu for all customers (champion = Production).
Step 3 is how any chef in the kitchen can grab the exact same recipe later with one command.
No more emailing model files around — everyone loads from the same registered source!
import mlflow
from mlflow.tracking import MlflowClient
client = MlflowClient()
# Step 1: Register a model from a run into the registry
model_uri = "runs:/abc123def456/model" # run ID from your best experiment
registered_model = mlflow.register_model(
model_uri=model_uri,
name="weather-prediction-model"
)
print(f"Registered as version {registered_model.version}")
# Output: Registered as version 1
# Step 2: Promote version 1 to Staging (internal testing)
client.set_registered_model_alias(
name="weather-prediction-model",
alias="staging",
version="1"
)
# Step 3: After testing passes — promote to Production!
client.set_registered_model_alias(
name="weather-prediction-model",
alias="champion", # 'champion' is the 2026 standard alias for production
version="1"
)
# Step 4: Load the production model anywhere in your codebase
champion_model = mlflow.pyfunc.load_model(
"models:/weather-prediction-model@champion"
)
predictions = champion_model.predict(new_weather_data)
print(predictions)
Querying the Registry for the Best Run Automatically
It searches ALL your experiments, finds the one with the lowest val_mae score (lowest error = best model), and then registers it straight into the Model Registry without any human clicking.
Think of it like a robot judge at a science fair that reads every project, picks the winner, and hands them the trophy — instantly! 🤖🏆
from mlflow.tracking import MlflowClient
client = MlflowClient()
# Find the single best run across ALL experiments — by lowest val_mae
best_run = client.search_runs(
experiment_ids=["1"],
filter_string="metrics.val_mae < 0.20",
order_by=["metrics.val_mae ASC"],
max_results=1
)[0]
print(f"Best Run ID: {best_run.info.run_id}")
print(f"Best MAE: {best_run.data.metrics['val_mae']:.4f}")
print(f"Parameters: {best_run.data.params}")
# Automatically register the best run in one step
mlflow.register_model(
model_uri=f"runs:/{best_run.info.run_id}/model",
name="weather-prediction-model"
)
Part 3 — Weights & Biases: The Developer-Favourite
W&B is loved by researchers for its beautiful real-time dashboards and zero-configuration setup.
Used by OpenAI, NVIDIA, Stability AI, Microsoft, and most of the top AI research labs worldwide.
Step 1: Install and Set Up W&B
Logging in gives W&B permission to save your experiments on their website at
wandb.ai.After this, every training run you do will appear on a beautiful online dashboard — like your own personal science fair display board that updates in real time while the model trains! 🎨
# Install
pip install wandb
# Log in (creates free account at wandb.ai)
wandb login
# Paste your API key from wandb.ai/authorize
Step 2: Track an Experiment with W&B
①
wandb.init() — presses the record button and opens a new page on wandb.ai for this experiment.② The
config dictionary — writes down all your settings before training starts (like filling in the top of a test paper: name, date, settings).③
wandb.log() inside the loop — after every single training lap, posts the latest scores to your live dashboard. You can watch the graphs update in real time from your phone! 📱④
wandb.Artifact — saves the trained model file so you can download it later from any computer.⑤
run.finish() — presses stop on the recording. The link in the output takes you straight to your results page.
import wandb
import torch
# Initialise a new experiment run
run = wandb.init(
project="weather-ai", # group experiments by project
name="lstm-v1-batch64", # human-readable run name
config={ # log all hyperparameters upfront
"learning_rate": 0.001,
"batch_size": 64,
"epochs": 50,
"hidden_size": 128,
"dropout": 0.2,
"dataset": "global_weather_2024_v3",
"architecture": "LSTM"
}
)
config = wandb.config
model = WeatherLSTM(
hidden_size=config.hidden_size,
dropout=config.dropout
)
optimizer = torch.optim.Adam(model.parameters(), lr=config.learning_rate)
for epoch in range(config.epochs):
train_loss = train_one_epoch(model, optimizer)
val_loss, val_mae = evaluate(model)
# Log metrics — they appear in real-time in your browser
wandb.log({
"train_loss": train_loss,
"val_loss": val_loss,
"val_mae": val_mae,
"epoch": epoch
})
# Save the model as a W&B Artifact
artifact = wandb.Artifact("weather-lstm", type="model")
torch.save(model.state_dict(), "model.pt")
artifact.add_file("model.pt")
run.log_artifact(artifact)
run.finish()
print(f"View your run at: {run.url}")
W&B Hyperparameter Sweeps — Find the Best Config Automatically
W&B Sweeps automatically search for the best hyperparameter combination using Bayesian optimisation.
Instead of guessing which learning rate to try next, the sweep learns from previous runs and suggests the most promising settings.
Bayesian optimisation is smarter — after eating a few flavours it starts to figure out which type you like (sweet vs sour, fruity vs creamy) and starts suggesting flavours more likely to be your favourite.
① The
sweep_config dictionary defines the "menu" of options to try — ranges for learning_rate, batch_size, hidden_size, and dropout.②
wandb.sweep() creates a manager that runs 20 experiments, learning from each one to suggest smarter settings for the next.③
wandb.agent() runs the 20 experiments one by one. At the end, W&B draws a chart showing which combination won! 🏆
import wandb
# Define the search space
sweep_config = {
"method": "bayes", # Bayesian optimisation — smarter than random!
"metric": {
"name": "val_mae",
"goal": "minimize" # find settings that minimise MAE
},
"parameters": {
"learning_rate": {
"distribution": "log_uniform_values",
"min": 1e-4,
"max": 1e-1
},
"batch_size": {
"values": [32, 64, 128, 256]
},
"hidden_size": {
"values": [64, 128, 256, 512]
},
"dropout": {
"min": 0.1,
"max": 0.5
}
}
}
# Create the sweep
sweep_id = wandb.sweep(
sweep=sweep_config,
project="weather-ai"
)
def train():
with wandb.init() as run:
config = run.config
# ... your training code using config values ...
wandb.log({"val_mae": val_mae})
# Launch 20 sweep runs — Bayesian will find the best combo
wandb.agent(sweep_id, function=train, count=20)
Part 4 — Data and Dataset Tracking
Tracking your model is not enough.
You must also track your data.
A model is only as good as the data it trained on — if your data changes, your model behaves differently.
Log Dataset Version with MLflow
mlflow.log_input() is like stapling the flour bag label to your recipe card.It records exactly which dataset file (and which version of it) was used to train this model — so you can always go back and use the exact same data again.
import mlflow
with mlflow.start_run():
# Log dataset as an input — tracks which data version trained this model
dataset = mlflow.data.from_numpy(
X_train,
targets=y_train,
name="weather_training_data",
source="s3://my-bucket/weather/train_v2.3.parquet"
)
mlflow.log_input(dataset, context="training")
# Also tag the run with useful data details
mlflow.set_tags({
"dataset.version": "v2.3",
"dataset.size": len(X_train),
"dataset.date_range": "2020-01-01 to 2024-12-31",
"git.commit": "abc123def456"
})
Track Datasets with DVC (Data Version Control)
DVC is the Git for your data files. It stores tiny pointer files in Git while the actual large data files live in cloud storage.
DVC is clever: it puts a tiny sticky note in Git that says "the real file is in S3 at this address with this checksum".
①
dvc init — sets up DVC in your project (like git init but for data).②
dvc remote add — tells DVC where to actually store the big files (your S3 bucket).③
dvc add — creates the tiny sticky note (.dvc file) for your big data file and adds the real file to .gitignore.④
git commit — saves the sticky note into Git.⑤
dvc push — uploads the actual big file to S3.Now anyone can run
git clone + dvc pull to get the exact same data. Perfect reproducibility! 🎯
# Install DVC
pip install dvc[s3] # for AWS S3 storage
# Initialise DVC in your project
dvc init
# Tell DVC where to store your large files
dvc remote add -d myremote s3://my-bucket/dvc-store
# Track your training dataset
dvc add data/weather_train_v2.3.parquet
# This creates a tiny .dvc pointer file — commit it to Git!
git add data/weather_train_v2.3.parquet.dvc .gitignore
git commit -m "data: add weather training data v2.3"
# Push the actual big data file to S3
dvc push
# Anyone on the team can now get the exact same data:
git clone your-repo
dvc pull
Part 5 — Automated Model Promotion in a Full MLOps Pipeline
In production, experiment tracking does not run manually — it is embedded in automated training pipelines.
Here is how it all connects:
New Data Arrives (daily/weekly)
↓
Data Validation + DVC versioning
↓
Training Pipeline triggered (GitHub Actions / Kubeflow)
↓
MLflow autolog() runs throughout training
↓
All params, metrics, artifacts logged to MLflow Tracking Server
↓
Automated model evaluation:
→ Is val_mae better than the current production model?
→ Does it pass fairness and bias checks?
↓
If YES → Register in Model Registry → deploy to production
If NO → Stays in Staging → alert ML team to investigate
Automated Model Promotion Script
Imagine the current production model is like the reigning champion at a sports tournament.
Every new training run is a challenger who wants to take the title.
① The function first checks the score of the current champion (the model already in production).
② It then checks the score of the new challenger (the model we just trained).
③ If the challenger's MAE is lower (better!) than the champion's — it wins! The function promotes it to "champion" automatically.
④ If the challenger is worse — the current champion keeps the title. Nothing changes in production.
⑤ At the very bottom, we call this function at the end of every training run — so this check happens automatically every single time.
import mlflow
from mlflow.tracking import MlflowClient
client = MlflowClient()
def promote_if_better(new_run_id, metric="val_mae"):
"""
Automatically promote a new model to champion
if it beats the current production model.
"""
# Step 1: Get the current production (champion) model metrics
try:
champion = client.get_model_version_by_alias(
name="weather-prediction-model",
alias="champion"
)
champion_run = client.get_run(champion.run_id)
champion_mae = champion_run.data.metrics[metric]
print(f"Current champion MAE: {champion_mae:.4f}")
except Exception:
champion_mae = float("inf") # no champion yet — promote anything
print("No champion yet — first model promoted automatically")
# Step 2: Get the challenger's metric
new_run = client.get_run(new_run_id)
challenger_mae = new_run.data.metrics[metric]
print(f"Challenger MAE: {challenger_mae:.4f}")
# Step 3: Compare and decide
if challenger_mae < champion_mae:
new_version = mlflow.register_model(
model_uri=f"runs:/{new_run_id}/model",
name="weather-prediction-model"
)
client.set_registered_model_alias(
name="weather-prediction-model",
alias="champion",
version=new_version.version
)
print(f"✅ Version {new_version.version} promoted to champion!")
return True
else:
print(f"❌ Challenger did not beat champion. Current champion stays.")
return False
# Step 5: Call this automatically at the end of every training run
promote_if_better(new_run_id="abc123def456")
Part 6 — GenAI and LLM Experiment Tracking
Many teams are training and fine-tuning Large Language Models (LLMs) — not just traditional ML models.
Experiment tracking for LLMs has some unique requirements:
- Prompt versioning → Track which prompt template produced which output
- Evaluation scores → Track BLEU, ROUGE, human eval scores, toxicity scores
- Token costs → Track how many tokens each run consumed (and the $ cost)
- Response samples → Log actual model responses as artifacts for manual review
The special LLM things we track are:
🔢 bleu_score and rouge_l — special scores that measure how close the model's answers are to correct human answers (like a spell-checker score for AI writing).
💰 total_tokens_trained — LLMs process millions of words at a time, which costs real money. Tracking tokens = tracking your cloud bill.
📝 sample_prediction.json — we save an actual example question and the model's real answer, so a human can review it later and say "yes that sounds right" or "no that's wrong".
📁 lora_adapters — LoRA is a fast fine-tuning method. Instead of saving the whole giant model (100GB!), we save only the small changes (a few MB). This gets saved as an artifact.
import mlflow
# MLflow 2.20+ has native LLM tracking support
with mlflow.start_run(run_name="llm-finetune-weather-v3"):
mlflow.log_params({
"base_model": "meta-llama/Llama-3-8B",
"fine_tune_method": "LoRA",
"lora_rank": 16,
"lora_alpha": 32,
"learning_rate": 2e-4,
"batch_size": 8,
"gradient_accumulation": 4,
"epochs": 3,
"prompt_template_version": "v2.1",
"training_dataset": "weather-qa-pairs-v3.jsonl"
})
# Track LLM-specific metrics
mlflow.log_metrics({
"train_loss": 0.423,
"eval_loss": 0.391,
"bleu_score": 0.782, # how accurate the text output is
"rouge_l": 0.841, # how complete the answers are
"perplexity": 12.4, # how "surprised" the model is by the test data
"tokens_per_second": 847, # training speed
"total_tokens_trained": 15_400_000 # cloud cost indicator
})
# Log an actual sample prediction for human review
sample_outputs = {
"question": "What will the temperature be in Mumbai tomorrow?",
"expected": "32°C with 80% humidity",
"model_output": "Tomorrow in Mumbai: 31°C, high humidity expected"
}
with open("sample_prediction.json", "w") as f:
import json
json.dump(sample_outputs, f, indent=2)
mlflow.log_artifact("sample_prediction.json")
# Log the fine-tuned adapter weights (small file — just the changes)
mlflow.log_artifact("lora_adapters/")
Part 7 — Setting Up a Production MLflow Server
Running MLflow locally is fine for learning. For a real team, you need a shared server that everyone connects to.
This
docker-compose.yaml sets up a proper shared MLflow server for your whole team, with two parts working together:🖥️ mlflow-server — the MLflow website that everyone visits (like a shared Google Drive for experiments).
🗄️ postgres — a real database that permanently stores all your experiment data (so it never gets lost even if the server restarts).
☁️ S3 bucket — stores the actual model files and artifacts in the cloud (because model files can be gigabytes).
The second code block shows how every data scientist on the team points their laptop at this shared server — so they all log to the same place.
# docker-compose.yaml for production MLflow
services:
mlflow-server:
image: ghcr.io/mlflow/mlflow:v2.20.0
ports:
- "5000:5000"
environment:
MLFLOW_BACKEND_STORE_URI: "postgresql://mlflow:${DB_PASSWORD}@postgres:5432/mlflowdb"
MLFLOW_DEFAULT_ARTIFACT_ROOT: "s3://my-company-mlflow-artifacts/"
AWS_ACCESS_KEY_ID: "${AWS_ACCESS_KEY_ID}"
AWS_SECRET_ACCESS_KEY: "${AWS_SECRET_ACCESS_KEY}"
command: >
mlflow server
--backend-store-uri postgresql://mlflow:${DB_PASSWORD}@postgres:5432/mlflowdb
--default-artifact-root s3://my-company-mlflow-artifacts/
--host 0.0.0.0
--port 5000
--workers 4
depends_on:
- postgres
postgres:
image: postgres:16-alpine
environment:
POSTGRES_USER: mlflow
POSTGRES_PASSWORD: "${DB_PASSWORD}"
POSTGRES_DB: mlflowdb
volumes:
- pgdata:/var/lib/postgresql/data
volumes:
pgdata:
①
docker-compose up -d — starts the server in the background (like pressing "on" on a router).②
export MLFLOW_TRACKING_URI — tells your laptop "when you log experiments, send them to THIS address instead of your own computer".Now when any teammate runs
python train.py, their experiments show up on the shared team dashboard automatically — no extra code needed! 🎉
# Start the production server
docker-compose up -d
# Every data scientist on the team runs this once:
export MLFLOW_TRACKING_URI=http://mlflow.my-company.com:5000
# Now all training runs automatically log to the shared server
python train.py
Every data scientist on the team now logs to the same shared tracking server. You can see each other's experiments, compare runs across team members, and the Model Registry becomes a company-wide asset.
Choosing the Right Tool — Decision Guide
Ask yourself these questions to pick the right experiment tracking tool:
- Want open-source, zero vendor lock-in, self-hosted? → Use MLflow. Industry standard, free forever. ✅
- Want the most beautiful UI and easiest setup for research? → Use Weights & Biases. Used by OpenAI. ✅
- Running 10,000+ experiments with strict governance? → Use Neptune.ai. ✅
- Want full pipeline + tracking in one open-source tool? → Use ClearML. ✅
- Already fully on AWS? → Use SageMaker Experiments. Zero extra setup. ✅
- Already fully on Google Cloud? → Use Vertex AI Experiments. ✅
- Fine-tuning LLMs and tracking prompts? → MLflow 2.20+ or W&B — both have LLM-native support. ✅
Quick Summary 📝
What we learned
- What experiment tracking is → The lab notebook for ML models. Records params, metrics, artifacts, and metadata from every training run.
- Why it matters → Without it, you cannot reproduce results, compare experiments, or safely promote models to production.
- The four things tracked → Parameters, Metrics, Artifacts, and Metadata (code commit, data version, environment).
- MLflow → Open-source standard. Tracking, Projects, Models, and Model Registry. 14M monthly downloads.
- Autologging → One line of code captures everything for sklearn, PyTorch, TensorFlow, XGBoost, and Hugging Face.
- Model Registry → Version and promote models from Staging to champion. Full lineage and rollback.
- W&B → Beautiful UI, best for research teams, Bayesian hyperparameter sweeps built in.
- Data tracking → DVC for dataset versioning. Log dataset inputs with
mlflow.log_input(). - GenAI tracking → Track prompts, BLEU/ROUGE scores, token costs, and response samples for LLM fine-tuning.
- Production setup → MLflow server + PostgreSQL + S3 = team-wide shared tracking server.
Happy tracking! 📊✨
Comments
Post a Comment