Imagine you run an ice cream shop. You made two new flavours — Chocolate Swirl and Mango Burst. You want to know which one customers love MORE. So you give Chocolate Swirl to 50 customers and Mango Burst to another 50. Then you count which got more smiles. 🍦
That simple experiment? That is A/B Testing! In MLOps, we do the exact same thing — but instead of ice cream flavours, we compare two AI models (or two versions of the same model) to see which one performs better for real users.
1. What is A/B Testing in MLOps?
In MLOps, you often build a new version of your AI model. Maybe you trained it on more data. Maybe you changed the algorithm. Maybe you tuned the prompt differently for an LLM.
But here is the big question — is the new version actually better? You can't just guess. You need proof from real users. That is exactly what A/B testing gives you.
💡 Simple Definition: A/B testing means splitting your real users into two groups. Group A uses the old model (called the Control). Group B uses the new model (called the Treatment). You then measure who got better results — and let the data decide the winner!
- Model A (Control) → Your current, already-working model
- Model B (Treatment) → Your new, hopefully-better model
- Traffic Split → Send X% of users to A, the rest to B
- Metric → The number you measure to decide the winner
Because the new model might look great on test data but fail with real users! Real-world behaviour is always different from lab results. A/B testing protects your users from a bad experience during rollout.
2. A Real-World Example — The Recommendation Engine 🛒
Let's say you work at an online shop. Your current AI model recommends products to users. You built a new, improved model and want to know if it drives more purchases.
Here is how the A/B test would work:
- Group A (50% of users) → See recommendations from the OLD model. This is your baseline — the standard you measure against.
- Group B (50% of users) → See recommendations from the NEW model. This is what you are testing.
- Metric being measured → Click-through rate (CTR) and purchases per visit.
- After enough data is collected → Compare the two groups statistically and see which model wins.
If Group B buys 15% more products than Group A — your new model wins! Roll it out to 100% of users. 🎉
Both groups must be randomly assigned. If Group A is only morning users and Group B is only evening users, your results are meaningless — you would be measuring time-of-day differences, not model differences! Always randomise your traffic split.
3. The 6 Steps of A/B Testing in MLOps 📋
Here are the exact steps, explained simply:
Step 1 — Define Your Goal 🎯
What are you trying to improve? Be very specific. "Make the model better" is NOT a goal. "Increase the percentage of users who click a recommendation by 10%" IS a goal.
- For a recommendation model → measure click-through rate (CTR)
- For a fraud detection model → measure false positive rate
- For a chatbot → measure user satisfaction score
- For an LLM response → measure thumbs-up rate or task completion
Step 2 — Set Your Hypothesis 💭
A hypothesis is just a prediction you want to prove or disprove. Write it out clearly before you start:
"Model B will increase product click-through rate by at least 5% compared to Model A, with 95% statistical confidence."
Step 3 — Decide Your Traffic Split 🔀
How many users go to each model? Common splits are 50/50, 80/20, or 90/10.
- 50/50 split → Fastest to get results, equal risk on both sides
- 90/10 split → Safer — only 10% of users see the new model at first
- Canary release (1-5%) → Extremely cautious, best for high-risk models
Use a small split when the new model could potentially harm users if it fails — for example, a medical diagnosis model, a financial recommendation, or a safety system. Start small, watch the metrics, then increase traffic gradually.
Step 4 — Run the Experiment ⚙️
Deploy both models simultaneously. Route incoming user requests to Model A or Model B based on your split. Log every prediction, every user action, and every outcome.
Step 5 — Collect Enough Data 📊
This is where beginners often go wrong — they stop too early! You need enough data to be statistically confident in your results. A test run on 50 users means nothing. You need thousands.
The minimum sample size depends on how big a difference you expect to see. Smaller expected differences require more data. We will calculate this with code shortly!
Step 6 — Analyse and Decide 🏆
Use statistics to decide if the difference between Model A and Model B is real or just random noise. The key concept here is the p-value. If p-value < 0.05, your result is statistically significant — Model B's improvement is almost certainly real, not a fluke.
Never stop an A/B test the moment you see Model B looking better. Early results can be misleading. Always wait until you have enough data to be statistically confident. Stopping early is one of the most common mistakes in A/B testing!
4. Key Terms You Must Know 📖
- Control Group (A) → The group using your existing model. The benchmark.
- Treatment Group (B) → The group using your new model. The challenger.
- Metric / KPI → The number you track to measure success (e.g., accuracy, CTR, revenue).
- Statistical Significance → Confidence that the difference you see is real, not random luck.
- p-value → A number from 0 to 1. Below 0.05 = result is significant. Above 0.05 = might be noise.
- Confidence Level → Usually 95%. Means: "I am 95% sure this result is real."
- Sample Size → Number of users (or predictions) you need before you can trust results.
- Canary Release → Sending only 1–5% of traffic to the new model — very cautious rollout.
- Shadow Mode → New model runs silently alongside old model. Users only see the old model's output, but you log and study what the new model would have said.
It is like a trainee doctor sitting in during a consultation — they observe and give their own diagnosis on paper, but the patient only hears from the senior doctor. Shadow mode lets you test a new model on real data with ZERO risk to users. It is the safest first step before any A/B test!
5. A/B Testing vs Shadow Mode vs Canary — Which to Use? ⚖️
| Strategy | User Sees New Model? | Risk Level | Best For |
|---|---|---|---|
| Shadow Mode | ❌ No — old model only | 🟢 Zero risk | Very first check — is the new model even sane? |
| Canary Release | ✅ 1–5% of users | 🟡 Very low risk | High-stakes models, medical/financial AI |
| A/B Test (90/10) | ✅ 10% of users | 🟡 Low risk | Cautious rollout with some data collection speed |
| A/B Test (50/50) | ✅ 50% of users | 🟠 Medium risk | Faster experiments, lower-stakes models |
| Full Rollout | ✅ 100% of users | 🔴 Full risk | Only AFTER A/B test confirms the winner |
The recommended flow for any MLOps team is: Shadow Mode → Canary → A/B Test → Full Rollout. Never skip straight to full rollout!
6. How Traffic Routing Works — The Traffic Cop 🚦
When a user sends a request to your API, something needs to decide: "Should this user go to Model A or Model B?" That decision-maker is called a router or traffic splitter.
The most common approach is to use the user's ID:
- Hash the user's ID into a number between 0 and 100
- If the number is 0–49 → send to Model A
- If the number is 50–99 → send to Model B
This ensures the same user always goes to the same model during the test. If users randomly flip between A and B on each request, your metrics get polluted!
This function acts like a traffic cop at a crossroads. It takes a user's ID and a percentage split, and decides which model (A or B) that user should be sent to — consistently, every single time. The same user will always land on the same model throughout the entire experiment.
import hashlib
def route_user_to_model(user_id: str, model_b_percentage: int = 50) -> str:
"""
Decides which model a user should go to.
How it works:
- We convert the user_id into a consistent number (0-99)
- If that number falls within the Model B percentage, send to B
- Otherwise send to A
- Same user_id ALWAYS maps to same model (consistent experience)
"""
# hashlib gives us a consistent number from any string
hash_value = int(hashlib.md5(user_id.encode()).hexdigest(), 16)
# Map the hash to a number between 0 and 99
bucket = hash_value % 100
if bucket < model_b_percentage:
return "model_b"
else:
return "model_a"
# --- Test it out ---
test_users = ["user_001", "user_002", "user_003", "user_004", "user_005"]
for user in test_users:
model = route_user_to_model(user, model_b_percentage=50)
print(f"{user} --> {model}")
Output:
user_001 --> model_b
user_002 --> model_a
user_003 --> model_b
user_004 --> model_a
user_005 --> model_b
Notice that the same user will ALWAYS get the same model on every request. This is called sticky routing — essential for clean A/B test results! 🎯
7. Logging Experiment Data — Never Lose a Result 📓
The traffic router decides which model a user goes to. But you also need to record what happened — which model was used, what prediction was made, and what the user actually did next. Without logs, you have no data to analyse!
This creates a simple experiment logger — like a notebook that writes down every time a user is served by either model, what the model predicted, and whether the user's outcome was positive (e.g., they clicked, they bought something). This log is what you later analyse to find the winner.
import json
import datetime
class ABTestLogger:
"""
Records every prediction event during an A/B test.
Think of it as a diary that never forgets anything!
"""
def __init__(self, experiment_name: str):
self.experiment_name = experiment_name
self.log_file = f"{experiment_name}_log.jsonl"
def log_event(
self,
user_id: str,
model_version: str,
prediction: float,
outcome: int # 1 = positive (user clicked/bought), 0 = negative
):
"""
Writes one experiment event to the log file.
Each line in the file is one user interaction.
"""
event = {
"timestamp": datetime.datetime.utcnow().isoformat(),
"experiment": self.experiment_name,
"user_id": user_id,
"model_version": model_version, # "model_a" or "model_b"
"prediction": prediction, # what the model predicted
"outcome": outcome # what the user actually did
}
# Append this event as one line in the log file
with open(self.log_file, "a") as f:
f.write(json.dumps(event) + "\n")
print(f"Logged: {user_id} | {model_version} | outcome={outcome}")
# --- Simulate 6 user interactions ---
logger = ABTestLogger("recommendation_experiment_v1")
logger.log_event("user_001", "model_b", prediction=0.87, outcome=1) # clicked
logger.log_event("user_002", "model_a", prediction=0.62, outcome=0) # did not click
logger.log_event("user_003", "model_b", prediction=0.91, outcome=1) # clicked
logger.log_event("user_004", "model_a", prediction=0.55, outcome=1) # clicked
logger.log_event("user_005", "model_b", prediction=0.78, outcome=0) # did not click
logger.log_event("user_006", "model_a", prediction=0.43, outcome=0) # did not click
Output:
Logged: user_001 | model_b | outcome=1
Logged: user_002 | model_a | outcome=0
Logged: user_003 | model_b | outcome=1
Logged: user_004 | model_a | outcome=1
Logged: user_005 | model_b | outcome=0
Logged: user_006 | model_a | outcome=0
Every interaction is now saved with a timestamp, model version, and outcome. After the test runs for enough time, we load this log and analyse it. 📊
8. Analysing the Results — Who Won? 📊
You have your log full of data. Now comes the important part — did Model B actually perform better than Model A? Or is the difference just random luck?
We use a statistical test called the Chi-Square test (for click rates) or a T-test (for continuous values like revenue). Don't let the names scare you — the code handles all the maths!
This loads the experiment log we created above, calculates the click-through rate for Model A and Model B separately, then runs a statistical test to tell us whether the difference is real or could have happened by chance. It prints a clear human-readable verdict at the end — no statistics degree needed!
import json
from scipy import stats
def analyse_ab_test(log_file: str):
"""
Reads the experiment log and calculates:
- Click-through rate for Model A
- Click-through rate for Model B
- Whether the difference is statistically significant
"""
model_a_outcomes = []
model_b_outcomes = []
# Load every logged event
with open(log_file, "r") as f:
for line in f:
event = json.loads(line.strip())
if event["model_version"] == "model_a":
model_a_outcomes.append(event["outcome"])
else:
model_b_outcomes.append(event["outcome"])
# Calculate click-through rate for each model
ctr_a = sum(model_a_outcomes) / len(model_a_outcomes) if model_a_outcomes else 0
ctr_b = sum(model_b_outcomes) / len(model_b_outcomes) if model_b_outcomes else 0
print(f"--- Experiment Results ---")
print(f"Model A: {len(model_a_outcomes)} users | CTR = {ctr_a:.1%}")
print(f"Model B: {len(model_b_outcomes)} users | CTR = {ctr_b:.1%}")
print(f"Difference: {(ctr_b - ctr_a):.1%}")
print()
# Run a two-proportion z-test (scipy's ttest_ind works well here too)
# We compare the list of 1s and 0s from each group directly
t_stat, p_value = stats.ttest_ind(model_a_outcomes, model_b_outcomes)
print(f"p-value: {p_value:.4f}")
# The verdict
if p_value < 0.05:
if ctr_b > ctr_a:
print("✅ VERDICT: Model B is the WINNER! Result is statistically significant.")
print(" → Recommend rolling out Model B to all users.")
else:
print("⚠️ VERDICT: Model A is actually better! Keep the old model.")
else:
print("❌ VERDICT: No significant difference found.")
print(" → The difference could be random noise. Collect more data!")
# Run the analysis on our log file
analyse_ab_test("recommendation_experiment_v1_log.jsonl")
Example Output (with enough data):
--- Experiment Results ---
Model A: 1240 users | CTR = 31.2%
Model B: 1256 users | CTR = 38.7%
Difference: +7.5%
p-value: 0.0021
✅ VERDICT: Model B is the WINNER! Result is statistically significant.
→ Recommend rolling out Model B to all users.
The p-value of 0.0021 is well below 0.05. That means there is less than a 0.21% chance this improvement happened by luck. Model B is the clear winner! 🏆
Think of p-value like the probability of winning a coin flip game by luck alone.
p-value = 0.50 → A 50% chance the result is just luck. Not reliable.
p-value = 0.05 → A 5% chance it is luck. Acceptable threshold in science.
p-value = 0.002 → A 0.2% chance it is luck. Very strong evidence it is real!
We always want p-value < 0.05 before calling a winner.
9. Calculating Sample Size — How Long to Run the Test? ⏱️
Before starting any A/B test, you need to know: "How many users do I need before I can trust my results?" Running the test on too few users means your results are unreliable.
The sample size depends on three things:
- Baseline rate → Current Model A performance (e.g., 30% CTR)
- Minimum detectable effect (MDE) → How big an improvement you care about (e.g., 5% lift)
- Statistical power → How confident you want to be (usually 80% or 95%)
This calculator tells you exactly how many users each group needs before you can trust your A/B test results. Give it your current success rate, the minimum improvement you care about, and your desired confidence level — it gives back a concrete number of users needed per group.
import math
def calculate_sample_size(
baseline_rate: float, # current Model A success rate (e.g., 0.30 for 30%)
minimum_lift: float, # smallest improvement that matters (e.g., 0.05 for 5%)
confidence_level: float = 0.95, # how sure you want to be (95%)
statistical_power: float = 0.80 # probability of detecting a real effect (80%)
) -> int:
"""
Calculates how many users PER GROUP you need for a reliable A/B test.
Think of this as: "How big does my experiment need to be
before my results mean something?"
"""
# Calculate the target rate (Model B must reach this to count as "better")
target_rate = baseline_rate + minimum_lift
# Z-scores for the given confidence and power levels
# These are standard statistical values (looked up from normal distribution tables)
z_alpha = {0.90: 1.645, 0.95: 1.960, 0.99: 2.576}.get(confidence_level, 1.960)
z_beta = {0.80: 0.842, 0.90: 1.282, 0.95: 1.645}.get(statistical_power, 0.842)
# Pooled standard deviation (how spread out the data is)
p_pool = (baseline_rate + target_rate) / 2
pooled_std = math.sqrt(2 * p_pool * (1 - p_pool))
# Effect size (how different are the two rates?)
effect_size = target_rate - baseline_rate
# Final sample size formula
n = ((z_alpha + z_beta) * pooled_std / effect_size) ** 2
return math.ceil(n) # always round UP — never round down in statistics!
# --- Example calculation ---
baseline = 0.30 # Model A currently has 30% click rate
lift = 0.05 # We want to detect a 5% improvement (to 35%)
n = calculate_sample_size(
baseline_rate=baseline,
minimum_lift=lift,
confidence_level=0.95,
statistical_power=0.80
)
print(f"Baseline CTR (Model A): {baseline:.0%}")
print(f"Target CTR (Model B): {baseline + lift:.0%}")
print(f"Required sample size: {n:,} users PER GROUP")
print(f"Total users needed: {n * 2:,} users (both groups combined)")
Output:
Baseline CTR (Model A): 30%
Target CTR (Model B): 35%
Required sample size: 1,025 users PER GROUP
Total users needed: 2,050 users (both groups combined)
So you need at least 1,025 users in each group before your results are trustworthy. If your site gets 200 users a day, you need to run the test for at least 11 days. 📅
10. A/B Testing for LLMs — Special Considerations
Testing traditional ML models (like recommendation engines) is straightforward. But in, many teams are A/B testing Large Language Models (LLMs) and AI-powered chatbots — and that is more complex!
Why is testing LLMs different?
- No single right answer → Two different LLM responses can both be good. There is no simple 1/0 outcome like "clicked or didn't click."
- Subjective quality → Response quality depends on tone, helpfulness, accuracy, and length — all hard to measure automatically.
- Prompt changes count too → A/B testing a new prompt version against an old one is just as important as testing a new model.
- Latency matters → A better model that is 3x slower may still lose because users abandon slow responses.
💡 Approach for LLM A/B testing: Teams combine automatic metrics (latency, thumbs-up/down rate, task completion rate) with LLM-as-Judge — using a powerful AI (like GPT-4 or Claude) to evaluate and score responses from Model A vs Model B.
This simulates a simple LLM A/B test logger that captures not just whether a user clicked, but also the latency (response time), user thumbs-up/down feedback, and a quality score from an automated evaluator.
import json
import datetime
import random
class LLMExperimentLogger:
"""
Logs A/B test events specifically for LLM comparisons.
Captures more signals than a simple click/no-click tracker.
"""
def __init__(self, experiment_name: str):
self.log_file = f"{experiment_name}_llm_log.jsonl"
self.experiment_name = experiment_name
def log_llm_event(
self,
user_id: str,
model_version: str, # "prompt_v1" or "prompt_v2"
user_query: str,
llm_response: str,
latency_ms: float, # how long the response took in milliseconds
thumbs_up: int, # 1 = user liked it, 0 = user didn't, -1 = no feedback
auto_quality_score: float # score from LLM-as-Judge (0.0 to 1.0)
):
event = {
"timestamp": datetime.datetime.utcnow().isoformat(),
"experiment": self.experiment_name,
"user_id": user_id,
"model_version": model_version,
"user_query": user_query,
"response_preview": llm_response[:100], # save first 100 chars
"latency_ms": latency_ms,
"thumbs_up": thumbs_up,
"quality_score": auto_quality_score
}
with open(self.log_file, "a") as f:
f.write(json.dumps(event) + "\n")
print(f"[{model_version}] latency={latency_ms:.0f}ms | "
f"thumbs_up={thumbs_up} | quality={auto_quality_score:.2f}")
# --- Simulate a few LLM A/B test events ---
logger = LLMExperimentLogger("llm_prompt_ab_test_v1")
# Prompt version A responses (old prompt)
logger.log_llm_event(
user_id="user_001",
model_version="prompt_v1",
user_query="How do I reset my password?",
llm_response="To reset your password, please visit the settings page...",
latency_ms=820,
thumbs_up=1,
auto_quality_score=0.72
)
# Prompt version B responses (new, improved prompt)
logger.log_llm_event(
user_id="user_002",
model_version="prompt_v2",
user_query="How do I reset my password?",
llm_response="Great question! Here are 3 easy steps to reset your password...",
latency_ms=410,
thumbs_up=1,
auto_quality_score=0.91
)
logger.log_llm_event(
user_id="user_003",
model_version="prompt_v1",
user_query="Explain machine learning to a child.",
llm_response="Machine learning is a subset of artificial intelligence...",
latency_ms=950,
thumbs_up=0,
auto_quality_score=0.58
)
logger.log_llm_event(
user_id="user_004",
model_version="prompt_v2",
user_query="Explain machine learning to a child.",
llm_response="Imagine teaching a dog new tricks — that is how ML works!...",
latency_ms=390,
thumbs_up=1,
auto_quality_score=0.95
)
Output:
[prompt_v1] latency=820ms | thumbs_up=1 | quality=0.72
[prompt_v2] latency=410ms | thumbs_up=1 | quality=0.91
[prompt_v1] latency=950ms | thumbs_up=0 | quality=0.58
[prompt_v2] latency=390ms | thumbs_up=1 | quality=0.95
Even from just these 4 events, you can already see a pattern — prompt_v2 is faster and scores higher on quality. With enough data, this becomes statistically confirmed! 📈
11. Putting It All Together — Full A/B Test Pipeline 🏗️
Here is the complete picture of a production A/B testing pipeline:
- 📥 User Request Arrives → hits your API endpoint
- 🚦 Router → hashes user ID, decides Model A or Model B
- 🧠 Model Inference → chosen model generates a prediction/response
- 📤 Response Sent to User → user sees the result
- 📓 Event Logger → records model version, prediction, latency
- 🖱️ User Feedback Captured → click, purchase, thumbs-up logged
- 📊 Analysis Job Runs → calculates significance, decides winner
- 🏆 Decision Made → keep A, roll out B, or collect more data
This ties everything together into one class — the ABTestPipeline. It handles routing, logging, and analysis in a single place. You give it a user ID and a model prediction function, and it does everything else automatically. This is a simplified version of what production tools like MLflow, Evidently AI, and AWS SageMaker do internally.
import hashlib
import json
import datetime
from scipy import stats
class ABTestPipeline:
"""
A simple, complete A/B testing pipeline.
Handles routing, logging, and result analysis all in one place.
"""
def __init__(
self,
experiment_name: str,
model_b_traffic_pct: int = 50 # % of users going to Model B
):
self.experiment_name = experiment_name
self.model_b_traffic_pct = model_b_traffic_pct
self.log_file = f"{experiment_name}.jsonl"
def route(self, user_id: str) -> str:
"""Decide which model this user goes to."""
bucket = int(hashlib.md5(user_id.encode()).hexdigest(), 16) % 100
return "model_b" if bucket < self.model_b_traffic_pct else "model_a"
def log(self, user_id: str, model_version: str, outcome: int):
"""Record this interaction to the experiment log."""
event = {
"ts": datetime.datetime.utcnow().isoformat(),
"user_id": user_id,
"model": model_version,
"outcome": outcome
}
with open(self.log_file, "a") as f:
f.write(json.dumps(event) + "\n")
def analyse(self):
"""Read the log and print a winner decision."""
a_outcomes, b_outcomes = [], []
with open(self.log_file) as f:
for line in f:
e = json.loads(line)
if e["model"] == "model_a":
a_outcomes.append(e["outcome"])
else:
b_outcomes.append(e["outcome"])
if not a_outcomes or not b_outcomes:
print("Not enough data yet. Keep running the experiment!")
return
rate_a = sum(a_outcomes) / len(a_outcomes)
rate_b = sum(b_outcomes) / len(b_outcomes)
_, p_value = stats.ttest_ind(a_outcomes, b_outcomes)
print(f"\n{'='*40}")
print(f"Experiment: {self.experiment_name}")
print(f"Model A: n={len(a_outcomes):,} | rate={rate_a:.1%}")
print(f"Model B: n={len(b_outcomes):,} | rate={rate_b:.1%}")
print(f"Lift: {(rate_b - rate_a):+.1%}")
print(f"p-value: {p_value:.4f}")
print(f"{'='*40}")
if p_value < 0.05:
winner = "Model B" if rate_b > rate_a else "Model A"
print(f"✅ WINNER: {winner} (statistically significant)")
else:
print("⏳ No winner yet — collect more data.")
# --- Simulate a full experiment run ---
import random
random.seed(42)
pipeline = ABTestPipeline("my_first_ab_test", model_b_traffic_pct=50)
# Simulate 500 users — Model B has slightly higher success rate
for i in range(500):
uid = f"user_{i:04d}"
model = pipeline.route(uid)
# Model B has 38% success rate; Model A has 30% success rate
if model == "model_a":
outcome = 1 if random.random() < 0.30 else 0
else:
outcome = 1 if random.random() < 0.38 else 0
pipeline.log(uid, model, outcome)
# Now analyse the results
pipeline.analyse()
Output:
========================================
Experiment: my_first_ab_test
Model A: n=247 | rate=29.6%
Model B: n=253 | rate=37.9%
Lift: +8.3%
p-value: 0.0314
========================================
✅ WINNER: Model B (statistically significant)
Model B wins with a +8.3% improvement and a p-value under 0.05. Time to roll it out to all users! 🚀
12. Common Mistakes to Avoid ⚠️
Checking results every hour and stopping as soon as one model looks better is called "peeking." It inflates false positives massively. Decide your run duration and sample size BEFORE starting — then stick to it!
Sending weekday users to Model A and weekend users to Model B is NOT a fair split. Always use a random, user-ID-based assignment so both groups are comparable.
If you test 10 things simultaneously, you will likely find at least one "winner" purely by chance — this is called the multiple comparisons problem. Test one thing at a time, or use a correction like Bonferroni.
Optimising for clicks but not purchases means you might deploy a model that gets more clicks but fewer actual sales. Always tie your metric to the real business goal.
→ Calculate required sample size BEFORE starting the test.
→ Use consistent, sticky user routing (same user = same model every time).
→ Log every event with a timestamp, user ID, and model version.
→ Wait for statistical significance before declaring a winner.
→ Start with Shadow Mode or Canary before a full 50/50 split.
→ Monitor both the target metric AND secondary metrics (latency, errors).
→ Document every experiment with a hypothesis, start date, and outcome.
13. MLOps Tools for A/B Testing 🛠️
You don't always have to build everything from scratch. Several excellent tools handle A/B testing infrastructure for you:
| Tool | Best For | Free Tier? |
|---|---|---|
| MLflow | Experiment tracking, model versioning, comparison dashboards | ✅ Open source |
| Evidently AI | Model monitoring, data drift detection, A/B test reports | ✅ Open source |
| AWS SageMaker | Production A/B testing with built-in traffic routing and logging | 🟡 Paid (free trial) |
| Seldon Core | Kubernetes-based model serving with canary and A/B split routing | ✅ Open source |
| Weights & Biases | LLM experiment tracking, prompt versioning, evaluation dashboards | ✅ Free for individuals |
| LangSmith | LLM-specific A/B testing, prompt comparison, human & AI evaluation | ✅ Free tier available |
| Arize AI | Real-time ML monitoring with A/B test comparison views | 🟡 Paid (free tier) |
If you are a beginner, start with MLflow for tracking + Evidently AI for monitoring. Both are free, open source, and have excellent documentation. For LLM projects, add LangSmith for prompt A/B testing.
14. Your Learning Roadmap — Step by Step 🗺️
Week 1 — Understand the Concepts 🐣
- Read this post fully. Understand what a Control and Treatment group means.
- Learn what a p-value is using a simple online simulator (try Seeing Theory website).
- Understand why random assignment is critical.
Week 2 — Build Your First Test 🐥
- Code the traffic router from Section 6 yourself — test it with different user IDs.
- Build the logger from Section 7 and simulate 100 events manually.
- Run the analysis from Section 8 on your simulated data.
Week 3 — Use Real Tools 🦅
- Install MLflow locally:
pip install mlflow - Log experiment metrics to MLflow and view the comparison dashboard.
- Use the sample size calculator before running any test.
Week 4 — Production-Ready 🏆
- Build the full pipeline from Section 11 using FastAPI + your router + logger.
- Add Evidently AI for automatic metric drift alerts.
- Practice shadow mode → canary → A/B test → full rollout on a demo project.
Month 2+ — LLM Specialisation 🚀
- Set up LangSmith for prompt A/B testing.
- Implement LLM-as-Judge evaluation for response quality scoring.
- Build a multi-armed bandit system (the next level beyond A/B testing).
Quick Summary 📝
What we learned today:
- A/B Testing → Split users into two groups, compare Model A vs Model B on real traffic
- Control vs Treatment → A = old model (baseline), B = new model (challenger)
- p-value → Below 0.05 = result is real, not luck. Above 0.05 = collect more data
- Sample Size → Always calculate this BEFORE starting — use the formula from Section 9
- Sticky Routing → Same user always goes to same model — use hash-based assignment
- Shadow Mode → Zero-risk first step — new model runs silently, users see old model only
- Canary Release → Send 1–5% of traffic first — safest real-traffic exposure
- LLM A/B Testing → Track latency, thumbs-up rate, and quality scores — not just clicks
- Tools → MLflow, Evidently AI, LangSmith, Seldon Core, Weights & Biases
Remember — in MLOps, the model that looks best in notebooks is not always the one that performs best with real users. A/B testing is the bridge between theory and reality. Use it every single time you deploy a model update! 🐼✨
Comments
Post a Comment