You've fine-tuned a language model. It generates responses. Now comes the hardest question in all of LLM engineering: how do you know if those responses are actually good?
Measuring "goodness" of text is surprisingly deep. Should you give each response a score from 1 to 5? Or should you compare two responses head-to-head and pick the winner? These two approaches — Absolute Scoring and Pairwise Ranking — are the two fundamental paradigms for evaluating AI outputs, and understanding their tradeoffs is one of the most important skills in LLM engineering.
💡 Think of it like: evaluating figure skating 🏂. A judge could give each skater a score out of 10 (absolute scoring), or watch two skaters back-to-back and pick the better one (pairwise ranking). Same goal. Very different process. Very different results.
Why Evaluation Is So Hard for Language Models
With most software, testing is binary — a function either returns the correct output or it doesn't. Language model outputs don't work like this. Two responses can both be factually correct, grammatically perfect, and still differ enormously in how helpful or appropriate they are.
Consider this prompt: "What should I eat before a workout?"
- Response A: "Eat a banana or a small bowl of oats 30–60 minutes before your workout for sustained energy."
- Response B: "Pre-workout nutrition depends on your training type, intensity, body weight, and metabolic rate. Generally, a combination of easily digestible carbohydrates with a modest amount of protein, consumed approximately 45 to 90 minutes prior to exercise onset, provides optimal glycogen replenishment while avoiding gastrointestinal discomfort..."
Both are correct. How do you programmatically measure which is better? This is the evaluation challenge. There are two main frameworks to solve it.
Framework 1 — Absolute Scoring
What Is It?
In absolute scoring, you evaluate each response in isolation and assign it a numerical score on a fixed scale — typically 1 to 5 or 1 to 10. The score represents how good the response is according to a defined rubric, without any reference to other responses.
Think of it like a restaurant health inspection 🍽️. An inspector visits each restaurant and rates it from 0 to 100 based on a checklist. They don't compare Restaurant A to Restaurant B — each gets its own score against the standard.
📐 How Absolute Scoring Works
💬 Prompt
"What should I eat before a workout?"
Response A → Score: 5/5 ✅
Evaluated alone against rubric
Response B → Score: 2/5 ❌
Evaluated alone against rubric
📊 Result
Each response gets its own score. No comparison needed.
The Absolute Scoring Rubric
Absolute scoring only works well when you have a clear, detailed rubric. Here is a standard 5-point rubric used in LLM evaluation:
- Score 5 — Excellent: Fully addresses the prompt. Accurate, clear, well-structured, appropriate length, safe, and directly useful to the user.
- Score 4 — Good: Addresses the prompt well with only minor gaps. Accurate and clear but perhaps slightly too long, slightly incomplete, or missing a small detail.
- Score 3 — Acceptable: Partially addresses the prompt. Some useful content, but also some inaccuracies, unnecessary padding, or missing key points.
- Score 2 — Poor: Mostly fails to address the prompt. Significant inaccuracies, very incomplete, irrelevant content, or inappropriate tone.
- Score 1 — Unacceptable: Completely fails. Wrong, harmful, nonsensical, or refuses a reasonable request without justification.
Absolute Scoring in Code — Single Response Evaluation
from openai import OpenAI
import json
client = OpenAI()
ABSOLUTE_SCORING_SYSTEM_PROMPT = """You are an expert evaluator of AI assistant responses.
Score each response on a scale of 1 to 5 based on the following rubric:
Score 5 - Excellent: Fully addresses the prompt. Accurate, clear,
appropriately concise, and directly useful.
Score 4 - Good: Addresses the prompt well with only minor gaps or imperfections.
Score 3 - Acceptable: Partially addresses the prompt with some useful content
but notable issues (inaccuracies, verbosity, missing details).
Score 2 - Poor: Mostly fails the prompt. Significant problems with accuracy,
completeness, or appropriateness.
Score 1 - Unacceptable: Completely fails. Wrong, harmful, or nonsensical.
Return ONLY a valid JSON object with this structure:
{
"score": ,
"reasoning": "",
"strengths": ["", ""],
"weaknesses": ["", ""]
}"""
def absolute_score(prompt: str, response: str) -> dict:
"""
Evaluate a single model response using absolute scoring.
Returns a structured score with reasoning.
"""
user_content = f"""PROMPT: {prompt}
RESPONSE TO EVALUATE:
{response}
Apply the rubric strictly and return the JSON score."""
result = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": ABSOLUTE_SCORING_SYSTEM_PROMPT},
{"role": "user", "content": user_content}
],
temperature=0,
response_format={"type": "json_object"}
)
return json.loads(result.choices[0].message.content)
# Example evaluation
score_result = absolute_score(
prompt="What should I eat before a workout?",
response="Eat a banana or a small bowl of oats 30–60 minutes before your workout for sustained energy."
)
print(json.dumps(score_result, indent=2))
Output:
{
"score": 5,
"reasoning": "Gives two specific, practical food options with a clear timing guideline —
directly and concisely answers the question.",
"strengths": [
"Specific food recommendations with no vagueness",
"Includes actionable timing guidance (30–60 minutes)"
],
"weaknesses": []
}
Evaluating Multiple Responses at Scale
import pandas as pd
from tqdm import tqdm
def evaluate_dataset_absolute(eval_data: list) -> pd.DataFrame:
"""
Run absolute scoring across a full evaluation dataset.
Each item in eval_data should have 'prompt' and 'response' fields.
Returns a DataFrame with scores and analysis.
"""
results = []
for item in tqdm(eval_data, desc="Scoring responses"):
try:
score_data = absolute_score(item["prompt"], item["response"])
results.append({
"prompt": item["prompt"][:60] + "...",
"response_preview": item["response"][:80] + "...",
"score": score_data["score"],
"reasoning": score_data["reasoning"],
"strengths": ", ".join(score_data.get("strengths", [])),
"weaknesses": ", ".join(score_data.get("weaknesses", []))
})
except Exception as e:
results.append({
"prompt": item["prompt"][:60] + "...",
"score": None,
"reasoning": f"Error: {str(e)}"
})
df = pd.DataFrame(results)
return df
# Example dataset
eval_data = [
{
"prompt": "What should I eat before a workout?",
"response": "Eat a banana or oats 30-60 minutes before for sustained energy."
},
{
"prompt": "How do I stay motivated when learning a new skill?",
"response": "Set small daily goals, track your progress visibly, and celebrate small wins.
Motivation follows action — start before you feel ready."
},
{
"prompt": "Explain what RAM is to a 10-year-old.",
"response": "RAM is your computer's short-term memory — it holds the things your computer is
currently working on so it can access them super fast. When you close a program,
that memory is cleared, like erasing a whiteboard."
}
]
df = evaluate_dataset_absolute(eval_data)
print(df[["prompt", "score", "reasoning"]].to_string())
print(f"\nAverage Score: {df['score'].mean():.2f} / 5.0")
print(f"Score Distribution:\n{df['score'].value_counts().sort_index()}")
Output:
prompt score reasoning
0 What should I eat before a workout?... 5 Specific, practical, and directly answers the question.
1 How do I stay motivated when learning... 5 Actionable advice with an insightful motivational reframe.
2 Explain what RAM is to a 10-year-old.... 5 Perfect analogy for the age group, accurate and memorable.
Average Score: 5.00 / 5.0
Score Distribution:
5 3
Name: score, dtype: int64
temperature=0.
Framework 2 — Pairwise Ranking
What Is It?
In pairwise ranking, instead of scoring each response individually, you present two responses side by side and simply ask: "Which one is better?"
This is how humans naturally make quality judgements. We are much better at comparing two things directly than we are at assigning an abstract number to a single thing in isolation.
💡 Think of it like: a cooking competition 👨🍳. A judge tastes two dishes and picks the tastier one. It's far easier than being asked "rate this dish out of 10 in absolute terms." Even professional food critics find comparative judgements more reliable than absolute ones.
📐 How Pairwise Ranking Works
💬 Prompt
"What to eat before a workout?"
Response A
"Banana or oats 30–60 min before..."
Response B
"Pre-workout nutrition depends on many complex factors including..."
🏆
A wins!
No score needed — just a winner
Pairwise Ranking in Code — Head-to-Head Comparison
from openai import OpenAI
import json
import random
client = OpenAI()
PAIRWISE_JUDGE_SYSTEM_PROMPT = """You are an expert evaluator comparing two AI assistant responses.
Given a prompt and two responses (A and B), determine which response is better overall.
Evaluate based on:
1. Accuracy — is the information correct?
2. Helpfulness — does it actually solve the user's need?
3. Clarity — is it easy to understand?
4. Appropriate length — not too short, not unnecessarily long?
5. Tone — professional, friendly, and appropriate for the context?
Return ONLY a valid JSON object with this exact structure:
{
"winner": "A" or "B" or "tie",
"confidence": "high" or "medium" or "low",
"reasoning": "",
"key_difference": ""
}"""
def pairwise_compare(prompt: str, response_a: str, response_b: str,
randomize_order: bool = True) -> dict:
"""
Compare two responses head-to-head using an LLM judge.
randomize_order: If True, randomly swaps A and B to reduce position bias.
The result is always returned in the original A/B order.
"""
# Track if we swapped
swapped = False
if randomize_order and random.random() > 0.5:
response_a, response_b = response_b, response_a
swapped = True
user_content = f"""PROMPT: {prompt}
RESPONSE A:
{response_a}
RESPONSE B:
{response_b}
Which response is better? Apply the evaluation criteria."""
result = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": PAIRWISE_JUDGE_SYSTEM_PROMPT},
{"role": "user", "content": user_content}
],
temperature=0,
response_format={"type": "json_object"}
)
judgement = json.loads(result.choices[0].message.content)
# Correct for the swap so results are always in original A/B terms
if swapped and judgement["winner"] == "A":
judgement["winner"] = "B"
elif swapped and judgement["winner"] == "B":
judgement["winner"] = "A"
return judgement
# Example comparison
result = pairwise_compare(
prompt="What should I eat before a workout?",
response_a="Eat a banana or a small bowl of oats 30–60 minutes before your workout for sustained energy.",
response_b="Pre-workout nutrition depends on your training type, intensity, body weight, and metabolic rate.
Generally, a combination of easily digestible carbohydrates consumed approximately 45 to 90 minutes prior
to exercise provides optimal glycogen replenishment..."
)
print(json.dumps(result, indent=2))
Output:
{
"winner": "A",
"confidence": "high",
"reasoning": "Response A gives two specific, actionable recommendations with a clear timing guideline.
Response B is accurate but buries practical advice under jargon and unnecessary complexity for this question.",
"key_difference": "Directness and actionability — A gets immediately to the point while B over-explains"
}
Clean, structured, explainable decision — made without any absolute score! 🎯
The Position Bias Problem — And How to Fix It
Position bias is one of the most studied problems in pairwise evaluation. Research consistently shows that LLM judges prefer whichever response appears first in the prompt — especially when responses are close in quality.
The fix is simple: run every comparison twice, in both orders, and apply consistency filtering.
def robust_pairwise_compare(prompt: str, response_a: str, response_b: str) -> dict:
"""
Robust pairwise comparison that eliminates position bias.
Runs each comparison twice (A-then-B and B-then-A).
Only declares a winner if both orderings agree.
Returns 'tie' if results are inconsistent — indicating responses are very close in quality.
"""
# First pass: A presented first
result_ab = pairwise_compare(prompt, response_a, response_b, randomize_order=False)
# Second pass: B presented first (labels are corrected back by function)
result_ba = pairwise_compare(prompt, response_b, response_a, randomize_order=False)
# Flip the labels for result_ba since we deliberately reversed the order
if result_ba["winner"] == "A":
result_ba["winner"] = "B"
elif result_ba["winner"] == "B":
result_ba["winner"] = "A"
# Only trust the result if both passes agree
if result_ab["winner"] == result_ba["winner"]:
final_winner = result_ab["winner"]
consistent = True
else:
final_winner = "tie"
consistent = False
return {
"winner": final_winner,
"consistent": consistent,
"pass_1": result_ab["winner"],
"pass_2": result_ba["winner"],
"reasoning": result_ab["reasoning"] if consistent else "Inconsistent — responses are very close in quality.",
"key_difference": result_ab.get("key_difference", "")
}
result = robust_pairwise_compare(
prompt="Explain photosynthesis simply.",
response_a="Plants make food from sunlight. They absorb water
from roots and carbon dioxide from air, then use sunlight energy
to convert these into glucose (their food) and oxygen — which they release into the air.",
response_b="Photosynthesis is the process where plants use light
energy to synthesise carbohydrates from carbon dioxide and water,
releasing oxygen as a by-product through a complex series of light-dependent and light-independent reactions."
)
print(json.dumps(result, indent=2))
Output:
{
"winner": "A",
"consistent": true,
"pass_1": "A",
"pass_2": "A",
"reasoning": "Response A uses accessible language with a clear cause-and-effect
structure, appropriate for a simple explanation.
Response B introduces technical terminology that contradicts the 'simply' instruction.",
"key_difference": "Instruction-following — A respects the 'simply' constraint, B does not"
}
Both passes agree: A wins consistently. We can be confident this isn't a position bias artefact! ✅
Side-by-Side Comparison — Absolute vs Pairwise
Now that you've seen both frameworks in action, let's put their core differences side by side to understand when each one is the right tool:
📊 Absolute Scoring vs Pairwise Ranking — Full Comparison
| Dimension | Absolute Scoring | Pairwise Ranking |
|---|---|---|
| What it measures | How good is this response on its own? | Which of these two responses is better? |
| Output type | A number (e.g., 4 out of 5) | A winner (A beats B, or tie) |
| Requires rubric? | ✅ Yes — essential for consistent scores | Optional — humans compare intuitively |
| Reliability | Moderate — scores drift without strong anchors | High — comparison is a more natural human task |
| Scalability | ✅ High — O(n) evaluations needed | ⚠️ Lower — O(n²) pairs at full scale |
| Best for | Monitoring model quality over time | Choosing between two model versions |
| Used in | Continuous evaluation dashboards | A/B testing, DPO dataset creation |
| Main weakness | Score inflation — everything drifts toward 4–5 | Position bias, can't compare across different evaluations |
The Biggest Weakness of Each — Score Inflation and the Intransitivity Problem
Absolute Scoring's Achilles Heel — Score Inflation
Over time, absolute scores from LLM judges drift upward. Response quality might stay constant, but GPT-4 or Claude tends to rate responses more generously across successive evaluation sessions. This is called score inflation — and it completely undermines longitudinal comparison.
If your average score goes from 3.8 last month to 4.2 this month, did your model improve — or did the judge get more generous? Without careful anchoring, you simply cannot tell.
import random
def detect_score_inflation(evaluator_fn, anchor_responses: list, n_rounds: int = 5) -> dict:
"""
Detect score inflation by tracking scores on fixed anchor responses over time.
Anchor responses are known-quality examples that should always score the same.
If their scores drift upward across rounds, you have score inflation.
"""
round_scores = {}
for round_num in range(1, n_rounds + 1):
scores_this_round = []
for anchor in anchor_responses:
score_data = evaluator_fn(anchor["prompt"], anchor["response"])
scores_this_round.append(score_data["score"])
avg_score = sum(scores_this_round) / len(scores_this_round)
round_scores[f"Round {round_num}"] = avg_score
print(f"Round {round_num}: Average anchor score = {avg_score:.2f}")
# Check for drift
first = round_scores["Round 1"]
last = round_scores[f"Round {n_rounds}"]
drift = last - first
print(f"\nScore drift over {n_rounds} rounds: {drift:+.2f}")
if abs(drift) > 0.3:
print("⚠️ Score inflation detected! Recalibrate your rubric or anchors.")
else:
print("✅ Scores are stable — no significant inflation detected.")
return round_scores
# Define anchor responses with known quality levels
anchor_responses = [
{
"prompt": "What is 15% of 200?",
"response": "30", # should always score 5
"expected_score": 5
},
{
"prompt": "Explain gravity to a child.",
"response": "Things fall because the earth pulls them down.", # should score 4
"expected_score": 4
},
{
"prompt": "What is photosynthesis?",
"response": "It's a thing plants do with sunlight somehow.", # should score 2
"expected_score": 2
}
]
results = detect_score_inflation(absolute_score, anchor_responses, n_rounds=3)
Output:
Round 1: Average anchor score = 3.67
Round 2: Average anchor score = 3.67
Round 3: Average anchor score = 3.67
Score drift over 3 rounds: +0.00
✅ Scores are stable — no significant inflation detected.
Pairwise Ranking's Achilles Heel — Intransitivity
Pairwise ranking has its own subtle problem: intransitivity. In a perfect world, if A beats B and B beats C, then A should also beat C. In practice, pairwise evaluations sometimes produce circular results: A beats B, B beats C, but C beats A.
This happens because different response pairs highlight different qualities. When you compare A to B, the judge focuses on clarity. When you compare B to C, the judge focuses on accuracy. When you compare C to A, the judge focuses on brevity — and suddenly C wins.
def check_transitivity(compare_fn, prompt: str,
response_a: str, response_b: str, response_c: str) -> dict:
"""
Check whether pairwise comparisons between three responses are transitive.
If A > B and B > C, we expect A > C.
Intransitivity indicates the responses are very close in quality
or that the judge is inconsistent across different quality dimensions.
"""
result_ab = compare_fn(prompt, response_a, response_b)
result_bc = compare_fn(prompt, response_b, response_c)
result_ac = compare_fn(prompt, response_a, response_c)
ab = result_ab["winner"] # "A", "B", or "tie"
bc = result_bc["winner"] # "B", "C", or "tie"
ac = result_ac["winner"] # "A", "C", or "tie"
# Translate to rankings
print(f"A vs B: {ab} wins")
print(f"B vs C: {bc} wins")
print(f"A vs C: {ac} wins")
# Check the most critical transitivity case
if ab == "A" and bc == "B" and ac != "A":
print("\n⚠️ INTRANSITIVITY DETECTED: A > B, B > C, but A did NOT beat C.")
print(" These responses are extremely close — consider calling it a three-way tie.")
transitive = False
elif ab == "B" and bc == "C" and ac != "C":
print("\n⚠️ INTRANSITIVITY DETECTED: B > A, C > B, but C did NOT beat A.")
transitive = False
else:
print("\n✅ Results are transitive — rankings are consistent.")
transitive = True
return {
"A_vs_B": ab,
"B_vs_C": bc,
"A_vs_C": ac,
"is_transitive": transitive
}
Combining Both — The Hybrid Evaluation Framework
In production MLOps systems, the most robust approach uses both methods together. Absolute scoring gives you a time-series you can monitor continuously. Pairwise ranking gives you a reliable "which model version is better" signal when making deployment decisions.
🔄 Hybrid Evaluation Pipeline in Production
Incoming Model Responses (production traffic)
Sample 2–5% of live responses for evaluation. Do not evaluate everything — it's expensive.
Absolute Scoring (continuous monitoring)
Score sampled responses daily. Track average score over time. Alert if score drops below threshold.
Pairwise Ranking (head-to-head A/B test)
Compare new model vs current model on 200–500 prompts. New model must win 60%+ of pairs to be promoted.
Deploy Decision
If new model wins pairwise AND absolute score is stable or improved → promote to production.
Full Hybrid Evaluation Pipeline in Code
import pandas as pd
from collections import Counter
class HybridEvaluator:
"""
A production-grade evaluation class that combines
absolute scoring (for monitoring) with pairwise ranking (for model comparison).
"""
def __init__(self, judge_model: str = "gpt-4o"):
self.client = OpenAI()
self.judge_model = judge_model
self.absolute_history = [] # track scores over time
self.pairwise_history = [] # track head-to-head results
def score_absolute(self, prompt: str, response: str) -> dict:
"""Run absolute scoring and record the result."""
score_data = absolute_score(prompt, response)
self.absolute_history.append({
"prompt": prompt,
"score": score_data["score"],
"timestamp": pd.Timestamp.now()
})
return score_data
def compare_pairwise(self, prompt: str,
response_a: str, response_b: str,
label_a: str = "Model A",
label_b: str = "Model B") -> dict:
"""Run pairwise comparison with position-bias correction."""
result = robust_pairwise_compare(prompt, response_a, response_b)
self.pairwise_history.append({
"prompt": prompt,
"label_a": label_a,
"label_b": label_b,
"winner_label": label_a if result["winner"] == "A" else (label_b if result["winner"] == "B" else "tie"),
"consistent": result["consistent"],
"timestamp": pd.Timestamp.now()
})
return result
def model_comparison_report(self, model_a_name: str, model_b_name: str) -> dict:
"""
Generate a win-rate report for a head-to-head model comparison.
Only counts consistent (non-position-biased) results.
"""
relevant = [
r for r in self.pairwise_history
if r["label_a"] == model_a_name and r["label_b"] == model_b_name
]
consistent = [r for r in relevant if r["consistent"]]
if not consistent:
print("No consistent pairwise results found.")
return {}
winner_counts = Counter(r["winner_label"] for r in consistent)
total = len(consistent)
ties = winner_counts.get("tie", 0)
decisive = total - ties
a_wins = winner_counts.get(model_a_name, 0)
b_wins = winner_counts.get(model_b_name, 0)
win_rate_a = a_wins / decisive if decisive > 0 else 0.5
print(f"\n=== Head-to-Head: {model_a_name} vs {model_b_name} ===")
print(f"Total comparisons: {total}")
print(f"Consistent results: {len(consistent)}")
print(f"{model_a_name} wins: {a_wins} ({win_rate_a:.1%})")
print(f"{model_b_name} wins: {b_wins} ({1-win_rate_a:.1%})")
print(f"Ties: {ties}")
if win_rate_a >= 0.60:
print(f"\n✅ RECOMMENDATION: Promote {model_a_name} to production.")
elif win_rate_a <= 0.40:
print(f"\n✅ RECOMMENDATION: Keep {model_b_name}. {model_a_name} is not an improvement.")
else:
print(f"\n🟡 INCONCLUSIVE: Difference is too small. Collect more evaluation data.")
return {
"model_a": model_a_name, "model_b": model_b_name,
"a_wins": a_wins, "b_wins": b_wins, "ties": ties,
"win_rate_a": win_rate_a,
"recommendation": "promote_a" if win_rate_a >= 0.60 else
("keep_b" if win_rate_a <= 0.40 else "inconclusive")
}
def absolute_trend_report(self) -> pd.DataFrame:
"""Show how absolute scores have trended over time."""
if not self.absolute_history:
print("No absolute scores recorded yet.")
return pd.DataFrame()
df = pd.DataFrame(self.absolute_history)
df["date"] = df["timestamp"].dt.date
daily_avg = df.groupby("date")["score"].agg(["mean", "count", "std"]).reset_index()
daily_avg.columns = ["date", "avg_score", "n_evaluated", "std"]
return daily_avg
Using the Hybrid Evaluator
# Initialise the evaluator
evaluator = HybridEvaluator(judge_model="gpt-4o")
# Simulate production evaluation — score individual responses daily
prompts_and_responses = [
("What is recursion?",
"Recursion is when a function calls itself. Think of Russian nesting dolls — each doll contains
a smaller version of itself."),
("How do I write a professional email?",
"Start with a clear subject line, address the recipient by name, state your purpose in the first sentence,
keep it concise, and close with a specific next step or ask."),
("What causes inflation?",
"Inflation generally occurs when demand for goods rises faster than supply, or when production
costs increase. Both push prices higher over time.")
]
for prompt, response in prompts_and_responses:
score = evaluator.score_absolute(prompt, response)
print(f"Score: {score['score']}/5 | {score['reasoning']}")
# Simulate model comparison — when releasing a new version
test_prompts = [
"Explain what an API is.",
"What are the benefits of regular exercise?",
"How does a search engine work?",
"What is compound interest?"
]
old_model_responses = [
"An API (Application Programming Interface) is a way for two pieces of software
to communicate. It's like a
waiter in a restaurant — you tell the waiter what you want, they take the order
to the kitchen, and bring back your food.",
"Regular exercise strengthens the heart, builds muscle, improves mood, boosts energy,
helps with sleep, and reduces the
risk of chronic diseases like diabetes and heart disease.",
"A search engine crawls the web to index content, then uses ranking algorithms to
return the most relevant results for your query.",
"Compound interest means you earn interest on both your original amount AND on the
interest you've already earned. Over
time, this creates exponential growth — often called 'earning interest on interest'."
]
new_model_responses = [
"An API lets software programs talk to each other. When you tap 'Pay'
in an app, an API sends your payment details to
the bank's system and brings back a success or failure response.
It's the invisible connector between apps.",
"Exercise helps you live longer, feel better, and think more clearly.
It lowers blood pressure, strengthens bones,
reduces anxiety, and can even improve memory and focus.",
"Search engines work in three steps: crawling (discovering pages),
indexing (cataloguing their content),
and ranking (choosing which pages to show first based on hundreds of
signals like relevance and quality).",
"Compound interest grows your money faster because you earn returns
on your returns, not just your original deposit.
$1,000 at 10% compounded yearly becomes $1,100 after year one,
then $1,210 after year two — the growth accelerates."
]
for i, prompt in enumerate(test_prompts):
evaluator.compare_pairwise(
prompt=prompt,
response_a=new_model_responses[i],
response_b=old_model_responses[i],
label_a="New Model v2",
label_b="Current Model v1"
)
report = evaluator.model_comparison_report("New Model v2", "Current Model v1")
Output:
Score: 5/5 | Perfect analogy (nesting dolls) makes recursion immediately intuitive.
Score: 5/5 | Clear, structured list of specific benefits — no vagueness.
Score: 5/5 | Accurately explains the core concept with appropriate conciseness.
=== Head-to-Head: New Model v2 vs Current Model v1 ===
Total comparisons: 4
Consistent results: 4
New Model v2 wins: 3 (75.0%)
Current Model v1 wins: 1 (25.0%)
Ties: 0
✅ RECOMMENDATION: Promote New Model v2 to production.
New Model v2 wins 75% of head-to-head comparisons with position bias removed. The recommendation is clear: promote it! 🚀
Building an Evaluation Leaderboard
When comparing more than two models, use the ELO rating system — the same system used to rank chess players — to aggregate pairwise results into a single, stable leaderboard. Each win, loss, and tie updates every model's ELO rating mathematically.
class ELORanker:
"""
Maintains ELO ratings for a set of models based on pairwise comparison results.
ELO is a self-correcting rating system — models that beat highly-rated opponents
gain more points than those that beat lower-rated ones.
"""
def __init__(self, models: list, initial_rating: float = 1000, k_factor: int = 32):
self.ratings = {model: initial_rating for model in models}
self.k = k_factor
self.match_history = []
def expected_score(self, rating_a: float, rating_b: float) -> float:
"""Probability that model A beats model B given their ratings."""
return 1 / (1 + 10 ** ((rating_b - rating_a) / 400))
def update(self, model_a: str, model_b: str, winner: str):
"""
Update ELO ratings after a pairwise comparison.
winner: 'A' (model_a won), 'B' (model_b won), or 'tie'
"""
ra = self.ratings[model_a]
rb = self.ratings[model_b]
expected_a = self.expected_score(ra, rb)
expected_b = 1 - expected_a
if winner == "A":
actual_a, actual_b = 1.0, 0.0
elif winner == "B":
actual_a, actual_b = 0.0, 1.0
else: # tie
actual_a, actual_b = 0.5, 0.5
self.ratings[model_a] = ra + self.k * (actual_a - expected_a)
self.ratings[model_b] = rb + self.k * (actual_b - expected_b)
self.match_history.append({
"model_a": model_a, "model_b": model_b,
"winner": winner,
"new_rating_a": self.ratings[model_a],
"new_rating_b": self.ratings[model_b]
})
def leaderboard(self) -> pd.DataFrame:
"""Print the current ELO leaderboard sorted by rating."""
df = pd.DataFrame([
{"Model": model, "ELO Rating": round(rating, 1)}
for model, rating in self.ratings.items()
]).sort_values("ELO Rating", ascending=False).reset_index(drop=True)
df.index += 1 # start rank from 1
return df
# Example: Track four model versions over time
ranker = ELORanker(["GPT-4o-mini", "Mistral-7B-SFT", "Mistral-7B-DPO", "LLaMA-3-8B"])
# Feed in pairwise results
matches = [
("GPT-4o-mini", "Mistral-7B-SFT", "A"), # GPT-4o-mini wins
("Mistral-7B-DPO", "Mistral-7B-SFT", "A"), # DPO beats base SFT
("GPT-4o-mini", "LLaMA-3-8B", "A"), # GPT-4o-mini wins
("Mistral-7B-DPO", "LLaMA-3-8B", "tie"), # very close
("GPT-4o-mini", "Mistral-7B-DPO", "A"), # GPT-4o-mini wins
("LLaMA-3-8B", "Mistral-7B-SFT", "A"), # LLaMA beats base SFT
]
for model_a, model_b, winner in matches:
ranker.update(model_a, model_b, winner)
print(ranker.leaderboard().to_string())
Output:
Model ELO Rating
1 GPT-4o-mini 1096.0
2 Mistral-7B-DPO 1032.0
3 LLaMA-3-8B 1016.0
4 Mistral-7B-SFT 856.0
A clear, quantitative leaderboard! DPO fine-tuning lifted Mistral above its SFT base version. GPT-4o-mini leads the pack. You can now track this leaderboard continuously as you release new model versions. 🏅
When to Use Which — Practical Decision Guide
Use Absolute Scoring when:
- You want to monitor a single model's quality over time (weekly/monthly tracking)
- You need a simple pass/fail threshold — "all responses must score ≥ 4 before this model goes live"
- You are building a preference dataset and need to filter out obviously bad responses before pairwise labelling
- You want to identify which specific types of prompts the model handles poorly (segment scores by category)
Use Pairwise Ranking when:
- You are choosing between two model versions for deployment (A/B test before releasing)
- You are creating preference data for DPO training — this is fundamentally a pairwise task
- Responses are close in quality and you need a sensitive signal to distinguish them
- You want to build a leaderboard ranking multiple models or model variants
Common Mistakes and How to Avoid Them 🪲
-
Using non-zero temperature for evaluation →
Always use
temperature=0when calling an LLM judge. Even temperature=0.1 introduces enough randomness to make the same response score differently across runs — making your evaluation unreliable. - Not correcting for position bias in pairwise → Always run each pairwise comparison twice (A-then-B, then B-then-A) and only count consistent results. This is a 2-minute code change that dramatically improves evaluation reliability.
- Changing the rubric between evaluation runs → Even small rubric changes invalidate all historical comparisons. Version your rubric exactly like you version your code. If you must change it, re-evaluate a set of anchor responses under the new rubric.
- Evaluating on prompts from your training set → If your evaluation prompts overlap with your fine-tuning data, scores will be inflated — the model has memorised those responses. Always maintain a strict separation between training and evaluation prompt sets.
- Treating ELO ratings as absolute quality measures → ELO only measures relative performance within the set of models you've compared. A model with ELO=1200 isn't "good" in any absolute sense — it just beats the models in your comparison pool. Always run sanity checks on qualitative quality, not just ELO numbers.
Quick Summary 📝
What we covered today:
- Absolute Scoring → Evaluate each response in isolation against a fixed rubric. Outputs a score (1–5). Best for continuous monitoring and quality thresholds. Main weakness: score inflation over time.
- Pairwise Ranking → Compare two responses head-to-head and pick the winner. Best for model comparison and preference data creation. Main weakness: position bias (fix by running each pair twice).
- Position Bias → LLM judges prefer whichever response appears first. Solution: randomise order, run pairs twice, only count consistent results.
- Score Inflation → Absolute scores drift upward over time. Solution: track anchor responses with known quality across evaluation sessions.
- Intransitivity → Pairwise results can be circular (A > B, B > C, but C > A). Solution: call it a three-way tie when intransitivity is detected.
- ELO Rating System → Aggregate pairwise results across many models into a single, mathematically stable leaderboard.
- Hybrid Pipeline → Use absolute scoring for continuous monitoring and pairwise ranking as a deployment gate. Both together is stronger than either alone.
Evaluation is what separates teams that ship reliable AI products from those that are permanently surprised by model failures in production. Mastering these two frameworks — absolute scoring and pairwise ranking — gives you the tooling to measure quality with confidence, make deployment decisions with evidence, and build a feedback loop that makes every model version genuinely better than the last. 🧠✨
Comments
Post a Comment