Imagine you built an AI model that detects whether a mango is ripe or not. 🥭
You're excited. You test it. It says "95% Accuracy!"
You celebrate. But wait — the model said every single mango is ripe. Including the green ones.
This is exactly why Accuracy alone is not enough to judge an AI model.
You need smarter tools — Precision, Recall, and F1 Score.
What is Model Evaluation?
When you build an AI model, it learns from data and makes predictions.
But how do you know if those predictions are good or bad?
Model Evaluation = the process of testing your AI model to see how well it performs.
Think of it like giving your model a report card 📝 after school.
Imagine you trained a dog to fetch only red balls from a pile of mixed coloured balls.
After training, you test the dog. Does it bring only red ones? Does it miss any red ones?
That testing is model evaluation. The dog is your AI model. 🐶
📦 Step 1: The Confusion Matrix — The Magic Box
Before learning Precision, Recall, and F1, you need to understand one very important tool:
the Confusion Matrix.
Don't worry — it's not confusing at all! Let's use a simple story. 🏥
🏥 The Story: AI Doctor Finding COVID Patients
An AI model looks at 100 people and decides: "Does this person have COVID or not?"
Let's say 40 people actually have COVID and 60 people are healthy.
The AI model makes predictions. Some are correct. Some are wrong.
The confusion matrix organises all results into a 2×2 table:
| — | AI Says: COVID ✅ | AI Says: Healthy ❌ |
|---|---|---|
| Actually: COVID 🤒 | ✅ True Positive (TP) = 35 | ❌ False Negative (FN) = 5 |
| Actually: Healthy 😊 | ❌ False Positive (FP) = 10 | ✅ True Negative (TN) = 50 |
🔑 Understanding Each Cell — In Simple Words:
- True Positive (TP) = 35 → AI said "COVID", and the person actually HAS COVID. ✅ Correct!
- True Negative (TN) = 50 → AI said "Healthy", and the person is actually Healthy. ✅ Correct!
- False Positive (FP) = 10 → AI said "COVID", but the person is actually Healthy. ❌ Wrong! (Alarm without fire 🚒)
- False Negative (FN) = 5 → AI said "Healthy", but the person actually HAS COVID. ❌ Very dangerous! (Missed the sick person 😟)
False Negative is the most dangerous error in medical AI.
It means you told a sick person "You're fine!" — and they never got treated.
This is why accuracy alone can be misleading.
🎯 Step 2: What is Precision?
Precision answers this question:
"When my AI model says YES — how often is it actually correct?"
📚 Real-Life Example First:
Imagine you are a treasure hunter. 🏴☠️
You dig 10 holes. You find treasure in 7 of them.
Your Precision = 7 out of 10 = 70%.
High precision = when you say "I found treasure here!" — you are usually right.
You don't waste effort digging empty holes.
🧮 The Precision Formula:
From our COVID example:
Precision = 35 ÷ (35 + 10) = 35 ÷ 45 = 0.78 or 78%
This means: when the AI says "this person has COVID", it is right 78% of the time.
22% of the time, it raised a false alarm (scared a healthy person needlessly 😬).
Use Precision when False Positives are costly.
Example: Spam Email Filter 📧 — you don't want real emails accidentally going to spam.
Example: Fraud Detection 💳 — you don't want to block a genuine customer's card.
📡 Step 3: What is Recall?
Recall answers this question:
"Out of ALL the actual YES cases — how many did my AI find?"
📚 Real-Life Example First:
Now imagine there are 40 gold coins hidden in a field. 🪙
You search and find 35 of them.
Your Recall = 35 out of 40 = 87.5%.
High recall = you don't miss many gold coins.
Even if you sometimes dig empty holes (false alarms), you don't leave gold behind.
🧮 The Recall Formula:
From our COVID example:
Recall = 35 ÷ (35 + 5) = 35 ÷ 40 = 0.875 or 87.5%
This means: out of the 40 people who actually had COVID,
the AI correctly identified 35 of them (87.5%).
It missed 5 sick people (those are the dangerous False Negatives!).
Use Recall when False Negatives are costly — when missing a case is dangerous.
Example: Cancer Detection 🏥 — don't miss a cancer patient!
Example: Airport Security 🛫 — don't miss a dangerous item in bags!
⚖️ Step 4: The Precision vs Recall Trade-off
Here is something very important that even experienced people get confused about:
Precision and Recall often fight each other. 🥊
When you try to make Precision go up → Recall usually comes down.
When you try to make Recall go up → Precision usually comes down.
If you use a very fine net (catch only big fish) → Precision is high, but you miss many small fish (low Recall).
If you use a very wide net (catch everything) → Recall is high, but you also catch garbage (low Precision). 🎣
So… which one should you use? That depends on your problem.
But wait — there is a smarter way! Enter the F1 Score. 🎉
🏆 Step 5: What is the F1 Score?
The F1 Score is the balanced referee between Precision and Recall. 🧑⚖️
It gives you a single number that considers BOTH at once.
Instead of a regular average, F1 uses a special kind called the Harmonic Mean.
(Don't worry, you don't need to memorise this term — just understand what it does!)
If Precision = 100% and Recall = 0%, regular average = 50% (looks okay!).
But F1 Score = 0% (correctly shows this model is useless!).
The Harmonic Mean punishes extreme imbalance — that's its superpower. 💪
🧮 The F1 Score Formula:
From our COVID example:
Precision = 0.78, Recall = 0.875
F1 = 2 × (0.78 × 0.875) ÷ (0.78 + 0.875)
F1 = 2 × 0.6825 ÷ 1.655
F1 = 0.825 or 82.5%
An F1 of 82.5% means our model is doing quite well overall.
It's a great single number to compare different models against each other!
📊 Quick Visual Comparison: All 3 Metrics Side by Side
| Metric | Question it Answers | Formula | Worry About |
|---|---|---|---|
| Precision 🎯 | When I say YES, am I right? | TP ÷ (TP + FP) | False Positives (FP) |
| Recall 📡 | Did I find all real YES cases? | TP ÷ (TP + FN) | False Negatives (FN) |
| F1 Score ⚖️ | Am I balanced between both? | 2×(P×R)÷(P+R) | Both FP and FN |
🐍 Step 6: Python Code — Hands-On with OCI Data Science
Now let's bring this to life with actual Python code!
We'll use Oracle Cloud Infrastructure (OCI) Data Science and the popular
scikit-learn library — which is already pre-installed in OCI Notebooks. 🎉
1. Log in to your OCI Console → Data Science → Create Project → Create Notebook Session.
2. Open a new Jupyter Notebook inside your session.
3. Copy-paste each code block below, one step at a time. That's it! ✅
📌 Code Block 1 — Import the Tools
This code brings in the libraries (tools) we need.
Think of it like taking your pencil box out before a math exam — you gather your tools first!
# Import the tools we need
from sklearn.metrics import (
precision_score, # To calculate Precision
recall_score, # To calculate Recall
f1_score, # To calculate F1 Score
confusion_matrix, # To create the confusion matrix
classification_report # Gives a full report card!
)
import numpy as np
import pandas as pd
📌 Code Block 2 — Create Sample Predictions
Here we are pretending to be a doctor.
y_true = the real answers (who actually has the disease — like the answer key to a test).y_pred = what our AI model predicted (the model's guesses).1 = has disease, 0 = healthy.
# y_true = the ACTUAL correct answers (ground truth)
# Think of this as: who ACTUALLY has the disease
y_true = [1, 0, 1, 1, 0, 1, 0, 0, 1, 1,
0, 1, 0, 1, 0, 0, 1, 1, 0, 1]
# y_pred = what our AI MODEL predicted
# Some right, some wrong - just like a student's exam answers
y_pred = [1, 0, 1, 0, 0, 1, 1, 0, 1, 1,
0, 1, 0, 0, 0, 1, 1, 0, 0, 1]
print("Total Patients:", len(y_true))
print("Actually Sick (1s) :", sum(y_true))
print("Actually Healthy (0s):", y_true.count(0))
Output:
Total Patients: 20 Actually Sick (1s) : 12 Actually Healthy (0s): 8
📌 Code Block 3 — Calculate Precision, Recall and F1
This is where the magic happens! 🎩
We feed our actual answers and AI predictions into scikit-learn's functions.
It automatically calculates Precision, Recall and F1 Score for us.
We're basically asking the library: "How good is our AI doctor?"
# Calculate Precision
precision = precision_score(y_true, y_pred)
# Calculate Recall
recall = recall_score(y_true, y_pred)
# Calculate F1 Score
f1 = f1_score(y_true, y_pred)
# Print the results nicely
print("=" * 40)
print(" 📊 MODEL EVALUATION RESULTS")
print("=" * 40)
print(f" 🎯 Precision : {precision:.4f} ({precision*100:.2f}%)")
print(f" 📡 Recall : {recall:.4f} ({recall*100:.2f}%)")
print(f" ⚖️ F1 Score : {f1:.4f} ({f1*100:.2f}%)")
print("=" * 40)
Output:
======================================== 📊 MODEL EVALUATION RESULTS ======================================== 🎯 Precision : 0.9091 (90.91%) 📡 Recall : 0.8333 (83.33%) ⚖️ F1 Score : 0.8696 (86.96%) ========================================
Our AI model has a Precision of ~91% (very few false alarms 🎉)
and a Recall of ~83% (found most sick patients ✅).
The F1 Score of ~87% tells us the model is well balanced overall!
📌 Code Block 4 — Print the Full Confusion Matrix
This prints the full confusion matrix — the 2×2 magic box we talked about earlier!
It shows exactly how many True Positives, False Positives, False Negatives and True Negatives there are.
It's like a detailed score sheet of every prediction our model made.
# Create the confusion matrix
cm = confusion_matrix(y_true, y_pred)
print("\n📦 CONFUSION MATRIX:")
print("-" * 35)
print(f" True Negative (TN) = {cm[0][0]} ✅")
print(f" False Positive (FP) = {cm[0][1]} ❌")
print(f" False Negative (FN) = {cm[1][0]} ❌")
print(f" True Positive (TP) = {cm[1][1]} ✅")
print("-" * 35)
Output:
📦 CONFUSION MATRIX: ----------------------------------- True Negative (TN) = 6 ✅ False Positive (FP) = 1 ❌ False Negative (FN) = 2 ❌ True Positive (TP) = 10 ✅ -----------------------------------
📌 Code Block 5 — The Full Classification Report (One Command!)
This is like getting your full report card in one shot! 🏆
The
classification_report function automatically calculates all metricsfor every class (healthy vs. sick) and also gives you averages.
This is one of the most-used functions in real AI projects.
# Full report — like a detailed school report card!
report = classification_report(
y_true,
y_pred,
target_names=["Healthy (0)", "Sick (1)"]
)
print("\n📋 FULL CLASSIFICATION REPORT:")
print(report)
Output:
📋 FULL CLASSIFICATION REPORT:
precision recall f1-score support
Healthy (0) 0.75 0.88 0.81 8
Sick (1) 0.91 0.83 0.87 12
accuracy 0.85 20
macro avg 0.83 0.85 0.84 20
weighted avg 0.85 0.85 0.85 20
The report shows metrics separately for each class (Healthy and Sick) — very useful when one class has more samples than another!
🌍 Real-World Use Cases: When to Use What?
| Use Case | What Matters More? | Why? |
|---|---|---|
| 🏥 Cancer Diagnosis | Recall | Missing a cancer patient (FN) is life-threatening |
| 📧 Email Spam Filter | Precision | Sending real emails to spam (FP) is very annoying |
| 💳 Bank Fraud Detection | Both (F1) | Missing fraud (FN) and blocking real customers (FP) both hurt |
| 🛫 Airport Security Scan | Recall | Missing a dangerous item (FN) is catastrophic |
| 🔎 Search Engine Results | Precision | Showing irrelevant results (FP) frustrates users |
🧩 Step 7: Advanced — Macro vs Weighted F1 (For Multi-Class Problems)
Everything we covered above was for binary classification — two classes (Sick vs Healthy).
But what if your model has 3 or more classes? Like: Cat 🐱 / Dog 🐶 / Bird 🐦?
In that case, you can compute F1 in three ways:
- Micro F1 — Counts all TP, FP, FN across all classes together, then calculates. Good when you want overall performance.
- Macro F1 — Calculates F1 for each class separately, then takes a simple average. Treats all classes equally — good when classes are equally important.
- Weighted F1 — Same as Macro, but weighs each class by its size (support). Good when classes have imbalanced samples (e.g. 900 cats, 50 dogs, 50 birds).
This code creates a 3-class classification problem (Animal Type) and shows you
how to compute Macro and Weighted F1 easily with one line of code each.
# Multi-class Example: Animal classifier (Cat=0, Dog=1, Bird=2)
y_true_multi = [0, 1, 2, 0, 1, 2, 0, 0, 1, 2,
1, 0, 2, 1, 0, 2, 1, 0, 2, 1]
y_pred_multi = [0, 1, 2, 0, 0, 2, 0, 1, 1, 2,
1, 0, 1, 1, 0, 2, 2, 0, 2, 1]
# Macro F1 — treats all classes equally
macro_f1 = f1_score(y_true_multi, y_pred_multi, average='macro')
# Weighted F1 — accounts for class sizes
weighted_f1 = f1_score(y_true_multi, y_pred_multi, average='weighted')
print(f"🐾 Macro F1 Score : {macro_f1:.4f}")
print(f"🐾 Weighted F1 Score : {weighted_f1:.4f}")
Output:
🐾 Macro F1 Score : 0.8095 🐾 Weighted F1 Score : 0.8150
✅ DO's and DON'Ts — Super Important!
- Always check all three metrics together — never rely on accuracy alone.
- Choose Recall when missing a positive case is costly or dangerous (medical, security).
- Choose Precision when raising a false alarm is costly or annoying (spam filter, ads).
- Use F1 Score when your dataset is imbalanced (more of one class than another).
- Always look at the confusion matrix — numbers tell the whole story.
- In OCI, use classification_report to get all metrics in one readable output.
- Don't rely on accuracy alone — especially with imbalanced data.
- Don't ignore False Negatives in critical use cases like disease detection.
- Don't assume a high F1 Score means a perfect model — always validate on new data!
- Don't forget to check F1 separately for each class in multi-class problems.
- Don't confuse Precision and Recall — they answer completely different questions.
🔁 Quick Revision
-
🎯 Precision = "I only point at things I'm sure about."
→ TP ÷ (TP + FP). Cares about False Positives. -
📡 Recall = "I try to find EVERY positive case — don't want to miss any."
→ TP ÷ (TP + FN). Cares about False Negatives. -
⚖️ F1 Score = "I'm the balanced scorecard between Precision and Recall."
→ 2 × (P × R) ÷ (P + R). Best for imbalanced datasets. - 📦 Confusion Matrix = The full picture of where your model is right and wrong.
Happy Learning! ✨
Comments
Post a Comment