Skip to main content

Precision, Recall & F1 Score: Essential Metrics for Model Evaluation

Calculating read time…

Imagine you built an AI model that detects whether a mango is ripe or not. 🥭
You're excited. You test it. It says "95% Accuracy!"
You celebrate. But wait — the model said every single mango is ripe. Including the green ones.

This is exactly why Accuracy alone is not enough to judge an AI model.
You need smarter tools — Precision, Recall, and F1 Score.




What is Model Evaluation?

When you build an AI model, it learns from data and makes predictions.
But how do you know if those predictions are good or bad?

Model Evaluation = the process of testing your AI model to see how well it performs.
Think of it like giving your model a report card 📝 after school.

💡 Simple Analogy:
Imagine you trained a dog to fetch only red balls from a pile of mixed coloured balls.
After training, you test the dog. Does it bring only red ones? Does it miss any red ones?
That testing is model evaluation. The dog is your AI model. 🐶

📦 Step 1: The Confusion Matrix — The Magic Box

Before learning Precision, Recall, and F1, you need to understand one very important tool:
the Confusion Matrix.

Don't worry — it's not confusing at all! Let's use a simple story. 🏥

🏥 The Story: AI Doctor Finding COVID Patients

An AI model looks at 100 people and decides: "Does this person have COVID or not?"
Let's say 40 people actually have COVID and 60 people are healthy.

The AI model makes predictions. Some are correct. Some are wrong.
The confusion matrix organises all results into a 2×2 table:

— AI Says: COVID ✅ AI Says: Healthy ❌
Actually: COVID 🤒 ✅ True Positive (TP) = 35 ❌ False Negative (FN) = 5
Actually: Healthy 😊 ❌ False Positive (FP) = 10 ✅ True Negative (TN) = 50

🔑 Understanding Each Cell — In Simple Words:

  • True Positive (TP) = 35 → AI said "COVID", and the person actually HAS COVID. ✅ Correct!
  • True Negative (TN) = 50 → AI said "Healthy", and the person is actually Healthy. ✅ Correct!
  • False Positive (FP) = 10 → AI said "COVID", but the person is actually Healthy. ❌ Wrong! (Alarm without fire 🚒)
  • False Negative (FN) = 5 → AI said "Healthy", but the person actually HAS COVID. ❌ Very dangerous! (Missed the sick person 😟)
🚨 Remember This!
False Negative is the most dangerous error in medical AI.
It means you told a sick person "You're fine!" — and they never got treated.
This is why accuracy alone can be misleading.

🎯 Step 2: What is Precision?

Precision answers this question:
"When my AI model says YES — how often is it actually correct?"

📚 Real-Life Example First:

Imagine you are a treasure hunter. 🏴‍☠️
You dig 10 holes. You find treasure in 7 of them.
Your Precision = 7 out of 10 = 70%.

High precision = when you say "I found treasure here!" — you are usually right.
You don't waste effort digging empty holes.

🧮 The Precision Formula:

Precision = TP ÷ (TP + FP)

From our COVID example:
Precision = 35 ÷ (35 + 10) = 35 ÷ 45 = 0.78 or 78%

This means: when the AI says "this person has COVID", it is right 78% of the time.
22% of the time, it raised a false alarm (scared a healthy person needlessly 😬).

✅ When should you care more about Precision?
Use Precision when False Positives are costly.
Example: Spam Email Filter 📧 — you don't want real emails accidentally going to spam.
Example: Fraud Detection 💳 — you don't want to block a genuine customer's card.

📡 Step 3: What is Recall?

Recall answers this question:
"Out of ALL the actual YES cases — how many did my AI find?"

📚 Real-Life Example First:

Now imagine there are 40 gold coins hidden in a field. 🪙
You search and find 35 of them.
Your Recall = 35 out of 40 = 87.5%.

High recall = you don't miss many gold coins.
Even if you sometimes dig empty holes (false alarms), you don't leave gold behind.

🧮 The Recall Formula:

Recall = TP ÷ (TP + FN)

From our COVID example:
Recall = 35 ÷ (35 + 5) = 35 ÷ 40 = 0.875 or 87.5%

This means: out of the 40 people who actually had COVID,
the AI correctly identified 35 of them (87.5%).
It missed 5 sick people (those are the dangerous False Negatives!).

✅ When should you care more about Recall?
Use Recall when False Negatives are costly — when missing a case is dangerous.
Example: Cancer Detection 🏥 — don't miss a cancer patient!
Example: Airport Security 🛫 — don't miss a dangerous item in bags!

⚖️ Step 4: The Precision vs Recall Trade-off

Here is something very important that even experienced people get confused about:
Precision and Recall often fight each other. 🥊

When you try to make Precision go up → Recall usually comes down.
When you try to make Recall go up → Precision usually comes down.

💡 Think of it like a fishing net:
If you use a very fine net (catch only big fish) → Precision is high, but you miss many small fish (low Recall).
If you use a very wide net (catch everything) → Recall is high, but you also catch garbage (low Precision). 🎣

So… which one should you use? That depends on your problem.
But wait — there is a smarter way! Enter the F1 Score. 🎉


🏆 Step 5: What is the F1 Score?

The F1 Score is the balanced referee between Precision and Recall. 🧑‍⚖️
It gives you a single number that considers BOTH at once.

Instead of a regular average, F1 uses a special kind called the Harmonic Mean.
(Don't worry, you don't need to memorise this term — just understand what it does!)

💡 Why not just use regular average?
If Precision = 100% and Recall = 0%, regular average = 50% (looks okay!).
But F1 Score = 0% (correctly shows this model is useless!).
The Harmonic Mean punishes extreme imbalance — that's its superpower. 💪

🧮 The F1 Score Formula:

F1 Score = 2 × (Precision × Recall) ÷ (Precision + Recall)

From our COVID example:
Precision = 0.78, Recall = 0.875
F1 = 2 × (0.78 × 0.875) ÷ (0.78 + 0.875)
F1 = 2 × 0.6825 ÷ 1.655
F1 = 0.825 or 82.5%

An F1 of 82.5% means our model is doing quite well overall.
It's a great single number to compare different models against each other!


📊 Quick Visual Comparison: All 3 Metrics Side by Side

Metric Question it Answers Formula Worry About
Precision 🎯 When I say YES, am I right? TP ÷ (TP + FP) False Positives (FP)
Recall 📡 Did I find all real YES cases? TP ÷ (TP + FN) False Negatives (FN)
F1 Score ⚖️ Am I balanced between both? 2×(P×R)÷(P+R) Both FP and FN

🐍 Step 6: Python Code — Hands-On with OCI Data Science

Now let's bring this to life with actual Python code!
We'll use Oracle Cloud Infrastructure (OCI) Data Science and the popular
scikit-learn library — which is already pre-installed in OCI Notebooks. 🎉

🛠️ Where to run this code?
1. Log in to your OCI Console → Data Science → Create Project → Create Notebook Session.
2. Open a new Jupyter Notebook inside your session.
3. Copy-paste each code block below, one step at a time. That's it! ✅

📌 Code Block 1 — Import the Tools

📝 What will this code do?
This code brings in the libraries (tools) we need.
Think of it like taking your pencil box out before a math exam — you gather your tools first!
# Import the tools we need
from sklearn.metrics import (
    precision_score,    # To calculate Precision
    recall_score,       # To calculate Recall
    f1_score,           # To calculate F1 Score
    confusion_matrix,   # To create the confusion matrix
    classification_report  # Gives a full report card!
)
import numpy as np
import pandas as pd

📌 Code Block 2 — Create Sample Predictions

📝 What will this code do?
Here we are pretending to be a doctor.
y_true = the real answers (who actually has the disease — like the answer key to a test).
y_pred = what our AI model predicted (the model's guesses).
1 = has disease, 0 = healthy.
# y_true = the ACTUAL correct answers (ground truth)
# Think of this as: who ACTUALLY has the disease
y_true = [1, 0, 1, 1, 0, 1, 0, 0, 1, 1,
          0, 1, 0, 1, 0, 0, 1, 1, 0, 1]

# y_pred = what our AI MODEL predicted
# Some right, some wrong - just like a student's exam answers
y_pred = [1, 0, 1, 0, 0, 1, 1, 0, 1, 1,
          0, 1, 0, 0, 0, 1, 1, 0, 0, 1]

print("Total Patients:", len(y_true))
print("Actually Sick (1s) :", sum(y_true))
print("Actually Healthy (0s):", y_true.count(0))

Output:

Total Patients: 20
Actually Sick (1s) : 12
Actually Healthy (0s): 8

📌 Code Block 3 — Calculate Precision, Recall and F1

📝 What will this code do?
This is where the magic happens! 🎩
We feed our actual answers and AI predictions into scikit-learn's functions.
It automatically calculates Precision, Recall and F1 Score for us.
We're basically asking the library: "How good is our AI doctor?"
# Calculate Precision
precision = precision_score(y_true, y_pred)

# Calculate Recall
recall = recall_score(y_true, y_pred)

# Calculate F1 Score
f1 = f1_score(y_true, y_pred)

# Print the results nicely
print("=" * 40)
print("  📊 MODEL EVALUATION RESULTS")
print("=" * 40)
print(f"  🎯 Precision : {precision:.4f}  ({precision*100:.2f}%)")
print(f"  📡 Recall    : {recall:.4f}  ({recall*100:.2f}%)")
print(f"  ⚖️  F1 Score  : {f1:.4f}  ({f1*100:.2f}%)")
print("=" * 40)

Output:

========================================
  📊 MODEL EVALUATION RESULTS
========================================
  🎯 Precision : 0.9091  (90.91%)
  📡 Recall    : 0.8333  (83.33%)
  ⚖️  F1 Score  : 0.8696  (86.96%)
========================================

Our AI model has a Precision of ~91% (very few false alarms 🎉)
and a Recall of ~83% (found most sick patients ✅).
The F1 Score of ~87% tells us the model is well balanced overall!

📌 Code Block 4 — Print the Full Confusion Matrix

📝 What will this code do?
This prints the full confusion matrix — the 2×2 magic box we talked about earlier!
It shows exactly how many True Positives, False Positives, False Negatives and True Negatives there are.
It's like a detailed score sheet of every prediction our model made.
# Create the confusion matrix
cm = confusion_matrix(y_true, y_pred)

print("\n📦 CONFUSION MATRIX:")
print("-" * 35)
print(f"  True Negative  (TN) = {cm[0][0]}  ✅")
print(f"  False Positive (FP) = {cm[0][1]}  ❌")
print(f"  False Negative (FN) = {cm[1][0]}  ❌")
print(f"  True Positive  (TP) = {cm[1][1]}  ✅")
print("-" * 35)

Output:

📦 CONFUSION MATRIX:
-----------------------------------
  True Negative  (TN) = 6   ✅
  False Positive (FP) = 1   ❌
  False Negative (FN) = 2   ❌
  True Positive  (TP) = 10  ✅
-----------------------------------

📌 Code Block 5 — The Full Classification Report (One Command!)

📝 What will this code do?
This is like getting your full report card in one shot! 🏆
The classification_report function automatically calculates all metrics
for every class (healthy vs. sick) and also gives you averages.
This is one of the most-used functions in real AI projects.
# Full report — like a detailed school report card!
report = classification_report(
    y_true,
    y_pred,
    target_names=["Healthy (0)", "Sick (1)"]
)

print("\n📋 FULL CLASSIFICATION REPORT:")
print(report)

Output:

📋 FULL CLASSIFICATION REPORT:
              precision    recall  f1-score   support

 Healthy (0)       0.75      0.88      0.81         8
    Sick (1)       0.91      0.83      0.87        12

    accuracy                           0.85        20
   macro avg       0.83      0.85      0.84        20
weighted avg       0.85      0.85      0.85        20

The report shows metrics separately for each class (Healthy and Sick) — very useful when one class has more samples than another!


🌍 Real-World Use Cases: When to Use What?

Use Case What Matters More? Why?
🏥 Cancer Diagnosis Recall Missing a cancer patient (FN) is life-threatening
📧 Email Spam Filter Precision Sending real emails to spam (FP) is very annoying
💳 Bank Fraud Detection Both (F1) Missing fraud (FN) and blocking real customers (FP) both hurt
🛫 Airport Security Scan Recall Missing a dangerous item (FN) is catastrophic
🔎 Search Engine Results Precision Showing irrelevant results (FP) frustrates users

🧩 Step 7: Advanced — Macro vs Weighted F1 (For Multi-Class Problems)

Everything we covered above was for binary classification — two classes (Sick vs Healthy).
But what if your model has 3 or more classes? Like: Cat 🐱 / Dog 🐶 / Bird 🐦?

In that case, you can compute F1 in three ways:

  • Micro F1 — Counts all TP, FP, FN across all classes together, then calculates. Good when you want overall performance.
  • Macro F1 — Calculates F1 for each class separately, then takes a simple average. Treats all classes equally — good when classes are equally important.
  • Weighted F1 — Same as Macro, but weighs each class by its size (support). Good when classes have imbalanced samples (e.g. 900 cats, 50 dogs, 50 birds).
📝 What will this code do?
This code creates a 3-class classification problem (Animal Type) and shows you
how to compute Macro and Weighted F1 easily with one line of code each.
# Multi-class Example: Animal classifier (Cat=0, Dog=1, Bird=2)
y_true_multi = [0, 1, 2, 0, 1, 2, 0, 0, 1, 2,
                1, 0, 2, 1, 0, 2, 1, 0, 2, 1]

y_pred_multi = [0, 1, 2, 0, 0, 2, 0, 1, 1, 2,
                1, 0, 1, 1, 0, 2, 2, 0, 2, 1]

# Macro F1 — treats all classes equally
macro_f1 = f1_score(y_true_multi, y_pred_multi, average='macro')

# Weighted F1 — accounts for class sizes
weighted_f1 = f1_score(y_true_multi, y_pred_multi, average='weighted')

print(f"🐾 Macro F1 Score    : {macro_f1:.4f}")
print(f"🐾 Weighted F1 Score : {weighted_f1:.4f}")

Output:

🐾 Macro F1 Score    : 0.8095
🐾 Weighted F1 Score : 0.8150

✅ DO's and DON'Ts — Super Important!

✅ DO's:
  • Always check all three metrics together — never rely on accuracy alone.
  • Choose Recall when missing a positive case is costly or dangerous (medical, security).
  • Choose Precision when raising a false alarm is costly or annoying (spam filter, ads).
  • Use F1 Score when your dataset is imbalanced (more of one class than another).
  • Always look at the confusion matrix — numbers tell the whole story.
  • In OCI, use classification_report to get all metrics in one readable output.
❌ DON'Ts:
  • Don't rely on accuracy alone — especially with imbalanced data.
  • Don't ignore False Negatives in critical use cases like disease detection.
  • Don't assume a high F1 Score means a perfect model — always validate on new data!
  • Don't forget to check F1 separately for each class in multi-class problems.
  • Don't confuse Precision and Recall — they answer completely different questions.

🔁 Quick Revision

  • 🎯 Precision = "I only point at things I'm sure about."
    → TP ÷ (TP + FP). Cares about False Positives.
  • 📡 Recall = "I try to find EVERY positive case — don't want to miss any."
    → TP ÷ (TP + FN). Cares about False Negatives.
  • ⚖️ F1 Score = "I'm the balanced scorecard between Precision and Recall."
    → 2 × (P × R) ÷ (P + R). Best for imbalanced datasets.
  • 📦 Confusion Matrix = The full picture of where your model is right and wrong.

Happy Learning! ✨

Comments