Skip to main content

How to Evaluate Deep Learning Models: Accuracy, Precision, Recall, F1, and Confusion Matrix

Calculating read time…

Evaluating a classification model means counting, in a structured way, exactly which predictions were right and which were wrong — and then computing accuracy, precision, recall, and F1 from those counts to answer different questions about the model's behavior, since no single number tells the whole story. The confusion matrix is the shared source all of these metrics are computed from. 🧮

A model that reports 99% accuracy can be either genuinely excellent or completely useless, depending entirely on what it's being evaluated against — and the only way to tell the difference is to look past accuracy into precision, recall, and the confusion matrix underneath them. This post builds each metric up from the raw counts, shows the documented definitions frameworks actually use, and covers the mistakes that make evaluation numbers quietly misleading. 🔍

A two-by-two confusion matrix grid with true positive, false negative, false positive, and true negative cells, with arrows leading to three formula boxes on the right for precision, recall, and F1 score, showing which cells feed each formula

🔀 Quick Comparison: Precision vs. Recall

Property Precision Recall
Formula TP / (TP + FP) TP / (TP + FN)
Question answered Of everything flagged positive, how much was correct? Of everything actually positive, how much did we catch?
Punishes False positives False negatives
Prioritize when A false alarm is costly (e.g., flagging a legitimate email as spam) A missed case is costly (e.g., missing a disease diagnosis)

1. Foundations: Why Accuracy Alone Can Lie

💭 Analogy first: imagine a smoke detector that never goes off. If real fires are rare, it will be "correct" the overwhelming majority of the time — silent on every ordinary day — while being useless at the one thing it exists to do. A model can hit the same trap: it looks accurate simply because the thing it's supposed to detect is rare.

What accuracy is: the fraction of all predictions that were correct — correct predictions divided by total predictions. It's the most intuitive metric, and often the least informative one on its own.

Why it fails on imbalanced data: if 99% of examples belong to one class, a model that always predicts that one class scores 99% accuracy while never correctly identifying a single example of the minority class — the exact class that usually matters most in real applications like fraud detection, rare disease screening, or defect detection.

What fails without looking deeper: teams that report only accuracy on an imbalanced problem can ship a model that appears excellent on a dashboard while being completely non-functional for its actual purpose, and nobody notices until the model is already in production and failing on exactly the cases it was built to catch.

🎯 Use this when someone hands you a single accuracy number and you need to know what question to ask next.

2. The Confusion Matrix: The Source of Every Metric

💭 Analogy first: a report card that only says "80% correct" tells you far less than one that says "you missed every fraction problem but got every addition problem right." The confusion matrix is that detailed breakdown for a classifier — not just how often it was right, but exactly which kinds of mistakes it made.

What it does: a confusion matrix counts, for every combination of true class and predicted class, how many examples fell into that combination. For binary classification, this produces four counts with well-known names: true positive (TP), false positive (FP), false negative (FN), and true negative (TN).

The documented convention. scikit-learn's own documentation defines the matrix precisely: entry C[i, j] equals the number of observations known to be in group i and predicted to be in group j — rows are the true class, columns are the predicted class. For binary classification specifically, the documentation states directly that true negatives are C[0,0], false negatives are C[1,0], true positives are C[1,1], and false positives are C[0,1]. This row/column convention matters because some other tools and references use the opposite arrangement, so always confirm which convention a given library or chart is using before reading numbers off it.

Why it's needed: every other metric in this post — accuracy, precision, recall, F1 — is just a different arithmetic combination of these four counts. Understanding the confusion matrix means understanding where every other number actually comes from, rather than treating each metric as a separate black box.

What fails without it: looking only at a single summary metric hides exactly which kind of mistake the model is making — a model with poor precision and one with poor recall can report similar-looking overall scores while failing in completely different, operationally distinct ways.

3. Precision and Recall: Two Different Kinds of "Correct"

💭 Analogy first: a fishing net with very small holes catches nearly every fish in the lake but also scoops up boots, seaweed, and driftwood — high recall, low precision. A net that only closes around a fish it's already certain about, ignoring anything even slightly ambiguous, keeps almost nothing but fish — high precision, low recall. Neither net is objectively "better"; it depends on whether you'd rather sort through junk later or risk missing fish.

scikit-learn's documentation defines both directly from the confusion matrix counts: precision is the ratio tp / (tp + fp), described as intuitively measuring the classifier's ability not to label a negative sample as positive. Recall is the ratio tp / (tp + fn), described as intuitively measuring the classifier's ability to find all the positive samples.

How they diverge in practice: a model that predicts "positive" for almost everything will have very high recall (it rarely misses an actual positive) but low precision (it also wrongly flags many negatives). A model that only predicts "positive" when extremely confident will have high precision but low recall, missing positives it wasn't sure enough about.

What fails without considering both: optimizing only for recall in a spam filter means marking most incoming mail as spam, burying real messages; optimizing only for precision in a disease-screening tool means only flagging the most obvious cases, missing early or ambiguous ones — exactly the cases where catching the disease matters most.

💡 Trade-off: precision and recall usually move in opposite directions as you adjust a model's decision threshold — raising the threshold for calling something "positive" tends to increase precision and decrease recall, and lowering it does the reverse. Neither number alone tells you where to set that threshold; that depends on the real-world cost of each type of mistake.

4. F1 Score: Balancing Precision and Recall

What it does: the F1 score condenses precision and recall into a single number by taking their harmonic mean rather than their simple average. scikit-learn's documentation describes the general F-beta score as a weighted harmonic mean of precision and recall, reaching its best value at 1 and worst at 0, with beta controlling how much more heavily recall is weighted relative to precision — beta equal to 1.0 (the standard F1 score) means precision and recall are weighted equally.

Why a harmonic mean instead of a simple average: a simple average would let a very high precision paired with a very low recall (or vice versa) still produce a moderate-looking combined score. The harmonic mean punishes that imbalance far more severely — F1 only reaches a high value when both precision and recall are reasonably high, which is exactly the behavior you want from a single summary metric.

What fails without F1 (or a similar combined metric): comparing two models by precision alone, or recall alone, can favor a model that's actually worse overall once you account for what it sacrifices on the other metric — F1 is specifically designed to resist that kind of one-sided comparison.

🎯 Use this when you need one number to summarize a trade-off between precision and recall, rather than two numbers that could each individually be misleading.

5. Multiclass Averaging: Micro, Macro, and Weighted

Precision, recall, and F1 are naturally defined per class in a multiclass problem — but reporting one number per class isn't always practical, so frameworks document specific ways to combine them into a single summary value. scikit-learn's documentation defines three of the most common:

  1. Micro averaging: pool the true positives, false positives, and false negatives across all classes first, then compute one precision, recall, and F-score from those combined totals. This treats every individual prediction equally, regardless of which class it belongs to.
  2. Macro averaging: compute the metric separately for each class, then take the unweighted average across classes. The documentation notes this explicitly does not account for label imbalance — a rare class counts exactly as much as a common one.
  3. Weighted averaging: compute the metric separately for each class, then average them weighted by each class's support (how many true instances it has). The documentation notes this can produce an F-score that falls outside the range between precision and recall, since it's not simply "macro, but fairer."

What fails without specifying which averaging method was used: two teams reporting "F1 = 0.85" on the same multiclass problem could be measuring genuinely different things if one used macro and the other used weighted averaging — especially on imbalanced datasets, where the choice can shift the reported number substantially. Always state which averaging method a reported multiclass metric uses.

✅ Worked example: on a three-class problem where one class has ten times more examples than the other two, macro-averaged F1 gives the rare classes equal say alongside the common one — useful for checking rare-class performance specifically — while weighted-averaged F1 will look much closer to the common class's own F1 score, since it's weighted by how often each class actually occurs.

6. Real Example: scikit-learn's Documented Metrics API

scikit-learn's precision_recall_fscore_support function documents an average parameter accepting exactly the strategies described above — 'micro', 'macro', 'weighted', and 'binary' — along with a value of None that returns the raw per-class metrics without any averaging applied, giving full visibility into every class individually.

The function also documents a support return value — the number of true occurrences of each class in the ground-truth labels — which is exactly what weighted averaging uses to weight each class's contribution, making the connection between the averaging strategies and the underlying class frequencies fully explicit rather than hidden.

7. Implementation: Computing Metrics in Code

The following illustrates both the underlying arithmetic and the documented library call that does it for you.

import numpy as np from sklearn.metrics import confusion_matrix, precision_recall_fscore_support y_true = [1, 0, 1, 1, 0, 1, 0, 0, 1, 0] y_pred = [1, 0, 0, 1, 0, 1, 1, 0, 1, 0] # The documented confusion matrix: rows = true, columns = predicted cm = confusion_matrix(y_true, y_pred) tn, fp, fn, tp = cm.ravel() print("confusion matrix:\n", cm) # Computing the metrics from raw counts, matching the documented formulas precision = tp / (tp + fp) recall = tp / (tp + fn) f1 = 2 * (precision * recall) / (precision + recall) accuracy = (tp + tn) / (tp + tn + fp + fn) print(f"accuracy={accuracy:.2f} precision={precision:.2f} " f"recall={recall:.2f} f1={f1:.2f}") # The equivalent, documented library call p, r, f, support = precision_recall_fscore_support( y_true, y_pred, average="binary" ) print(f"sklearn -> precision={p:.2f} recall={r:.2f} f1={f:.2f}")

Confirming the manual calculation against the library call is a useful habit whenever building custom evaluation code — a mismatch usually means a bug in the manual formula, an averaging-strategy mismatch, or a class-label ordering issue rather than a real disagreement about the model's performance.

8. Enterprise Rollout: Choosing and Governing the Right Metric

Which metric a team optimizes for is a decision with real business consequences, and it deserves the same rigor as any other production decision.

Tie the metric to a real cost, before training starts: decide, with stakeholders, what a false positive and a false negative actually cost in the deployed system — lost revenue, wasted reviewer time, missed fraud, a user's bad experience — and choose precision, recall, F1, or a custom weighted combination accordingly, rather than defaulting to accuracy out of habit.

Fix and version the test set: the exact examples and their true labels used to compute a reported metric should be versioned like code, so a metric reported this quarter is genuinely comparable to one reported next quarter, rather than silently measured against a different, shifted sample.

Report per-class and per-subgroup metrics, not just an aggregate: an aggregate F1 score can look healthy while performance on a specific important subgroup (a minority class, a particular user segment) is much worse — this is exactly the kind of gap micro/macro/weighted averaging choices can also obscure if only one is reported.

Monitor metrics in production, not just at training time: the confusion matrix computed once on a held-out test set doesn't guarantee anything about live traffic, where the underlying data distribution can drift over time. Recompute and dashboard these metrics continuously wherever ground-truth labels become available.

Document the threshold and averaging choice alongside every reported number: a precision or recall figure without its decision threshold, and a multiclass F1 without its averaging method, are both incomplete numbers that invite miscommunication between teams.

✅ Practical pattern: report the full confusion matrix alongside any headline metric in model documentation and dashboards, not just the single summary number. Anyone reviewing the model later can recompute precision, recall, or F1 for their own specific concern directly from it.

9. Common Mistakes

Reporting only accuracy on imbalanced data. As Section 1 covers, this can make a model that never detects the minority class look excellent, hiding the exact failure that matters most for the application.

Confusing which axis of the confusion matrix is which. Since row/column conventions differ across tools — scikit-learn documents rows as true and columns as predicted, but other references and libraries can use the opposite — misreading a matrix can silently swap precision and recall, or false positives and false negatives, in an analysis.

Comparing multiclass F1 scores computed with different averaging methods. As Section 5 covers, macro and weighted averaging can produce meaningfully different numbers on the same predictions, and comparing them as if they were interchangeable produces conclusions that don't actually hold.

Leaving the decision threshold unexamined. Precision and recall are both computed at whatever threshold is currently used to convert a model's probability output into a hard positive/negative decision; reporting them without stating or tuning that threshold hides an entire axis of control over the trade-off described in Section 3.

Evaluating on data that leaked into training. Metrics computed on examples the model has already seen, directly or indirectly (through data preprocessing fit on the full dataset, or duplicate records split across train and test), will look better than the model's real, honest performance on new data.

❓ FAQ

If I can only report one metric, should it always be F1?

Not necessarily. F1 is a reasonable default when precision and recall matter roughly equally, but if one type of mistake is genuinely far more costly than the other in your application, a weighted F-beta score, or reporting precision and recall directly at a specific business-relevant threshold, is often more honest than forcing everything into one balanced number.

What counts as a "good" precision or recall score?

There's no universal threshold — it depends entirely on the problem's difficulty and the cost of each error type. A score that's excellent for a highly ambiguous task (like sentiment analysis on sarcastic text) might be unacceptably low for a task with a clear, unambiguous answer, so always compare against a relevant baseline rather than an absolute number.

How is the confusion matrix different for more than two classes?

It becomes an n-by-n grid instead of 2-by-2, with one row and one column per class, and the diagonal representing correct predictions. Precision, recall, and F1 are then computed per class from that row and column, and combined using micro, macro, or weighted averaging as covered in this post.

Why does the F1 score use a harmonic mean instead of a regular average?

A regular average lets a very high score on one metric compensate too easily for a very low score on the other. The harmonic mean is more sensitive to small values, so F1 only scores well when both precision and recall are reasonably good, which better reflects genuinely balanced performance.

Do these metrics apply to deep learning models specifically, or only traditional machine learning?

These metrics are entirely about comparing predicted labels to true labels, so they apply identically regardless of what produced the predictions — a deep neural network, a decision tree, or a simple rule-based system. What changes with deep learning is typically the scale of the data and the need for automated, continuous evaluation, not the metrics themselves.

🔗 References & Further Reading

scikit-learn is a trademark of its respective owners.

📝 Summary

  • Accuracy alone can be misleading on imbalanced datasets, where a model that ignores the minority class can still score very high.
  • The confusion matrix counts true/false positives and negatives, and is the shared source every other classification metric is computed from.
  • Precision measures how trustworthy a positive prediction is; recall measures how completely positives are found — they typically trade off against each other.
  • F1 combines precision and recall via a harmonic mean, punishing imbalance between the two more severely than a simple average would.
  • Multiclass problems require choosing an averaging strategy — micro, macro, or weighted — and that choice can meaningfully change the reported number.
  • scikit-learn's documented API implements all of these definitions directly and consistently.
  • Enterprise evaluation requires tying metric choice to real business costs, versioning test sets, and monitoring metrics continuously in production.
  • Most evaluation mistakes are about which number to trust and how it was computed, not about the arithmetic itself.

Thanks for reading — may your confusion matrices be boring and your F1 scores be high. 🚀

Comments