Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A confusion matrix shows how a classifier’s predictions compare with known labels. Its cells count correct predictions and each kind of mistake, giving developers a practical starting point for evaluating a model. For scikit-learn’s convention, rows are actual classes and columns are predicted classes; check the label order before interpreting any matrix.

How do you read a binary confusion matrix?

For a binary classifier, choose which class is positive and which is negative. With negative labeled 0 and positive labeled 1, scikit-learn places actual classes in rows and predicted classes in columns:

Actual Predicted Negative (0) Positive (1)
Negative (0) True negative (TN) False positive (FP)
Positive (1) False negative (FN) True positive (TP)

Each cell counts observations for one actual/predicted pair. In scikit-learn’s API, C[i,j] counts observations known to be in group i and predicted to be in group j. Thus, for this ordering, TN=C[0,0], FP=C[0,1], FN=C[1,0], and TP=C[1,1]. The [actual-row, predicted-column] convention is documented in the scikit-learn confusion_matrix API; other tools may orient their tables differently.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What do TP, FP, TN, and FN mean?

  • True positive (TP): an actual positive predicted positive.
  • False positive (FP): an actual negative predicted positive.
  • True negative (TN): an actual negative predicted negative.
  • False negative (FN): an actual positive predicted negative.

“True” means the prediction matches the known label; “false” means it does not. “Positive” and “negative” identify the class assigned to the observation, not whether a prediction is good or bad. For example, a false negative is a missed positive.

Which metrics can you calculate from the counts?

These measures answer different questions about the same predictions. Let TP, FP, TN, and FN be the counts in the matrix:

Metric Formula What it tells you
Accuracy (TP + TN) / (TP + TN + FP + FN) The share of all predictions that are correct.
Precision TP / (TP + FP) Among predicted positives, the fraction that are actually positive.
Recall (true positive rate) TP / (TP + FN) Among actual positives, the fraction the classifier finds.
False positive rate FP / (FP + TN) Among actual negatives, the fraction incorrectly predicted positive.
F1 2TP / (2TP + FP + FN) The harmonic mean of precision and recall.

Precision versus recall

Precision is useful when false alarms are costly or positive predictions need to be trustworthy. Recall matters when missing an actual positive is costly. A model can improve one while worsening the other, so neither number should be read as a complete measure of performance. The Google for Developers metrics guide explains these measures and their relationship to the classification threshold.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Why accuracy can mislead

Accuracy treats every observation as part of one overall correct-prediction share, so it can conceal poor results on a rare class. Google gives an illustrative, hypothetical case: if positives occur 1% of the time, a classifier that always predicts negative can reach 99% accuracy while detecting none of those positives. This is an example, not a reported dataset result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What F1 does—and does not—capture

Standard F1 gives precision and recall equal relative contribution through their harmonic mean. It does not directly use true negatives or encode application-specific costs for false positives and false negatives. A strong F1 score therefore does not by itself establish that a model is appropriate for a particular use.

Handle zero denominators explicitly

A metric is undefined when its denominator is zero—for example, precision when there are no predicted positives. Libraries may apply different conventions or allow a configurable result. Scikit-learn’s f1_score documentation describes its zero_division parameter and handling for an absent class. State the convention used in reports instead of presenting an undefined case as an ordinary score.

How should you choose a metric?

Choose measures in light of the decision the classifier supports, rather than selecting a score in isolation:

  • Compare error costs: decide whether false positives or false negatives are more harmful. Precision emphasizes the former; recall emphasizes the latter.
  • Check class prevalence: when one class is uncommon, do not rely on accuracy alone. Inspect class-specific performance as well.
  • Account for the threshold: metrics such as precision and recall are calculated at a chosen classification threshold. Changing that threshold can trade one against the other.
  • Choose the reporting level: decide whether you need per-class results or an aggregate, and explain how the aggregate is calculated.

F1 is a useful precision-and-recall summary when that balance fits the task, but it is not a substitute for considering error costs. The scikit-learn model evaluation guide documents metrics and scoring options, including averaging modes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How does a confusion matrix work for multiple classes?

A multiclass confusion matrix has one row and one column per class. Its diagonal cells count correct classifications; off-diagonal cells show which actual class was predicted as which other class. This makes it possible to spot specific confusions that a single aggregate score can hide.

Precision, recall, and F-measures can be calculated for each class. If you report one summary across classes, name the averaging convention: macro averaging gives classes equal weight, while weighted averaging accounts for their support (the number of actual observations in each class). Different averaging choices can produce different summaries, particularly when class frequencies differ.

How to calculate a confusion matrix in scikit-learn

Use ground-truth labels as y_true and model predictions as y_pred:

from sklearn.metrics import confusion_matrix

cm = confusion_matrix(y_true, y_pred)
  1. Confirm label semantics and order. Check which label represents the positive class and inspect the order used in the matrix. The labels argument can specify or reorder labels.
  2. Calculate counts. The default output is a matrix of counts. For a binary task using labels 0 and 1 in that order, read rows as actual and columns as predicted under scikit-learn’s convention.
  3. Normalize only when useful. The normalize argument can request normalized values, but keep raw counts available: proportions alone do not show how many observations contributed to a cell.
  4. Report class-level metrics where needed. Pair the matrix with precision, recall, or F-score results, stating the averaging mode and any zero-division convention.

The API accepts y_true, y_pred, optional labels and sample_weight, and a normalize option. Its documented signature and parameter behavior are in the scikit-learn API reference.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.