The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A confusion matrix shows how a classifier’s predictions compare with known labels. Its cells count correct predictions and each kind of mistake, giving developers a practical starting point for evaluating a model. For scikit-learn’s convention, rows are actual classes and columns are predicted classes; check the label order before interpreting any matrix.
How do you read a binary confusion matrix?
For a binary classifier, choose which class is positive and which is negative. With negative labeled 0 and positive labeled 1, scikit-learn places actual classes in rows and predicted classes in columns:
| Actual Predicted | Negative (0) | Positive (1) |
|---|---|---|
| Negative (0) | True negative (TN) | False positive (FP) |
| Positive (1) | False negative (FN) | True positive (TP) |
Each cell counts observations for one actual/predicted pair. In scikit-learn’s API, C[i,j] counts observations known to be in group i and predicted to be in group j. Thus, for this ordering, TN=C[0,0], FP=C[0,1], FN=C[1,0], and TP=C[1,1]. The [actual-row, predicted-column] convention is documented in the scikit-learn confusion_matrix API; other tools may orient their tables differently.
What do TP, FP, TN, and FN mean?
- True positive (TP): an actual positive predicted positive.
- False positive (FP): an actual negative predicted positive.
- True negative (TN): an actual negative predicted negative.
- False negative (FN): an actual positive predicted negative.
“True” means the prediction matches the known label; “false” means it does not. “Positive” and “negative” identify the class assigned to the observation, not whether a prediction is good or bad. For example, a false negative is a missed positive.
#1 Best Overall
Which metrics can you calculate from the counts?
These measures answer different questions about the same predictions. Let TP, FP, TN, and FN be the counts in the matrix:
| Metric | Formula | What it tells you |
|---|---|---|
| Accuracy | (TP + TN) / (TP + TN + FP + FN) | The share of all predictions that are correct. |
| Precision | TP / (TP + FP) | Among predicted positives, the fraction that are actually positive. |
| Recall (true positive rate) | TP / (TP + FN) | Among actual positives, the fraction the classifier finds. |
| False positive rate | FP / (FP + TN) | Among actual negatives, the fraction incorrectly predicted positive. |
| F1 | 2TP / (2TP + FP + FN) | The harmonic mean of precision and recall. |
Precision versus recall
Precision is useful when false alarms are costly or positive predictions need to be trustworthy. Recall matters when missing an actual positive is costly. A model can improve one while worsening the other, so neither number should be read as a complete measure of performance. The Google for Developers metrics guide explains these measures and their relationship to the classification threshold.
Rank #2
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Why accuracy can mislead
Accuracy treats every observation as part of one overall correct-prediction share, so it can conceal poor results on a rare class. Google gives an illustrative, hypothetical case: if positives occur 1% of the time, a classifier that always predicts negative can reach 99% accuracy while detecting none of those positives. This is an example, not a reported dataset result.
What F1 does—and does not—capture
Standard F1 gives precision and recall equal relative contribution through their harmonic mean. It does not directly use true negatives or encode application-specific costs for false positives and false negatives. A strong F1 score therefore does not by itself establish that a model is appropriate for a particular use.
Rank #3
Handle zero denominators explicitly
A metric is undefined when its denominator is zero—for example, precision when there are no predicted positives. Libraries may apply different conventions or allow a configurable result. Scikit-learn’s f1_score documentation describes its zero_division parameter and handling for an absent class. State the convention used in reports instead of presenting an undefined case as an ordinary score.
How should you choose a metric?
Choose measures in light of the decision the classifier supports, rather than selecting a score in isolation:
Rank #4
- Compare error costs: decide whether false positives or false negatives are more harmful. Precision emphasizes the former; recall emphasizes the latter.
- Check class prevalence: when one class is uncommon, do not rely on accuracy alone. Inspect class-specific performance as well.
- Account for the threshold: metrics such as precision and recall are calculated at a chosen classification threshold. Changing that threshold can trade one against the other.
- Choose the reporting level: decide whether you need per-class results or an aggregate, and explain how the aggregate is calculated.
F1 is a useful precision-and-recall summary when that balance fits the task, but it is not a substitute for considering error costs. The scikit-learn model evaluation guide documents metrics and scoring options, including averaging modes.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteHow does a confusion matrix work for multiple classes?
A multiclass confusion matrix has one row and one column per class. Its diagonal cells count correct classifications; off-diagonal cells show which actual class was predicted as which other class. This makes it possible to spot specific confusions that a single aggregate score can hide.
Best Value
Precision, recall, and F-measures can be calculated for each class. If you report one summary across classes, name the averaging convention: macro averaging gives classes equal weight, while weighted averaging accounts for their support (the number of actual observations in each class). Different averaging choices can produce different summaries, particularly when class frequencies differ.
How to calculate a confusion matrix in scikit-learn
Use ground-truth labels as y_true and model predictions as y_pred:
from sklearn.metrics import confusion_matrix
cm = confusion_matrix(y_true, y_pred)
- Confirm label semantics and order. Check which label represents the positive class and inspect the order used in the matrix. The
labelsargument can specify or reorder labels. - Calculate counts. The default output is a matrix of counts. For a binary task using labels 0 and 1 in that order, read rows as actual and columns as predicted under scikit-learn’s convention.
- Normalize only when useful. The
normalizeargument can request normalized values, but keep raw counts available: proportions alone do not show how many observations contributed to a cell. - Report class-level metrics where needed. Pair the matrix with precision, recall, or F-score results, stating the averaging mode and any zero-division convention.
The API accepts y_true, y_pred, optional labels and sample_weight, and a normalize option. Its documented signature and parameter behavior are in the scikit-learn API reference.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

