The F1 score is the harmonic mean of a classifier’s precision and recall. It summarizes how well a model balances false positives and false negatives in one number, but it hides the trade-off between those two measures. Interpret it alongside precision, recall, the averaging method, and the evaluation context.
What the F1 score measures
Precision and recall describe different kinds of classification errors:
- Precision = TP / (TP + FP). Of the cases the model predicted positive, this is the share that really was positive.
- Recall = TP / (TP + FN). Of the cases that really were positive, this is the share the model found.
Here, TP means true positives, FP false positives, and FN false negatives. The F1 score combines precision and recall using their harmonic mean:
F1 = 2 × (precision × recall) / (precision + recall)
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
Using confusion-matrix counts, the same formula is F1 = 2TP / (2TP + FP + FN). Scikit-learn defines F1 on a scale from 0 (worst) to 1 (best) and notes that precision and recall contribute equally in relative terms. See the scikit-learn F1 score documentation.
How to interpret an F1 score
The harmonic mean is pulled toward the smaller input. A model therefore generally needs both precision and recall to be strong to achieve a high F1. The score is useful when you want one summary that reflects both false positives and false negatives.
Rank #2
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
That summary does not show which error is more common or more costly. Two models can have the same F1 but different precision and recall, so report both component values when interpreting or comparing results. F1 also does not include true negatives; it is not a complete measure of performance for every task.
Is F1 useful for imbalanced data?
F1 can be more informative than accuracy when a class imbalance makes accuracy misleading, because it focuses on precision and recall for the positive class (or on per-class values, depending on the averaging method). But it is not automatically the right metric for every imbalanced problem: it does not account for true negatives, and it gives precision and recall equal relative weight. Choose metrics based on the costs, benefits, and risks of the task. Google’s classification metrics guidance explains why accuracy can mislead with imbalanced classes and why precision and recall respond differently to threshold changes.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsRank #3
Which F1 averaging method should you report?
For multiclass or multilabel classification, “F1” is incomplete unless the averaging method is clear. Scikit-learn documents these options:
| Method | How it is calculated | What it emphasizes |
|---|---|---|
| Binary | Calculates F1 for one selected positive class. | The selected class; this is scikit-learn’s default for the average parameter. |
| Micro | Adds TP, FP, and FN across labels, then calculates F1. | Aggregate decisions across labels. |
| Macro | Calculates F1 for each label, then takes the unweighted arithmetic mean. | Each class equally, regardless of its number of true instances. |
| Weighted | Averages per-class F1 weighted by each class’s support, or count of true instances. | Class results in proportion to class frequency. The result can fall outside the interval between aggregate precision and aggregate recall. |
| Samples | Calculates a score for each instance, then averages those scores. | Per-instance results; documented as meaningful for multilabel classification. |
These methods answer different questions, so name the method when reporting multiclass or multilabel results. Scikit-learn’s metrics and scoring documentation describes these aggregation distinctions.
Rank #4
How the decision threshold changes F1
A classifier’s decision threshold determines which cases it labels positive. Changing the threshold can change TP, FP, and FN, so precision, recall, and F1 can change as well. A metric value is tied to the threshold at which predictions were evaluated; it is not an inherent, threshold-independent property of the model.
Choose an operating threshold using suitable validation data and the relative costs of false positives and false negatives. When the threshold is relevant to a result, report it alongside precision and recall. Scikit-learn documents precision-recall curves for examining these metrics across decision thresholds.
Best Value
Undefined cases and zero-division handling
If the calculation’s denominator is zero—for example, when there are no predicted positives and no actual positives—F1 is undefined without a chosen convention. Scikit-learn’s f1_score API provides a zero_division parameter: its default warns and uses 0, while documented alternatives include np.nan. State the convention when a result includes a class or sample with no predicted or actual positives.
What to include when comparing models
For a useful comparison, evaluate models under the same conditions and report the details that make the score interpretable:
Quick Recap
- Precision, recall, and F1—not F1 alone.
- The averaging method for multiclass or multilabel results.
- The decision threshold used.
- Confusion-matrix counts, so readers can see the underlying false positives and false negatives.
- Class-level results or the chosen averaging approach when class imbalance matters.
- The application’s relative costs for false positives and false negatives.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

