Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The F1 score is the harmonic mean of a classifier’s precision and recall. It summarizes how well a model balances false positives and false negatives in one number, but it hides the trade-off between those two measures. Interpret it alongside precision, recall, the averaging method, and the evaluation context.

What the F1 score measures

Precision and recall describe different kinds of classification errors:

  • Precision = TP / (TP + FP). Of the cases the model predicted positive, this is the share that really was positive.
  • Recall = TP / (TP + FN). Of the cases that really were positive, this is the share the model found.

Here, TP means true positives, FP false positives, and FN false negatives. The F1 score combines precision and recall using their harmonic mean:

F1 = 2 × (precision × recall) / (precision + recall)

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Using confusion-matrix counts, the same formula is F1 = 2TP / (2TP + FP + FN). Scikit-learn defines F1 on a scale from 0 (worst) to 1 (best) and notes that precision and recall contribute equally in relative terms. See the scikit-learn F1 score documentation.

How to interpret an F1 score

The harmonic mean is pulled toward the smaller input. A model therefore generally needs both precision and recall to be strong to achieve a high F1. The score is useful when you want one summary that reflects both false positives and false negatives.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

That summary does not show which error is more common or more costly. Two models can have the same F1 but different precision and recall, so report both component values when interpreting or comparing results. F1 also does not include true negatives; it is not a complete measure of performance for every task.

Is F1 useful for imbalanced data?

F1 can be more informative than accuracy when a class imbalance makes accuracy misleading, because it focuses on precision and recall for the positive class (or on per-class values, depending on the averaging method). But it is not automatically the right metric for every imbalanced problem: it does not account for true negatives, and it gives precision and recall equal relative weight. Choose metrics based on the costs, benefits, and risks of the task. Google’s classification metrics guidance explains why accuracy can mislead with imbalanced classes and why precision and recall respond differently to threshold changes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which F1 averaging method should you report?

For multiclass or multilabel classification, “F1” is incomplete unless the averaging method is clear. Scikit-learn documents these options:

Method How it is calculated What it emphasizes
Binary Calculates F1 for one selected positive class. The selected class; this is scikit-learn’s default for the average parameter.
Micro Adds TP, FP, and FN across labels, then calculates F1. Aggregate decisions across labels.
Macro Calculates F1 for each label, then takes the unweighted arithmetic mean. Each class equally, regardless of its number of true instances.
Weighted Averages per-class F1 weighted by each class’s support, or count of true instances. Class results in proportion to class frequency. The result can fall outside the interval between aggregate precision and aggregate recall.
Samples Calculates a score for each instance, then averages those scores. Per-instance results; documented as meaningful for multilabel classification.

These methods answer different questions, so name the method when reporting multiclass or multilabel results. Scikit-learn’s metrics and scoring documentation describes these aggregation distinctions.

How the decision threshold changes F1

A classifier’s decision threshold determines which cases it labels positive. Changing the threshold can change TP, FP, and FN, so precision, recall, and F1 can change as well. A metric value is tied to the threshold at which predictions were evaluated; it is not an inherent, threshold-independent property of the model.

Choose an operating threshold using suitable validation data and the relative costs of false positives and false negatives. When the threshold is relevant to a result, report it alongside precision and recall. Scikit-learn documents precision-recall curves for examining these metrics across decision thresholds.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Undefined cases and zero-division handling

If the calculation’s denominator is zero—for example, when there are no predicted positives and no actual positives—F1 is undefined without a chosen convention. Scikit-learn’s f1_score API provides a zero_division parameter: its default warns and uses 0, while documented alternatives include np.nan. State the convention when a result includes a class or sample with no predicted or actual positives.

What to include when comparing models

For a useful comparison, evaluate models under the same conditions and report the details that make the score interpretable:

  • Precision, recall, and F1—not F1 alone.
  • The averaging method for multiclass or multilabel results.
  • The decision threshold used.
  • Confusion-matrix counts, so readers can see the underlying false positives and false negatives.
  • Class-level results or the chosen averaging approach when class imbalance matters.
  • The application’s relative costs for false positives and false negatives.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.