Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a chosen positive class, calculate precision as TP ÷ (TP + FP), recall as TP ÷ (TP + FN), and F1 as 2TP ÷ (2TP + FP + FN). With imbalanced classes, report per-class results and identify the averaging method: macro gives classes equal weight, while weighted averaging reflects their true support and can let the majority class dominate.

How do you calculate precision and recall?

Start by identifying the positive class. In a binary task, “positive” is the class you want to detect; it does not have to be labeled 1. Then count outcomes for that class:

  • True positive (TP): the model predicts the positive class and the true label is positive.
  • False positive (FP): the model predicts positive, but the true label is negative.
  • False negative (FN): the true label is positive, but the model predicts negative.
  • True negative (TN): the model predicts negative and the true label is negative.

Precision and recall use different denominators, so they answer different questions. Scikit-learn defines them in its metrics and scoring guide and F1 API reference.

  • Precision = TP ÷ (TP + FP): among instances predicted positive, what fraction truly belongs to the positive class? Low precision means many false alarms.
  • Recall = TP ÷ (TP + FN): among actual positive instances, what fraction did the model find? Low recall means many missed positives.

Neither metric is inherently better. If false alarms are especially costly, precision may be a priority; if missing a positive case is more costly, recall may matter more.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

How do you calculate F1 from TP, FP, and FN?

F1 is the harmonic mean of precision and recall. You can calculate it from the two metrics or directly from the confusion-matrix counts:

  • F1 = 2 × precision × recall ÷ (precision + recall)
  • F1 = 2TP ÷ (2TP + FP + FN)

The direct formula avoids rounding precision and recall before combining them. F1 balances the two metrics equally; a high score generally requires both to be high. It does not include TN, so it is not a general measure of all correct predictions.

Worked example

Suppose the true labels are [1,1,1,0,0,0,0,0,0,0] and the predictions are [1,0,1,1,0,0,0,0,0,0]. If label 1 is the positive class, the counts are TP = 2, FP = 1, FN = 1, and TN = 6.

  • Precision = 2 ÷ (2 + 1) = 0.667
  • Recall = 2 ÷ (2 + 1) = 0.667
  • F1 = 2 × 2 ÷ (2 × 2 + 1 + 1) = 0.667

Accuracy is 8 ÷ 10 = 0.8, but that single figure does not show which errors the model made. These values are calculated from the listed labels, not from a published dataset.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should you report scores for imbalanced classes?

In multiclass classification, treat each class as the positive class in turn and calculate its precision, recall, and F1 against all other classes. Then report how you combine the class-level scores. Scikit-learn documents these averaging choices in its metrics guide and F-beta API reference.

Method How it combines classes What it tells you—and can hide
Macro Calculates the metric for each class, then takes the arithmetic mean. Every class has equal weight. Makes weak performance on a rare class more visible, but does not reflect how common each class is.
Weighted Averages per-class scores using each class’s true support (the number of actual instances in that class) as its weight. Reflects performance in proportion to observed class counts, but a majority class can dominate the result.
Micro Adds TP, FP, and FN across classes first, then calculates the metric from those totals. In ordinary single-label multiclass classification with all classes included, micro precision, recall, and F1 equal accuracy; that can obscure minority-class failures.

Two details are easy to miss: weighted recall equals accuracy, and weighted F1 is not guaranteed to fall between weighted precision and weighted recall. These behaviors are documented by scikit-learn in its metrics guide and F-beta reference.

Choosing what to show

There is no universally best average for imbalanced data. Choose based on the decision the score needs to support:

  • Use macro when each class matters independently or when you want rare classes to count as much as common ones.
  • Use weighted when the overall score should reflect the class frequencies in the evaluated data, while recognizing that common classes can overwhelm rare-class results.
  • Use micro when the aggregate count of correct positive decisions is relevant, but do not treat it as evidence that every class performs well.
  • Show per-class precision, recall, F1, and support when a single average could conceal important differences. Scikit-learn’s classification report provides per-class values and support alongside macro and weighted averages; its micro row is conditional.

When reporting a score, name the positive class for binary results, the averaging method for multiclass results, and the supports when class prevalence matters. Interpret the scores against the relative costs of false positives and false negatives.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When should you use F-beta instead of F1?

F1 weights precision and recall equally. F-beta generalizes the F-measure: a beta greater than 1 gives recall more weight, while beta below 1 gives precision more weight. Choose beta to reflect whether missing relevant cases or raising false alarms is more costly, and state the beta value when reporting the result. Scikit-learn describes F-measures as weighted harmonic means in its metrics guide and F-beta reference.

How do decision thresholds and undefined scores affect the result?

Thresholds change hard-label scores

When a model produces scores or probabilities, precision and recall for hard predictions depend on the threshold used to assign class labels. Changing the threshold changes the balance between false positives and false negatives. If threshold selection is important, compare the precision-recall curve across thresholds and choose an operating point that fits the task’s error costs; scikit-learn documents this approach in its metrics guide.

Zero denominators need an explicit convention

A metric can be undefined when its denominator is zero. For example, if a class is absent from both the true labels and predictions, its F1 has no defined value from the formula. Scikit-learn’s F1 API defaults this all-zero case to 0.0 with a warning; its zero_division setting controls alternatives, and support for np.nan was added in scikit-learn 1.3. The F-beta and recall APIs also document zero-division behavior. State the convention or software setting used so readers can distinguish a chosen display value from a measured failure.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.