iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
sklearn.metrics.accuracy_score reports the share of evaluated samples whose predicted labels match their true labels. That makes it easy to read, but a single accuracy value can hide poor results for minority classes, costly types of error, or partial matches in a multilabel task. Pair it with metrics that reflect what matters in your evaluation.
What does accuracy_score return?
The documented call is sklearn.metrics.accuracy_score(y_true, y_pred, *, normalize=True, sample_weight=None). In ordinary binary or multiclass classification, it compares each predicted label with the corresponding true label and aggregates the matches.
- With the default
normalize=True, the result is the fraction of correctly classified samples, from 0 to 1. - With
normalize=False, it returns the number of correct samples. sample_weightlets you weight samples in the calculation; explain the reason for the weights when reporting the result.
The API example returns 0.5 for two correct predictions among four, and 2.0 for the same predictions with normalize=False. See the scikit-learn accuracy_score API documentation.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For the score to be meaningful, y_true and y_pred must correspond sample by sample and use the intended label representation. The API accepts one-dimensional labels as well as multilabel indicator arrays or matrices. Accuracy alone does not identify which classes were missed, distinguish types of error, or measure whether predicted probabilities are calibrated.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
How does accuracy work in multilabel classification?
In multilabel classification, accuracy_score uses subset accuracy: a sample counts as correct only when its entire predicted label set exactly matches its true label set. If a sample has several true labels and the model gets all but one right, that sample is still counted as incorrect.
Consequently, do not describe this result as the percentage of individual labels predicted correctly. If partial matches matter, report the strict subset score alongside per-label precision, recall, or F1, or Hamming loss. The API defines the behavior in its accuracy_score documentation.
Rank #2
When can accuracy be misleading?
Class imbalance
Accuracy weights samples, not classes. When one class dominates the evaluation set, a classifier can score well by predicting the common class while performing poorly on a rare class. That is especially problematic when the rare class is important to detect. The scikit-learn model evaluation guide describes balanced accuracy as a measure that avoids inflated performance estimates on imbalanced datasets.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Accuracy is not inherently invalid: it can be useful when the evaluation examples and the relative importance of their errors reflect the decision you need to make. But give the class distribution and relevant per-class results when an aggregate could obscure the pattern. See the scikit-learn model evaluation guide.
Rank #3
Different errors have different costs
A false positive and a false negative both count as incorrect predictions in ordinary accuracy, even if one is much more costly in your application. Accuracy does not encode that difference. Choose and report measures that expose the error type and class of interest instead of assuming the aggregate captures the practical cost.
A high score is not a future-performance guarantee
accuracy_score summarizes predictions against labels in the data you evaluated. It is not proof of how the model will perform on future examples. Say how predictions were generated and identify the evaluation data or cross-validation design. Scikit-learn’s guide discusses scoring in cross-validation and model-selection tools.
Rank #4
Which metric should you add?
| Evaluation need | Useful measure | How to interpret or report it |
|---|---|---|
| Give each class’s ability to be found equal weight | Balanced accuracy | Average recall across classes; scikit-learn also documents it as accuracy with class-balanced sample weights. |
| See false positives and false negatives separately | Precision and recall | Report class-specific values or explain the averaging method. |
| Combine precision and recall in one summary | F1 | State the averaging choice and recognize that a single summary hides the precision–recall trade-off. |
| Evaluate ranking from prediction scores, not only final labels | ROC AUC | State the class setup and multiclass configuration; the API has restrictions and configuration parameters. |
| Allow one of several top-ranked classes to count as correct | Top-k accuracy | Define k; a prediction counts when the true class is among the k highest-scored classes. |
| Inspect label-level errors in multilabel tasks | Per-label precision, recall, or F1; Hamming loss | Use alongside subset accuracy to reveal partial matches and label-specific performance. |
For precision, recall, and F1, the averaging method changes the question answered: macro gives classes equal weight, weighted accounts for class support, and micro pools contributions across sample-class pairs. The scikit-learn guide explains these distinctions. Metric choice should follow the class importance, error costs, need to assess ranking scores, and tolerance for partial multilabel matches.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteHow to report an accuracy score clearly
- Confirm that true and predicted labels align sample by sample and use the intended representation.
- Choose whether to report a fraction (the default) or a count with
normalize=False. If you usesample_weight, state why that weighting is appropriate. - For imbalanced data, report class distribution and add a class-sensitive measure such as balanced accuracy or per-class recall.
- For multilabel results, name the metric as subset accuracy and consider adding label-level measures.
- Describe the evaluation data or cross-validation design, and explain how predictions were produced.
The stable documentation can advance independently of an installed project version; check the scikit-learn version used by your project when behavior or API details are version-sensitive. The cited documentation search showed version 1.9.1.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

