Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no single authoritative, verified list of exactly 23 types of data bias in machine learning and deep learning. A 2020 article attributed to Ajit Jaokar and Data Science Central is described as a 23-source list, but its complete entries could not be verified. Rather than invent the missing items, this guide explains the common, documented forms of bias and how to distinguish them.

What data bias means in machine learning

Data bias occurs when data or the methods used to collect, process, interpret, or apply it systematically misrepresent the population, situation, or outcome a system is meant to address. A model trained on such data can make unreliable or uneven predictions, even if its code works as designed. The American Academy of Actuaries describes both unrepresentative datasets and flawed data methods as sources of data bias (2023 brief).

Bias is not only a property of a dataset or algorithm. NIST’s 2022 AI bias framework emphasizes that bias also manifests in the societal context in which AI systems are developed and used (NIST announcement). A technically accurate model can still create harm if its intended use, decision process, or effects on people are not considered.

Common types of bias and where they enter

The categories below are useful working labels, not a universal taxonomy. They overlap: a dataset can be selectively collected, poorly measured, and shaped by historical inequality at the same time. IBM’s 2024 overview describes these as commonly discussed forms of bias, rather than claiming to supply a definitive list (IBM, “What is Data Bias?”).

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Type How it arises Example or signal
Selection bias The process for including examples systematically favors some cases over others. A dataset excludes people who are less likely to appear in the source records.
Sampling bias A sample does not adequately represent the population the model will encounter; IBM treats it as a form of selection bias. A medical prediction model trained on a narrow patient population may perform poorly for patients unlike those represented.
Measurement bias A variable is measured inaccurately or differently across people or settings. A proxy measurement may not capture the underlying trait equally well for every group.
Reporting bias Some events, opinions, or outcomes are more likely to be recorded than others. A sentiment model trained on reviews may overrepresent people with especially strong opinions.
Historical or temporal bias Past patterns in data carry forward inequalities or circumstances that have changed. A hiring model trained on historical employment patterns may reproduce past disparities.
Exclusion bias Relevant people, variables, or cases are left out during collection or preparation. Omitting a population from the data can make the model’s predictions less useful for it.
Cognitive, confirmation, or implicit bias Human expectations influence which data is gathered, how labels are assigned, or which results are accepted. Annotators or designers may interpret ambiguous cases in ways that reinforce prior assumptions.
Automation bias People give automated outputs undue weight, including when evidence should prompt review. A decision-maker may accept a model recommendation without checking a questionable input or result.

How to examine a dataset and its use

Start with the decision the model is meant to support, not with a checklist of labels. Define the population and conditions in which it will operate, then examine how the data and outcomes were created.

  1. Define the target population and task. Specify who will be affected, what situations the model will face, and what outcome it is supposed to predict or inform.
  2. Trace data inclusion. Identify where examples came from and who could be missing because of access, eligibility, participation, or record-keeping practices.
  3. Review measures and labels. Check whether variables and outcome labels mean the same thing across groups and settings, and whether proxies genuinely represent the intended concept.
  4. Compare coverage and performance. Look for gaps in representation and evaluate errors for relevant groups and conditions, rather than relying only on an overall score.
  5. Assess context and consequences. Consider how people use the system, what happens when its output is wrong, and whether the deployment setting changes the risks.

This approach reflects the broader socio-technical perspective in NIST’s Special Publication 1270, “Towards a Standard for Identifying and Managing Bias in Artificial Intelligence” (March 2022). Finding a bias label does not by itself establish that a model is unfair or fair; the relevant evidence depends on the system’s intended population, use, and outcomes.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why “23 types” is not a settled classification

Lists of bias terms vary because they group different kinds of things together: data-collection mechanisms, statistical effects, human behaviors, and system-level consequences. A partial secondary reproduction of the 2020 list attributed to Jaokar includes labels such as aggregation bias, population bias, Simpson’s paradox, longitudinal data fallacy, behavioral bias, popularity bias, algorithmic bias, emergent bias, omitted-variable bias, cause-effect bias, and funding bias. Because the complete list and exact wording are not verified, these entries should not be presented as the definitive 23 or treated as directly comparable categories.

For practical work, organize the question by where the problem enters the lifecycle, how it arises, whose experiences are missing or mischaracterized, how measures and labels were formed, and what evidence would reveal the issue. This makes the analysis actionable without implying that every source uses the same vocabulary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.