What are some famous AI disasters? They range from a chatbot manipulated into posting offensive messages to biased risk scores, unsafe advice, non-consensual image abuse and a fatal crash during automated-driving tests. These ten widely discussed cases are a curated selection, not a ranking of the worst AI failures. They also differ in evidence and deployment: some involved products in public use, while others were experiments, evaluations or test programs.
What counts as an AI disaster?
Here, “AI disaster” means a consequential failure involving an AI-enabled system or its use—not necessarily an autonomous machine acting alone. Harm can come from generated content, a biased prediction, misleading advice or physical operation. In each case, the system sits within a larger process: its data and design, safeguards, monitoring, human decisions and organizational incentives can all matter.
The cases below are grouped by year. Their evidence is not interchangeable: an official accident investigation, a peer-reviewed study, a company postmortem, investigative reporting and an incident summary each support different kinds of claims.
Ten widely discussed AI failures
1. Microsoft Tay chatbot (2016): adversarial abuse
Microsoft’s Tay was a conversational bot launched on Twitter. Within its first 24 hours, coordinated users exploited weaknesses in how it interacted with and learned from messages, leading it to post offensive content. This was not evidence of a chatbot spontaneously developing beliefs; it was a failure to anticipate adversarial interaction and put adequate safeguards around a public-facing system.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Microsoft vice president Peter Lee wrote in the company’s postmortem: “We take full responsibility for not seeing this possibility ahead of time.” Microsoft’s March 25, 2016 postmortem describes the company’s response.
2. COMPAS risk scoring (2016 public scrutiny): disputed racial bias findings
COMPAS is a criminal-risk assessment tool. In a 2016 analysis of more than 7,000 Broward County, Florida, risk scores, ProPublica reported that Black defendants were more likely to be incorrectly labeled high risk, while white defendants were more likely to be incorrectly labeled low risk. In the analyzed sample, ProPublica said the tool correctly predicted recidivism 61% of the time. Its analysis also found Black defendants were nearly twice as likely as white defendants to be labeled higher risk without subsequently reoffending.
These are findings from ProPublica’s analysis, not universal performance estimates for every jurisdiction or use of COMPAS. Northpointe disputed ProPublica’s methodology, and the disagreement is part of the controversy. Read ProPublica’s analysis and its explanation of the competing claims.
3. Amazon’s experimental recruiting model (reported 2018): gendered resume patterns
Reuters reported that Amazon scrapped an experimental resume-screening system after discovering that it had learned patterns that disadvantaged some resumes associated with women. The tool was not used in production hiring, so this was not a case of the system screening real applicants at scale. It is a cautionary example of how historical patterns in training data can shape an employment model even when the system is meant to rank qualifications.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteRank #2
Reuters’ report describes the experimental system and its limitations.
4. IBM Watson for Oncology: unsafe recommendations in evaluation
STAT reported in 2018 that internal documents described unsafe and incorrect treatment recommendations during evaluation of IBM Watson for Oncology. The important qualification is that this was a reported evaluation finding—not a regulator’s determination that patients were harmed by a deployed system. The case illustrates why clinical decision-support tools need rigorous validation and clinician oversight before their recommendations are trusted.
STAT’s report details what the internal documents described.
5. Uber automated-driving test crash (Tempe, 2018): a pedestrian was killed
On March 18, 2018, an Uber developmental automated-driving test vehicle struck and killed pedestrian Elaine Herzberg in Tempe, Arizona. A human safety operator was in the vehicle; this was a developmental test, not a driverless commercial ride. The National Transportation Safety Board investigated the crash. Its report is the primary source for the sequence of events and findings about the vehicle, safety operator and testing program.
Read the NTSB report. The case shows why physical safety depends on the broader testing and operating system, not just the automated-driving software.
6. Google Photos mislabels Black people as “gorillas” (2015): representational harm
Google Photos’ image recognition labeled images of Black people as “gorillas.” Google apologized and removed the label category. That response removed a harmful label, but it does not establish that the underlying recognition system was comprehensively corrected. The incident is a stark example of how a classification error can reproduce demeaning stereotypes and damage trust.
The incident record summarizes the case.
7. DeepNude (2019): non-consensual synthetic images
DeepNude was an app that generated fake nude images of women from clothed photos. The incident record says its creator pulled it after media exposure, but copies proliferated. The central harm was non-consensual sexualized image abuse: a person’s photograph could be used to fabricate an intimate image without permission.
The incident record summarizes the app and its aftermath.
8. Healthcare risk prediction (2019 study): cost used as a proxy for need
A widely used healthcare risk-prediction algorithm used past healthcare cost as a proxy for medical need. A peer-reviewed Science paper reported that this choice meant Black patients were under-referred for additional care. Spending does not equal need: differences in access and care can make cost a misleading stand-in for health. The lesson is not simply to remove a protected attribute; apparently neutral proxy variables can still produce unequal outcomes.
The study, published in Science, examines the algorithm and its proxy.
9. Robert Williams wrongful arrest (2020): facial-recognition error
Detroit police arrested Robert Williams after an incorrect facial-recognition identification. The incident record says he was detained for roughly 30 hours before his release. The case raises questions about how a probabilistic match is treated by investigators and what human review and accountability should look like before an identification contributes to an arrest. The details here are those recorded in the incident database, not a broader claim about a legal finding.
The incident record summarizes the arrest.
10. Air Canada chatbot refund misinformation (2022): a company held responsible
Air Canada’s chatbot gave a customer incorrect information about a bereavement-fare refund. A British Columbia tribunal held the airline liable for the chatbot’s statements, according to the incident record. The case is a practical reminder that putting advice behind a chatbot does not necessarily remove a company’s responsibility for what customers are told.
Best Value
The incident record summarizes the dispute and tribunal decision.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What these incidents show—and what they do not
These cases are not all failures of the same kind. A manipulated chatbot, an experimental hiring model and a test vehicle call for different safeguards. Nor do they prove that every AI system is unsafe or biased. They show why risk needs to be evaluated in context: what data or proxy a system uses, how it behaves under abuse, whether its output is independently checked, who can act on it, and what happens when it is wrong.
- Bias can enter indirectly. The healthcare study’s cost proxy shows that a system can produce unequal outcomes without explicitly using race as an input.
- Human involvement does not eliminate risk. A person may operate, review or act on a system’s output; the quality of that role and the surrounding process still matter.
- Deployment status changes the claim. Amazon’s system was experimental, Watson’s recommendations were reported in evaluation, and Uber’s vehicle was in developmental testing. Those cases should not be described as equivalent to a widely deployed product harming users.
- Evidence strength varies. The NTSB investigated the Tempe crash; Microsoft published a postmortem; ProPublica published a contested analysis; STAT reported on internal evaluation documents; other entries here rely on the incident record’s summaries. The source type and its limits matter when drawing conclusions.
Other widely discussed incidents include the reversal of the UK’s 2020 A-level grading algorithm, fabricated legal citations in Mata v. Avianca, the suspension of NEDA’s Tessa wellness chatbot in 2023, and errors in Google’s 2024 AI Overviews. The ten above are a selection, not a definitive ranking.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

