The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Famous technology failures teach developers to look beyond the bug. Ariane 5 Flight 501 exposed the danger of carrying software assumptions into a new operating environment and relying on identical redundant systems. The Therac-25 accidents showed why software must be protected by system-level safeguards, oversight and evidence-gathering. Together, they point to a practical discipline: test realistic conditions, contain failures, make systems observable, and treat incident response as engineering work.
Why famous technology failures matter to developers
A failure rarely has just one useful explanation. A coding error may be the trigger, but design choices, operating assumptions, testing gaps, safeguards and response practices shape what happens next. Looking at those factors is more useful than assigning blame to one engineer or treating a failure as proof that a whole technology is unsafe.
The Ariane 5 and Therac-25 cases are technically and historically distinct; their human consequences should not be treated as a scorecard. Their value for software teams is in comparing how assumptions crossed system boundaries, what protections were available, and whether the system made its condition visible enough to investigate.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Why Ariane 5 Flight 501 failed
Ariane 5 Flight 501, its first flight, failed on 4 June 1996. The European Space Agency’s inquiry summary attributes the loss of guidance and attitude information to specification and design errors in inertial reference system software. It also found that reviews and tests had not adequately analyzed the inertial reference system and the complete flight control system to detect the failure. ESA’s inquiry-board summary reports that guidance and attitude information were completely lost 37 seconds after the main-engine ignition sequence began, or 30 seconds after lift-off.
#1 Best Overall
How an inherited assumption became a flight failure
The inquiry report explains that the inertial reference system software had been carried over from Ariane 4. An alignment function intended for pre-launch operation continued running after lift-off. During Ariane 5’s flight, an internal value became too large to convert to a 16-bit signed integer. That conversion raised an Operand Error. Both the active and backup inertial reference systems had identical software and encountered the same exception; guidance software then treated diagnostic data from the failed system as flight data. The Ariane 5 Flight 501 Inquiry Board report describes this chain.
What the case says about reuse and redundancy
The lesson is not to reject code reuse. It is to revalidate the assumptions behind reused code when the mission, inputs, operating conditions or failure consequences change. Teams should ask whether inherited functions are still needed, whether input ranges remain valid, and what the system will do when a function fails.
Rank #2
- Supplies and preparations
- Energy, heat and power
- Low-tech medicine and healing
- Water quality and treatment
- Food, shelter and first aid
Redundancy also needs independence. Two components that share the same software and assumptions may fail together; having a backup does not by itself prevent a common-mode failure. The inquiry board recommended switching off unneeded functions after lift-off, reviewing critical software and double-failure handling, improving telemetry collection, and using representative equipment and simulated trajectories in qualification. It summarized its stance on critical software this way: “The Board is in favour of the opposite view, that software should be assumed to be faulty until applying the currently accepted best practice methods can demonstrate that it is correct.”
What Therac-25 teaches about safety
Nancy Leveson and Clark S. Turner’s investigation of the Therac-25 accidents treats safety as a system property, not a matter of checking whether code appears correct in isolation. They point to the protective role of hardware interlocks in the earlier Therac-20, which mitigated the consequence of the same software error implicated in the Tyler deaths. Their central formulation is: “Safety is a quality of the system in which the software is used; it is not a quality of the software itself.” The authors’ investigation, reprinted from IEEE Computer in July 1993, is available through MIT.
Build protection beyond the software
Software quality practices matter, but they cannot be the only barrier between an error and harm. Leveson and Turner argue for safety assurance at the system level even when software errors occur. That means considering independent safeguards, the way components interact, the conditions under which operators use the system, and the procedures for reporting problems.
The authors also recommend documentation, simple designs, software quality assurance, and extensive testing and formal analysis at both module and software levels. Audit trails should be designed in from the beginning so that a later investigation is not forced to rely on incomplete recollections or transient system state. User and government oversight are part of the safety picture, not substitutes for sound engineering.
Rank #4
- Author: Kranz, Gene.
- Publisher: Simon & Schuster
- Pages: 416
- Publication Date: 2009
- Binding: Paperback
What production incidents teach about coping with failure
In a 2020 qualitative study, Jonathan Sillito and Esdras Kutomi examined 30 software incidents: 15 drawn from in-depth interviews with engineers and 15 from published incident reports. The study considers how failures occur, how teams detect them, how they investigate, and how they mitigate their effects. Its cases are not a statistically representative estimate of software failures, but they show why incident response should be treated as deliberate engineering work. The paper also notes that failures can cascade and that teams may not understand system scaling limits until they are exceeded. Read the study on arXiv.
A practical response sequence
- Mitigate immediate impact. Choose a containment action suited to the incident. Rolling back a deployment is one possible mitigation, not a universal answer.
- Keep observing the system. Continue monitoring behavior after mitigation; a reduction in visible symptoms does not establish that the underlying cause is understood.
- Preserve evidence. Retain relevant logs, telemetry and other records before they are lost or overwritten. Audit trails designed into a system make this work more reliable.
- Investigate contributing conditions. Reconstruct the sequence of events and examine assumptions, interfaces, safeguards and operating conditions—not only the component where an error surfaced.
- Turn findings into reviewable changes. Record what should change in the design, tests, safeguards or operating procedures, and subject corrective work to review. An incident report alone does not prevent recurrence.
How to compare failures without flattening them
Ariane 5 and Therac-25 do not share one cause or call for one remedy. A useful comparison asks the same engineering questions while respecting the differences between a flight-control system and a medical device.
| Question | Ariane 5 Flight 501 | Therac-25 analysis | Development takeaway |
|---|---|---|---|
| What assumptions crossed a boundary? | Software carried over from Ariane 4 continued an alignment function after lift-off; Ariane 5’s operating conditions produced a value outside the conversion range. | Leveson and Turner caution that prior software use does not guarantee safety in a different system; they note the earlier Therac-20’s hardware interlocks mitigated a software error’s consequence. | Reassess assumptions when software moves to a new context, and check whether protections from the old system remain present. |
| Could failure be contained? | Active and backup inertial reference systems had identical software and encountered the same exception. | The authors emphasize system-level safety and the protective role of hardware interlocks. | Test whether safeguards and redundancy are genuinely independent, and define safe behavior when a component fails. |
| Did testing represent real conditions? | The inquiry found inadequate analysis and testing of the inertial reference system and complete flight control system; it recommended representative qualification and simulated trajectories. | The analysis recommends extensive testing and formal analysis at module and software levels. | Combine component checks with end-to-end tests that exercise realistic inputs, operating conditions and failure paths. |
| Could teams observe and learn? | The board recommended improved telemetry collection. | The authors recommend audit trails, reporting procedures and user oversight. | Design evidence collection and reporting into the system and the organization, rather than improvising them after an incident. |
Turn failure analysis into safer engineering
For a development team, these cases suggest a repeatable review habit: identify the assumptions embedded in software and system design, examine whether independent safeguards can contain errors, and ask whether tests cover realistic system behavior rather than only isolated components. Then consider how an operator or investigator would know what happened and what evidence would be available.
When an incident occurs, separate immediate mitigation from causal investigation. Preserve evidence, examine contributing conditions across boundaries, and make corrective changes reviewable. That approach does not guarantee that failures will never happen; it gives teams a better chance to limit their consequences and learn from them.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

