The most damaging cloud outages are not always caused by a provider losing every server. They happen when a small defect or change reaches a shared dependency—such as storage, networking, a control plane, or a backup path—and spreads farther than the systems designed to contain it. This representative editorial ranking weighs blast radius, duration, cascading dependencies, and lasting engineering lessons; it is not an objective league table. Several older incidents have limited documented detail here, so their exact scope or duration is identified as unresolved rather than guessed.
How the incidents compare
The comparison below distinguishes confirmed figures from details that are not established in the cited accounts. “Not established” does not mean an incident had no impact; it means a reliable scope or duration is not available here.
| Incident | Trigger or failure mode established | Scope and duration established |
|---|---|---|
| Windows Azure, 2012 leap day | Date-handling failure | Exact scope and duration not established; historical incident index |
| AWS S3, February 28, 2017 | Authorized operator command removed more servers than intended from an S3 billing subsystem | Severe four-hour US-EAST-1 disruption; effects lasted up to 11 hours for some websites, according to a 2018 Lloyd’s / Singapore Reinsurers report |
| Azure Storage, 2017 | Storage and regional-dependency incident; detailed trigger not established | Exact scope and duration not established; historical incident index |
| Google Cloud asia-northeast1, June 8, 2017 | Topology upgrade, routing misconfiguration, and manual change | 62 minutes of lost external connectivity in the region |
| GitLab database and backup loss, January 31, 2017 | Accidental PostgreSQL deletion; backups were missing | Database loss and failed backup recovery documented in GitLab’s public postmortem; no comparable duration figure is stated here |
| AWS us-east-1, 2019 | Automation, cascading failure, and cloud configuration | Major regional incident; exact scope and duration not established |
| Cloudflare global WAF, July 2, 2019 | One WAF regular expression drove CPU to 100% worldwide | About 30 minutes of global 502 errors; traffic dropped 82% at the worst point |
| Google Cloud networking, 2020 | Networking and control-plane dependency case | Exact impact figures not established |
| Fastly CDN, June 8, 2021 | Edge-network concentration and a latent-bug/configuration case; exact trigger not established here | Widely cited as about a one-hour internet disruption; exact customer count not established |
| AWS us-east-1 networking, December 7, 2021 | Regional networking failure | Widespread event; exact duration and blast radius not established here |
Ten outages and the reliability lessons they expose
1. Windows Azure’s leap-day disruption (2012)
A leap-year date-handling failure is a reminder that calendar logic can be a production dependency, not harmless bookkeeping. The historical incident index confirms the event, but the available account does not establish its precise duration or a detailed service-by-service scope. It would be misleading to assign either a figure.
Lesson: test date boundaries—including leap days, time-zone transitions, and year changes—in systems that make deployment, certificate, billing, or resource-lifecycle decisions. A managed cloud’s control plane can be vulnerable to assumptions that look routine in application code.
#1 Best Overall
- FAST 15-MINUTE DEPLOYMENT – Provision and configure in just 15 minutes (down from 40+ minutes with previous models). Perfect for field technicians who need to get sites up and running quickly without deep networking expertise.
- UPGRADED PERFORMANCE – Powered by the Allwinner H618 processor with 1GB LPDDR4 RAM (double the previous generation). Enables accurate speed tests on gigabit connections and supports SNMP v3 encryption for enhanced security monitoring.
- PLUG-AND-PLAY SIMPLICITY – No complex configuration required. Simply connect to your network via the Gigabit Ethernet port, power up with the included USB-C cable, and start monitoring. Multi-VLAN support with just a few clicks in the interface.
- RISK MITIGATION FOR MSPs – Domotz maintains the operating system and security updates, transferring liability concerns away from your organization. Eliminates the security risks of deploying monitoring software on customer-managed servers or domain controllers.
- UNIVERSAL CONNECTIVITY – USB-C power port (more durable and universal than previous micro USB), Gigabit Ethernet port, and USB 2.0 port for future expansion. Premium casing designed for rack mounting or standalone deployment in professional environments.
2. AWS S3 in US-EAST-1 (February 28, 2017)
AWS’s account says an authorized operator ran a command intended to remove a small number of servers from an S3 billing subsystem. The command removed more capacity than intended. S3 APIs became unavailable, and services that depended on S3 or related regional functions were affected: the S3 console, new EC2 instance launches, EBS snapshot-backed volume operations, and Lambda. A 2018 Lloyd’s / Singapore Reinsurers report describes a severe four-hour disruption in US-EAST-1 and effects lasting as long as 11 hours for some websites.
What it shows: a maintenance action in a subsystem can have a much wider blast radius when the command’s scope is not tightly bounded or when multiple products share regional dependencies. The incident is a clear answer to “what caused the AWS S3 outage”: an operator command removed too much capacity, and dependent services felt the loss. The accounts summarized here do not establish detailed first-signal, rollback, or customer-communication timelines.
Control to prioritize: cap the amount of infrastructure a single command can affect, require staged execution and independent review for high-impact actions, and rehearse rollback paths. Treat dependency mapping as part of capacity planning, not only application architecture.
3. Azure Storage (2017)
A 2017 Azure Storage incident is identified in the historical incident index as a storage-control-plane and regional-dependency case. The available incident detail does not establish the exact triggering change, affected services, duration, detection path, recovery sequence, or communications timeline.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteRank #2
- Hardware Controller with Professional Network Management-Centralized management for up to 100 Omada devices including Omada access points, Omada Security Gateways and Jetstream switches.
- Premium Hardware Design-Industry-leading flexible Rackmount/Desktop design with a powerful chipset, durable metal casing, 2 fast ethernet ports and 1 USB 2.0 port for auto backup.
- Dual power selection-Support PoE (802.3af/802.3at) and micro USB for flexible installations.
- Easy Network Monitor & Maintenance-The easy-to-use dashboard makes it simple to see your real-time network status and improve network maintenance for peace of mind.
- Cloud Access with No License Fee-Enjoy cloud service with no license fee with the use of OC200. Remote Cloud access and Omada app brings centralized cloud management of the whole network from different sites—all controlled from a single interface anywhere, anytime.
Lesson: storage reliability includes the control plane that provisions, routes, and manages storage, not just the durability of stored data. Customers should map which application functions depend on regional storage management and decide what can continue in degraded or read-only mode if those operations are unavailable.
4. Google Cloud asia-northeast1 networking (June 8, 2017)
Google reports 62 minutes of lost external connectivity in asia-northeast1, affecting Compute Engine, App Engine, Cloud SQL, Cloud Datastore, and Cloud Storage. During a topology upgrade, existing links were decommissioned before replacement links could carry traffic. A routing misconfiguration and a manual change that bypassed per-zone restrictions caused the region’s zones to fail together. Google also found that a load-balancing health-detection feature failed to identify unhealthy backends.
Google’s incident report states: “We recognize we failed to deliver the regional reliability that multiple zones are meant to achieve.” The important point is that a nominal isolation boundary did not hold under a real manual change, and the health signal did not correctly identify the failure.
Controls to prioritize: validate topology changes before removing existing capacity, enforce per-zone limits even for manual operations, and test health checks against realistic failure states. An isolation policy that can be bypassed in an emergency is not a reliable isolation boundary.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
- 【Hardware Controller with Greater Network Management】Latest Omada SDN hardware controller provides centralized management for up to 500 Omada devices including Omada access points, Omada switches and Omada routers.
- 【Premium Hardware Design】Industry-leading flexible Rackmount/Desktop design with a powerful chipset, durable metal casing, 2 * gigabit ports and 1 * USB 3.0 port for auto backup.
- 【Easy Network Monitor & Maintenance】The easy-to-use dashboard makes it simple to see your real-time network status and improve network maintenance for peace of mind.
- 【Cloud Access with No License Fee】Enjoy cloud service with no license fee with the use of OC300. Remote Cloud access and Omada app brings centralized cloud management of the whole network from different sites—all controlled from a single interface anywhere, anytime.
- 【SDN Compatibility】For SDN usage, make sure your devices/controllers are either equipped with or can be upgraded to SDN version. OC300 work only with SDN APs, Switches and Gateways. For devices that are compatible with SDN firmware, please visit TP-Link website.
5. GitLab’s database and backup loss (January 31, 2017)
Google’s reliability guidance points to GitLab’s public postmortem: PostgreSQL data was accidentally deleted, and the team discovered that backups were missing. This was not simply a cloud provider’s regional outage; it belongs in a cloud reliability list because it exposes a core failure in recovery readiness. A backup strategy that cannot produce a usable restore when needed is not a working recovery strategy.
Controls to prioritize: keep backup copies independent of the systems and credentials that could compromise production, monitor backup completion, and routinely restore into a separate environment. Measure recovery time and data loss in exercises rather than assuming that a successful backup job proves recoverability.
6. AWS us-east-1 (2019)
The historical incident index identifies a major us-east-1 AWS event involving automation, cascading failure, and cloud configuration. It does not establish a sufficiently detailed primary account here for an exact customer count, duration, or product-by-product impact, so those figures should not be inferred.
Lesson: automation can accelerate both recovery and failure. A configuration or control-plane dependency that is shared across services can turn a bounded issue into a cascade. Put limits on automated actions, provide a known-good configuration to roll back to, and ensure recovery does not depend solely on the same impaired control plane.
Rank #4
7. Cloudflare’s global WAF outage (July 2, 2019)
Cloudflare reports that a single Web Application Firewall regular expression drove CPU to 100% worldwide. The rule was deployed globally in one step. Global 502 errors lasted about 30 minutes, and traffic fell 82% at the worst point. Cloudflare concluded that its testing and deployment controls were insufficient.
This is a striking example of a small change with a global failure domain: the rule’s deployment scope exceeded the scope in which its effects were safely tested. The published figures establish the traffic impact and initial outage duration; they do not, in the material cited here, provide a full detection, rollback, or customer-communication timeline.
Controls to prioritize: validate high-cost rules against representative traffic, canary them on a small slice, and expand only after health signals remain acceptable. A global configuration should have an automated, independently operable path back to a known-good version.
8. Google Cloud networking (2020)
The historical incident index identifies Google networking incidents in 2020 as major cloud events and supports treating this as a networking and control-plane dependency case. Exact impact figures and a detailed trigger or recovery account are not established here.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
Lesson: networking is a dependency shared by otherwise unrelated products. Review what happens when name resolution, routing, or management functions are impaired, and define which customer-facing functions must remain available without those control-plane operations.
9. Fastly’s CDN disruption (June 8, 2021)
The incident is widely cited as an approximately one-hour disruption to internet services, and it illustrates concentration in edge delivery alongside a latent-bug or configuration failure mode. The available account here does not establish the exact trigger, affected-customer count, detection sequence, rollback details, or communications timeline.
Lesson: a CDN can sit in front of many unrelated sites, so a provider-level edge failure can look like many simultaneous application failures. For critical services, understand how to bypass or replace edge delivery safely, while recognizing that an alternate path only helps if it is maintained and tested before an incident.
10. AWS us-east-1 networking (December 7, 2021)
The historical incident index identifies a widespread AWS regional networking event. Exact duration and blast-radius figures require a detailed primary incident report and are not established in the account summarized here.
Lesson: products that appear independent may still share a provider region’s network and management dependencies. A second availability zone inside the same region may not protect against a regional networking problem; resilience depends on the failure domain the alternate path actually escapes.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What the outages have in common
- Shared dependencies defeat apparent independence. A provider’s region or edge network can affect many products because they share control planes, identity, DNS, storage, or networking. Separate product names do not prove separate failure domains.
- Deployment scope should match tested scope. Cloudflare’s WAF rule went global in one step, while Google’s topology incident showed how a manual action could bypass intended per-zone isolation. Use canaries, staged rollout, and enforced cell or zone boundaries.
- Detection is part of the system. Google found that a load-balancing health-detection feature did not identify unhealthy backends. Monitoring should reflect user-visible outcomes, not only whether components report themselves healthy.
- Backups must be restorable. GitLab’s experience shows why backup jobs, copies, and restore exercises need independent verification. A recovery plan should specify who can restore, from where, and how much data loss and downtime are acceptable.
- Postmortems must produce controlled change. Google SRE engineers Adrian Hilton and Gwendolyn Stockman write that blameless postmortems help teams learn and make systems more reliable. Findings matter when owners track prioritized corrective actions to completion.
How to protect an application from a cloud outage
Start with the service users depend on, then work outward through its dependencies. Multi-region deployment can reduce some risks, but it does not automatically prevent downtime: identity, DNS, deployment tooling, data replication, or failover decisions may remain shared. The design must identify which failure domain is being escaped and prove that the alternate path can serve traffic.
- Define user-facing service objectives. Set service-level indicators and objectives for real user outcomes, such as successful requests and latency. Use error-budget accounting to decide when feature changes should yield to reliability work.
- Draw the dependency and failure-domain map. Include regional and global control planes, storage, identity, DNS, network paths, third-party edge services, and the systems needed to deploy or recover. Mark dependencies shared by supposedly independent regions.
- Limit the blast radius of changes. Deploy configuration and code through canaries or staged regions. Use per-zone or per-cell permissions and limits, and prevent a single automated or manual operation from changing every cell at once.
- Automate safe rollback. Define health gates based on customer-facing signals and retain a known-good configuration. Verify that rollback remains possible if the primary management plane is degraded.
- Exercise failover and restore. Practice regional failover, traffic shifting, and database recovery. Test that backups are complete, isolated, and restorable, and record achieved recovery time and data loss against objectives.
- Prepare incident communications. Establish who declares an incident, where status is published, how often updates are issued, and how customer support receives the same operational facts. The incident accounts summarized above do not provide complete communication timelines for every event; teams should make this an explicit readiness exercise.
- Track blameless postmortem actions. Record impact, timeline, contributing conditions, detection gaps, recovery decisions, and concrete prevention or mitigation actions. Assign owners and due dates, then verify the controls through testing.
Does multi-region really prevent downtime?
It can help when the failed component is genuinely regional and the application has independent, tested capacity elsewhere. It will not help if both regions depend on the same global control plane, credentials, DNS configuration, data store, network service, or faulty rollout. Nor is failover useful if the replica is stale, the application cannot write there, or traffic-shifting procedures have never been exercised. Design for explicit failure domains and prove the failover path before relying on it.
Quick Recap
What should an SRE postmortem include?
- A clear summary of user impact, affected services, and the time window.
- A timestamped sequence covering first symptoms, detection, escalation, mitigation, restoration, and communications.
- The initiating event and contributing conditions, including shared dependencies and controls that failed or were bypassed.
- What monitoring showed—and what it failed to show—at the time.
- Recovery choices, rollback or restore results, and any remaining risk.
- Blameless corrective actions ranked by impact, each with an owner, a target date, and a way to verify completion.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

