In short
- In all four outages the trigger was an ordinary change: a content update, a policy row, an automation run, a database permission.
- What made them large was reach. The change landed in many places at once with no automatic gate in between.
- Roll out data and configuration in rings with an automated health gate, the same way you roll out code.
- Treat generated config as untrusted input: validate it, keep the last good copy and make writes conditional.
- Limit the blast radius with cells, and alert on error budget burn so you find out in minutes.
Four outages, one shape
Between July 2024 and November 2025, four of the largest outages in cloud and security did not start with an attack or a hardware failure. Each started with a normal change, and each reached far more customers than anyone intended, quickly. The postmortems are public and they are worth reading in full.
CrowdStrike
19 July 2024
A content update for the Falcon sensor defined 21 input fields while the sensor code supplied 20. The sensor read past the end of an array and crashed Windows machines. The update was live for about 1 hour 22 minutes before it was reverted.
Read the reportGoogle Cloud
12 June 2025
A quota policy feature shipped on 29 May without error handling and without a feature flag. On 12 June a policy row with blank fields was written to a regional database, replicated, and crashed Service Control, which sits in front of Google Cloud's APIs.
Read the reportAWS
19 and 20 October 2025
A race condition in the automation that manages DNS for DynamoDB left the regional endpoint with no IP addresses. DynamoDB was reachable again by 2:40 AM PDT, but EC2 was not fully recovered until 1:50 PM.
Read the reportCloudflare
18 November 2025
A database permission change made a query return duplicate rows, a generated file grew past a limit of 200 features, and the proxy code hit an unhandled error. Core traffic was largely flowing again by 14:30 UTC and all systems by 17:06.
Read the reportThe details differ. The shape is the same: a change that was safe in the place it was tested and unsafe in the place it landed.
Blast radius is a design decision
If every change goes to every customer at once, your blast radius is 100% before anything has gone wrong. Progressive delivery changes that default. You ship to a small ring first, wait for an automated check, and only then widen.
Swipe sideways to see the whole figure
It works the same whether you use Argo Rollouts or Flagger on Kubernetes, a feature-flag service or deploy stages in a pipeline: an internal ring, a small share of real traffic, a bigger share, then everyone. The gate matters more than the rings. A person watching a dashboard at 3 a.m. is a slow gate, and a gate that compares error rate and latency against the old version is a fast one.
Data and configuration need the same treatment as code. Google's report says the bad policy data was replicated across regions, and its fix list includes incremental replication with time to validate. CrowdStrike's report says template instances should have staged deployment. Both are the code-rollout lesson applied to content.
Treat generated config as untrusted input
Cloudflare's Bot Management system has a limit of 200 machine learning features at runtime, well above the roughly 60 in use, because memory for them is preallocated. A database permission change made the query that builds the feature file return duplicate rows, and the file grew past the limit. The proxy code that loaded it hit an unhandled error and returned HTTP 5xx.
The file was rebuilt every five minutes, and only part of the database cluster had the new permissions, so the network swung between good and bad files. That is why the failure came and went. The figure replays that pattern with two kinds of proxy.
Swipe sideways to see the whole figure
A program that reads generated config should validate it before use, with checks on schema, size and count, keep the last known good version, and degrade or fail open instead of crashing. Cloudflare's own list of fixes starts with hardening how it ingests its generated configuration files, and adds more global kill switches.
let active = loadFromDisk(); // last known good
function onNewConfig(raw) {
const next = parse(raw);
if (!schemaOk(next) || next.features.length > LIMIT) {
metrics.increment("config.rejected");
alert("bad config, still serving version " + active.version);
return; // keep serving with the old one
}
active = next;
}Race conditions hide in automation
AWS's DynamoDB outage began in DNS management. A planner produces plans for the service's endpoints, and several enactors apply them. One enactor was slow, another applied a newer plan and started cleaning up old ones, and the slow one then wrote its old plan over the new one. The cleanup deleted that old plan, and with it every IP address for the regional endpoint.
Swipe sideways to see the whole figure
The first enactor's check, is my plan newer than what is live, was true when it ran and false by the time it wrote. A check followed by a write leaves a gap. Closing it means making the write itself conditional on the version you saw, with compare-and-swap or a conditional update, so a stale writer fails instead of winning.
Recovery from the DNS fix was a load event too. In AWS's account, the EC2 lease manager fell into what it calls congestive collapse from restart queues, and engineers throttled requests to get out. Network Load Balancer health checks failed on instances that had not yet received their network state, and automatic failover kept removing capacity until engineers turned it off. EC2 was not fully recovered until 1:50 PM, almost eleven and a half hours after the DNS record was repaired.
Make failures small, then make them visible
Cells are the architectural version of rings. Customers are split across independent copies of the stack, each with its own share of users, and a bad deploy or bad data lands in one cell at a time. AWS describes the pattern in its Well-Architected material on reducing the scope of impact with cell-based architecture.
Swipe sideways to see the whole figure
Once the blast radius is small, speed of detection decides the rest. An alert on raw error counts is either noisy or late. Google's SRE Workbook recommends alerting on how fast the error budget is burning instead: page at 14.4 times the planned burn over one hour, which spends 2% of a 30-day budget, page at 6 times over six hours, and open a ticket at 1 times over three days.
Swipe sideways to see the whole figure
With a 99.9% target, a failure that hits 10% of requests burns the whole month's budget in 7.2 hours, and the first alert fires in under ten minutes. A failure that hits 0.5% never pages anyone, and that is correct: it is a ticket, and it takes six days to use the budget.
A checklist for your own platform
When I review a platform, these are the questions I start with.
- Can any change reach more than 5% of production without passing an automated health gate?
- Are config and data deployed through the same pipeline as code, with the same validation?
- Does every kill switch exist, and has anyone used one this quarter?
- Is there a last-known-good fallback for every generated file the system reads?
- Do retries use jitter and a budget?
- Do alerts fire on burn rate rather than on raw error counts?
- When did you last roll back on purpose, in production, on a quiet day?
Sources
- CrowdStrike, Channel File 291 Incident Root Cause Analysis
- Google Cloud incident report, 12 June 2025
- AWS, Summary of the Amazon DynamoDB service disruption in the Northern Virginia (us-east-1) Region
- Cloudflare, Cloudflare outage on November 18, 2025
- Google SRE Workbook, Alerting on SLOs
- AWS Well-Architected, Reducing the scope of impact with cell-based architecture
The animated figures are simplified illustrations that use the numbers stated in the text. Where a figure is modelled on a real incident, the sources above are the full accounts.
