DevOps and reliabilityTechnical article

How one config change takes down production, and how to contain yours

CrowdStrike, Google Cloud, AWS and Cloudflare each had a major outage between July 2024 and November 2025, and none started with an attack. Each began with a routine change that reached too many places too fast. Here is what happened, what it teaches and what to put in place.

In short

  • In all four outages the trigger was an ordinary change: a content update, a policy row, an automation run, a database permission.
  • What made them large was reach. The change landed in many places at once with no automatic gate in between.
  • Roll out data and configuration in rings with an automated health gate, the same way you roll out code.
  • Treat generated config as untrusted input: validate it, keep the last good copy and make writes conditional.
  • Limit the blast radius with cells, and alert on error budget burn so you find out in minutes.

Four outages, one shape

Between July 2024 and November 2025, four of the largest outages in cloud and security did not start with an attack or a hardware failure. Each started with a normal change, and each reached far more customers than anyone intended, quickly. The postmortems are public and they are worth reading in full.

CrowdStrike

19 July 2024

A content update for the Falcon sensor defined 21 input fields while the sensor code supplied 20. The sensor read past the end of an array and crashed Windows machines. The update was live for about 1 hour 22 minutes before it was reverted.

Read the report

Google Cloud

12 June 2025

A quota policy feature shipped on 29 May without error handling and without a feature flag. On 12 June a policy row with blank fields was written to a regional database, replicated, and crashed Service Control, which sits in front of Google Cloud's APIs.

Read the report

AWS

19 and 20 October 2025

A race condition in the automation that manages DNS for DynamoDB left the regional endpoint with no IP addresses. DynamoDB was reachable again by 2:40 AM PDT, but EC2 was not fully recovered until 1:50 PM.

Read the report

Cloudflare

18 November 2025

A database permission change made a query return duplicate rows, a generated file grew past a limit of 200 features, and the proxy code hit an unhandled error. Core traffic was largely flowing again by 14:30 UTC and all systems by 17:06.

Read the report

The details differ. The shape is the same: a change that was safe in the place it was tested and unsafe in the place it landed.

Blast radius is a design decision

If every change goes to every customer at once, your blast radius is 100% before anything has gone wrong. Progressive delivery changes that default. You ship to a small ring first, wait for an automated check, and only then widen.

Figure 1Blast radius: everyone at once, or in rings
Push to everyone at onceone step, no waiting between groupsRelease in ringseach ring waits for a health gatefailing now 100% · worst 100%failing now 0% · worst 4%waiting for a person to notice and roll backgate failed, rolled back automatically1%passed5%failed25%waiting100%waiting

Swipe sideways to see the whole figure

The change is
Each square is a group of customers. Grey runs the old version, green runs the new version and is healthy, red runs the new version and is failing. A gate is an automatic check of error rate and latency between rings.

It works the same whether you use Argo Rollouts or Flagger on Kubernetes, a feature-flag service or deploy stages in a pipeline: an internal ring, a small share of real traffic, a bigger share, then everyone. The gate matters more than the rings. A person watching a dashboard at 3 a.m. is a slow gate, and a gate that compares error rate and latency against the old version is a fast one.

Data and configuration need the same treatment as code. Google's report says the bad policy data was replicated across regions, and its fix list includes incremental replication with time to validate. CrowdStrike's report says template instances should have staged deployment. Both are the code-rollout lesson applied to content.

Treat generated config as untrusted input

Cloudflare's Bot Management system has a limit of 200 machine learning features at runtime, well above the roughly 60 in use, because memory for them is preallocated. A database permission change made the query that builds the feature file return duplicate rows, and the file grew past the limit. The proxy code that loaded it hit an unhandled error and returned HTTP 5xx.

The file was rebuilt every five minutes, and only part of the database cluster had the new permissions, so the network swung between good and bad files. That is why the failure came and went. The figure replays that pattern with two kinds of proxy.

Figure 2A generated file that every proxy loads
Databasebuilds the filereturns duplicate rowsFeature file v4every 5 minuteslimit 200about 60 featuresproxy fleetv4 is fine, so every proxy loads itTraffic flows.The last six files, one every five minutesv1goodv2goodv3bad filev4goodv5bad filev6bad fileOnly part of the database cluster had the new permissions yet,so good and bad files alternate and the failure comes and goes.The fleet flips between healthy and failing with each new file.

Swipe sideways to see the whole figure

The proxy
Modelled on the Cloudflare incident of 18 November 2025, simplified. A good file has about 60 features and the proxy has room for 200. A bad file is over the limit. Compare what the fleet does when it trusts the file and when it checks it first.

A program that reads generated config should validate it before use, with checks on schema, size and count, keep the last known good version, and degrade or fail open instead of crashing. Cloudflare's own list of fixes starts with hardening how it ingests its generated configuration files, and adds more global kill switches.

let active = loadFromDisk(); // last known good

function onNewConfig(raw) {
  const next = parse(raw);
  if (!schemaOk(next) || next.features.length > LIMIT) {
    metrics.increment("config.rejected");
    alert("bad config, still serving version " + active.version);
    return; // keep serving with the old one
  }
  active = next;
}

Race conditions hide in automation

AWS's DynamoDB outage began in DNS management. A planner produces plans for the service's endpoints, and several enactors apply them. One enactor was slow, another applied a newer plan and started cleaning up old ones, and the slow one then wrote its old plan over the new one. The cleanup deleted that old plan, and with it every IP address for the regional endpoint.

Figure 3Two writers and a stale check
The record is empty and each enactor now sees a state it does not expectPlannerEnactor AEnactor Bplan 100plan 101applying plan 100, slowwrites plan 100plan 101 livecleanup of older plansdeletes plan 100THE DNS RECORD FOR THE SERVICE ENDPOINTNo IP addressesNothing can find the service, and the automation cannot repair what it did not expect.The check “is my plan newer than what is live?” was true when A ran it. By the time A wrote, it was false.A write that is conditional on the version it saw would have failed instead of winning.

Swipe sideways to see the whole figure

Simplified from AWS's account of the DynamoDB outage on 19 and 20 October 2025. The details of the real system are richer. The shape is the usual one: a check, a gap, then a write that acts on a check that is no longer true.

The first enactor's check, is my plan newer than what is live, was true when it ran and false by the time it wrote. A check followed by a write leaves a gap. Closing it means making the write itself conditional on the version you saw, with compare-and-swap or a conditional update, so a stale writer fails instead of winning.

Recovery from the DNS fix was a load event too. In AWS's account, the EC2 lease manager fell into what it calls congestive collapse from restart queues, and engineers throttled requests to get out. Network Load Balancer health checks failed on instances that had not yet received their network state, and automatic failover kept removing capacity until engineers turned it off. EC2 was not fully recovered until 1:50 PM, almost eleven and a half hours after the DNS record was repaired.

Make failures small, then make them visible

Cells are the architectural version of rings. Customers are split across independent copies of the stack, each with its own share of users, and a bad deploy or bad data lands in one cell at a time. AWS describes the pattern in its Well-Architected material on reducing the scope of impact with cell-based architecture.

Figure 4Cells shrink the blast radius
One cell fails: 12.5% of users feel it240 users, 8 cells, about 30 users in eachThe failing cell moves every few seconds so you can see each one. One cell is the same as having no cells.

Swipe sideways to see the whole figure

Each box is a full copy of the stack with its own share of users. A bad deploy or bad data lands in one cell at a time. The price is real: every cell is something to run, patch and pay for.

Once the blast radius is small, speed of detection decides the rest. An alert on raw error counts is either noisy or late. Google's SRE Workbook recommends alerting on how fast the error budget is burning instead: page at 14.4 times the planned burn over one hour, which spends 2% of a 30-day budget, page at 6 times over six hours, and open a ticket at 1 times over three days.

Figure 5How fast the error budget burns
30-day budget at a 99.9% targetburn rate 100×After 7.2 hours of this · budget gone, the target is missedall of it is gone in 7.2 hoursWHEN THE ALERTS FIREPage14.4× over 1 hourspends 2% of the budget in that windowfiring after about 9 minutesPage6× over 6 hoursspends 5% of the budget in that windowfiring after about 22 minutesTicket1× over 3 daysspends 10% of the budget in that windowfiring after about 43 minutesFires after about = long window × alert rate ÷ burn rate, for a failure that starts suddenly and stays.

Swipe sideways to see the whole figure

A 99.9% target leaves 0.1% of requests, about 43 minutes of a full outage, as the 30-day budget. Burn rate is the failure share divided by 0.1%. The alert rules are the starting points in Google's SRE Workbook.

With a 99.9% target, a failure that hits 10% of requests burns the whole month's budget in 7.2 hours, and the first alert fires in under ten minutes. A failure that hits 0.5% never pages anyone, and that is correct: it is a ticket, and it takes six days to use the budget.

A checklist for your own platform

When I review a platform, these are the questions I start with.

  1. Can any change reach more than 5% of production without passing an automated health gate?
  2. Are config and data deployed through the same pipeline as code, with the same validation?
  3. Does every kill switch exist, and has anyone used one this quarter?
  4. Is there a last-known-good fallback for every generated file the system reads?
  5. Do retries use jitter and a budget?
  6. Do alerts fire on burn rate rather than on raw error counts?
  7. When did you last roll back on purpose, in production, on a quiet day?

Sources

  1. CrowdStrike, Channel File 291 Incident Root Cause Analysis
  2. Google Cloud incident report, 12 June 2025
  3. AWS, Summary of the Amazon DynamoDB service disruption in the Northern Virginia (us-east-1) Region
  4. Cloudflare, Cloudflare outage on November 18, 2025
  5. Google SRE Workbook, Alerting on SLOs
  6. AWS Well-Architected, Reducing the scope of impact with cell-based architecture

The animated figures are simplified illustrations that use the numbers stated in the text. Where a figure is modelled on a real incident, the sources above are the full accounts.

The service behind this articleDevOps Architecture & Platform ReliabilityRelated case studyPrologue Partnerships: an author's book, turned into a paid AI coach
Aslam Sarfraz profile photo

Aslam Sarfraz

AI, DevOps and full-stack engineer

More articles

AI systems · 10 min read

How AI backends handle millions of requests without falling over

A chat answer keeps a GPU, a block of memory and a connection busy for many seconds. What that does to a backend, how serving systems cope, and what a startup should copy.

Full-stack engineering · 8 min read

Why startup backends break at the first big traffic spike

Ticketmaster took 3.5 billion requests in one sale and Segment merged more than 140 services into one. The failures behind spikes are rarely exotic, and each one has arithmetic you can do before launch.

LinkedIn outreach · 7 min read

Designing a LinkedIn outreach pipeline: a fixed capacity and an unseen limit

Outreach is a pipeline with a fixed capacity, a limit nobody publishes and very small samples. The planning numbers I use and the engineering behind them.