The Update That Stopped the World
01:14 · Point of failure
On 19 July 2024, a flawed CrowdStrike content update crashed Windows machines around the world. The failure was not an attack: it was a release path that trusted a validator which had accepted the configuration it was meant to stop.

Editions
- Watch
- The episode on YouTube1:27
- Read
- The Largest IT Outage in History Was a File Full of Zerostechnical debrief · zof.ai
Incident
- System
- CrowdStrike Falcon Sensor Rapid Response Content delivery path
- Date
- 19 July 2024
- Location
- Worldwide
- Toll
- Microsoft's estimate was about 8.5 million affected Windows devices.
- Topics
- deployment · security software · configuration · validation · resilience
Point of failure
A bug in the Content Validator let a problematic Template Instance pass validation, so the release process treated a failing configuration as safe to deploy.
Transcript
A routine update
Eight and a half million computers crashed at the same time. Planes grounded. Hospitals on paper. TV stations dark. This is The Point of Failure, episode one.
July 19th, 2024. CrowdStrike, one of the biggest security companies on Earth, pushed a routine update to its software. The kind of update that goes out all the time. Nobody thinks twice.
The deepest layer
But this update has a defect. And the software it updates sits in the deepest and most trusted layer of Windows. So when it fails, it does not crash an app. It crashes an entire system.
Blue screens
Within minutes, blue screens ripple across the entire planet. Airports in every time zone, hospital systems, broadcasters live on air, banks, supermarket checkouts, all down.
Not from an attack. From an update meant to protect them.
And here is the brutal part. The fix could not be pushed remotely, because the machines could not stay on long enough to receive it.
Millions of computers had to be fixed by hand. One at a time.
The estimated damage: billions. One airline alone said it lost around $500 million. All of it traced back to a single file that shipped without catching one defect.
What actually broke
Point of failure
The biggest IT outage in human history was not a hack. It was a test that did not happen.
Every failure has a story. Every story was preventable. I'm Kevin.
Sources
CrowdStrike's preliminary account (24 July 2024) of the Content Validator, the problematic Template Instance, the crash path, and its "How Do We Prevent This From Happening Again?" measures: "Due to a bug in the Content Validator, one of the two Template Instances passed validation despite containing problematic content data." Describes the release as "a content configuration update for the Windows sensor" made "as part of regular operations".
External Technical Root Cause Analysis — Channel File 291
The full root cause analysis (dated 2024-08-06 on every page), which names the missing test: "In summary, it was the confluence of these issues that resulted in a system crash: the mismatch between the 21 inputs validated by the Content Validator versus the 20 provided to the Content Interpreter, the latent out-of-bounds read issue in the Content Interpreter, and the lack of a specific test for non-wildcard matching criteria in the 21st field." Findings 3 to 6 ("Template Type testing should cover a wider variety of matching criteria", "The Content Validator contained a logic error", "Template Instance validation should expand to include testing within the Content Interpreter", "Template Instances should have staged deployment") set out the testing, validation and deployment gaps that let the Template Instance ship, and are the basis for "a test that did not happen".
Falcon Content Update Remediation and Guidance Hub
The remediation guidance, updated in place through 6 August 2024: hosts "unable to stay online to receive the Channel File update", the branch "If the host crashes again on reboot:" leading to bootable recovery images or a manual process, the instruction for cloud volumes to "Locate the files matching “C-00000291*.sys”, and delete them", and "Bitlocker-encrypted hosts may require a recovery key." Also carries the plain statement "This was not a cyberattack." It is the record of why the fix was applied by hand, one machine at a time.
Helping our customers through the CrowdStrike outage
Microsoft's estimate of 8.5 million affected Windows devices.
A worldwide IT outage disrupted airlines, banks, hospitals and businesses today
Contemporary reporting of the day: "Thousands of flights canceled worldwide, including more than 2,600 flights that begin or end at U.S. airports"; "employees of airlines, banks, hospitals and emergency services staring at the dreaded blue screen of death"; the sectors dependent on the one vendor: "hospitals, retail businesses, broadcasters, ports, government agencies and emergency responders"; "CrowdStrike says this was not a cyberattack"; and on the fix, "You often have to have physical access to the device."
Global technology outage disrupts flights, banks and companies around the world
A second, independent report of the same disruption across airports, hospitals and broadcasters.
Delta CEO says CrowdStrike-Microsoft outage cost the airline $500 million
Delta's chief executive putting the airline's own cost at $500 million over five days, the "one airline alone" in the episode. This is the airline's statement, not an adjudicated figure.
Delta's CEO says the CrowdStrike outage cost the airline $500 million in 5 days
NPR's report of the same statement, which it attributes to the CNBC appearance ("Speaking on CNBC on Wednesday"): "Half a billion dollars in five days." A second outlet's account of the same interview, not an independent figure.
CrowdStrike's Impact on the Fortune 500
An insurer's estimate of $5.4 billion in direct losses to US Fortune 500 companies alone, excluding Microsoft, which is the basis for "the estimated damage: billions". No precise global total exists; the episode does not state one.
Related
- 03
The Song That Broke YouTube's Math
Gangnam Style's view count approached the maximum value of a signed 32-bit integer. YouTube widened the counter before it overflowed, a friendly example of a failure mode that is much less friendly in systems that cannot be patched in time.
- 04
The Bank That Sent $900 Million by Accident
Citibank meant to send about $7.8 million of interest on Revlon's loan. Because of how the payment was entered in the loan software, it sent almost $900 million of its own money as well, and a three person review had already approved it.