Skip to content
The Point of Failure

Episode 19 · Cloud and Internet Backbone · 2:19

The Morning the Internet's Address Book Failed

00:35 · Point of failure

On 19 and 20 October 2025 a latent race condition between two components of AWS's DNS automation left the DynamoDB endpoint in us-east-1 with an empty DNS record. Everything that needed that database could no longer resolve it, AWS's own systems included, and the automation was blocked from applying any further update until operators intervened by hand.

Incident

Topics
automation · dns · race condition · shared dependencies
Point of failure
The endpoint's DNS records were maintained by automation whose enactors could race each other, and the design allowed an empty record to be written and then blocked every later update, so the automation could not repair what it had produced.

Transcript

310 words · 2 min read

No address

One database service in one region of one cloud lost its internet address for a while. You probably felt it, wherever you live. This is The Point of Failure, episode nineteen.

The oldest region

AWS, us-east-1: the oldest, busiest patch of cloud on earth. Inside it, DynamoDB, a database so foundational that AWS's own services are built on it. And in front of DynamoDB, like in front of everything on the internet, sits DNS: the address book that turns a name into a location.

One planner, many enactors

Point of failure

That address book was maintained by automation: one component planning updates, others applying them. For years, fine. Then, per Amazon's own postmortem, the components interacted in a sequence nobody had hit before, a race, and the result was surreal: the address record for the region's DynamoDB endpoint came out... empty. Not wrong. Blank. The service was healthy. It just had no address.

The cascade

Now the cascade, because us-east-1 is a load bearing wall of the consumer internet. Systems that needed that database stumbled, including AWS's own machinery: new servers struggled to launch, load balancers second guessed healthy targets. Apps, banks, smart homes, games around the world spent the day degraded. And the automation that managed the addresses could not repair the blank it had created. Humans had to reach in by hand.

Faster than anyone can watch

AWS published everything: the race, the fix, the guardrails added. Credit, as always, for the homework. But the day's lesson sits above any one company: we have built automation that operates faster than humans can watch, and the failure modes are now sequences no human ever rehearsed. The blank page was legal. The system just never imagined writing it.

Every failure has a story. Every story was preventable. I'm Kevin. See you at the next one.

Sources

4 sources

  1. Summary of the Amazon DynamoDB Service Disruption in the Northern Virginia (US-EAST-1) Region

    Amazon Web Services · 2025

    AWS dates the disruption to 19 and 20 October 2025, beginning at 11:48 PM PDT on 19 October and ending at 2:20 PM PDT on 20 October, in three distinct periods of impact rather than one continuous day. It records the cause as a latent race condition in the DynamoDB DNS management system that left an incorrect empty record for the regional endpoint and that the automation failed to repair, requiring manual operator intervention. On remediation its wording is future tense: the DNS Planner and DNS Enactor automation was disabled worldwide, and the race condition fix and the additional protections are what AWS said it would do before re-enabling it. The failure was in resolving the endpoint, not in the database itself.

  2. In the works – New Availability Zone in Maryland for US East (Northern Virginia) Region

    AWS News Blog · 2025

    AWS states that Northern Virginia was the first region it launched, which is what backs the oldest. No AWS document states that it is the busiest.

  3. AWS services recover after daylong outage hits major sites

    CNBC · 2025

    Lloyds Banking Group confirmed that some of its services were affected. The rest of the consumer picture in this report rests on user reports rather than statements from the companies named.

  4. Amazon identifies the issue that broke much of the internet, says AWS is back to normal

    TechCrunch · 2025

    Filed the day after, from the status page rather than the full summary, so its mitigation time differs slightly from the figures above.

  1. 12

    The Expired Certificate That Silenced 11 Countries

    On 6 December 2018, an expired certificate in Ericsson core-network software disrupted operators across 11 countries. O2 and SoftBank were among the networks affected by the same dated dependency.

  2. 06

    The Company That Locked Itself Out of Its Own Building

    On 4 October 2021 a maintenance command disconnected Facebook's backbone, its DNS servers responded by withdrawing the routes that tell the internet where Facebook is, and the same outage took down the internal tools and building access the engineers needed to put it back.

  3. 17

    Sixteen Hours of Tay

    On 23 March 2016 Microsoft released Tay, a chatbot designed to learn from the people who talked to it, onto Twitter. Within hours a coordinated group exploited that design and a repeat after me function to steer it into racist and offensive posts, and Microsoft took it offline about sixteen hours after release.