Designing for a Region Outage: What Recent us-east-1 Incidents Teach
What the 2021, 2023 and 2025 AWS us-east-1 outages teach about control plane dependencies, static stability, multi-Region failover and resilience governance.
On October 20, 2025, a race condition in DynamoDB's DNS automation left the regional endpoint in us-east-1 with an empty DNS record. The DNS was restored in under three hours. EC2 instance launches in the Region did not return to normal for another eleven hours, and the event ended 14 hours and 32 minutes after it began.
That gap between the trigger and full recovery is the most useful thing an engineering leader can learn from recent us-east-1 incidents. This guide reviews what AWS's own post-event summaries say happened, draws the design lessons, and sets out the architecture, governance and testing an enterprise needs to keep serving customers when a Region, or a service inside it, is impaired.
Executive summary
- Most Region "outages" are dependency failures. In October 2025 a DNS fault in one service cascaded into EC2 launches, Network Load Balancer health checks, Lambda, STS and container services.
- Control planes fail first and recover last. In both the 2021 and 2025 events, running workloads largely kept working while the ability to launch, change or scale resources was impaired for hours.
- Recovery that needs new resources is the riskiest recovery. Design for static stability: capacity, DNS records and credentials are in place before the incident, and failover is a data plane action.
- Some global services have control planes in us-east-1. IAM, Organizations, Route 53 record changes and CloudFront configuration all depend on it, even for workloads that run elsewhere.
- Multi-Region is a business decision, not a default. Tier workloads by business impact, and reserve active-active or warm standby for the services that justify the cost and complexity.
| Workload profile | Recommended posture | Why |
|---|---|---|
| Revenue or safety critical, minutes of downtime are material | Multi-Region active-active or hot standby, with pre-provisioned capacity | Failover is a routing decision, not a build |
| Important, hours of downtime are tolerable | Warm standby in a second Region, scaled down, with a rehearsed scale-up | Lower cost, but scale-up depends on the standby Region's control plane |
| Internal or batch workloads | Multi-AZ in one Region, plus pilot light or backup and restore | Covers the far more common single-AZ failure at low cost |
| Anything regulated | The posture above, plus documented decision rights and test evidence | Auditors ask for proof, not architecture diagrams |
1. What actually happened in us-east-1
Four recent events show the range of failures a Region can experience. All times below are as AWS or the cited report published them.
| Date | Trigger | Scope | Duration |
|---|---|---|---|
| December 7, 2021 | An automated scaling activity overwhelmed devices connecting AWS's internal network to the main network | Regional; control planes and several data paths | About 7 hours for network recovery; some services longer |
| June 13, 2023 | A latent defect in Lambda's frontend fleet when one cell crossed a new capacity threshold | Regional; Lambda and dependent services | 3 hours 48 minutes |
| October 19–20, 2025 | A race condition in DynamoDB's DNS management left the regional endpoint with an empty record | Regional; cascaded across many services | 14 hours 32 minutes |
| May 7–8, 2026 | Rising temperatures in one data center led to loss of power for affected hardware | One Availability Zone (use1-az4); EC2 and EBS | Recovery continued into the next day |
October 2025: one DNS record, fourteen hours
AWS's post-event summary describes two independent DNS Enactors racing each other. A delayed Enactor applied an older plan after a newer one was already in place, and the newer plan's cleanup then deleted it, removing every IP address for dynamodb.us-east-1.amazonaws.com. Automation could not repair the inconsistent state, and operators had to intervene.
The DNS record was restored by 2:25 AM PDT. The damage was in what depended on DynamoDB. EC2's internal DropletWorkflow Manager lost its leases, and when DynamoDB returned, the backlog drove it into congestive collapse. Network configuration then queued behind it, new instances launched without connectivity, and Network Load Balancer health checks began removing healthy capacity. Each step had its own recovery time.
<!-- region-incident-timeline -->December 2021: running workloads survived, changes did not
In 2021, AWS reported that running EC2 instances, existing load balancers, existing DNS answers and Lambda invocations were largely unaffected. What failed were the control planes: EC2 APIs, Route 53 record changes, console sign-in and STS. Teams that needed to launch capacity, edit DNS or sign in to the console to recover were blocked for hours. AWS also reported that network congestion prevented its Service Health Dashboard tooling from failing over to a standby Region.
June 2023: a cell boundary held, but not everywhere
A Lambda frontend cell crossed a capacity threshold it had never reached before and hit a latent defect. Other cells were unaffected, which shows cellular architecture working. Dependent services still felt it: STS errors throttled SAML federation sign-in, EventBridge delivery was delayed, and the us-east-1 console returned errors while consoles in other Regions were fine.
May 2026: the common case is still one Availability Zone
The most recent event was not regional. A thermal incident in one data center in use1-az4 impaired EC2 instances and EBS volumes on the affected hardware. Workloads spread across Availability Zones with spare capacity rode through it. Single-AZ failures remain far more frequent than Region-wide ones, so multi-AZ static stability is the first investment, before any multi-Region spend.
2. Five design lessons from the incidents
- Map dependencies, not just services. Your workload may not call DynamoDB directly, but in October 2025 EC2 launches, load balancer health and Lambda did. Ask what your recovery path depends on, two and three levels down.
- Assume control planes are unavailable during an incident. Launching instances, changing DNS records, creating IAM roles and editing CloudFront origins are all control plane actions. A runbook built on them may not run.
- Plan for recovery to take longer than the trigger. Backlogs, retries and throttles extend impact long after the root cause is fixed. Measure your RTO against the full tail, not the headline fix time.
- Know your hidden us-east-1 dependencies. AWS documents that IAM, Organizations and Account Management have control planes in us-east-1, and that Route 53 record, hosted zone and health check changes, CloudFront configuration and S3 bucket creation depend on it. A workload in eu-west-1 can still be affected.
- Keep your response tooling outside the blast radius. Monitoring, paging, runbooks, status pages and break-glass access should not live only in the Region they protect. AWS's own dashboard and support tooling were affected in 2021, 2023 and 2025.
3. Control plane, data plane and static stability
AWS's Builders' Library defines a statically stable system as one that keeps doing what it was doing when a dependency is impaired, even if it cannot receive updates. The control plane makes changes; the data plane keeps existing resources running. The data plane is simpler and built for higher availability, so recovery that only uses the data plane has fewer ways to fail.
Turn on the impairment below to see which common recovery actions still work.
<!-- region-plane-simulator -->AWS's fault isolation guidance translates this into specific rules for recovery paths:
- Credentials: Use Regional STS endpoints, not the global endpoint. Newer SDK major versions default to Regional endpoints; for older tools such as the AWS CLI v1 and the AWS SDK for Java 1.x, set
sts_regional_endpoints = regional. - DNS: Do not create or edit Route 53 records, hosted zones or health checks to fail over. Pre-create failover records with health checks, or use Application Recovery Controller (ARC) routing controls, whose Regional cluster endpoints are on the data plane.
- Edge: Do not change CloudFront origins during an incident. Configure origin failover groups in advance.
- Capacity: Pre-provision load balancers, buckets and enough compute in the recovery location. If recovery needs control plane data, cache it in a data plane store such as Parameter Store, DynamoDB or S3.
- Access: Allow SAML sign-in through more than one Regional endpoint, and keep tested break-glass access in case your identity provider is also affected.
4. Choosing a multi-Region strategy
AWS's disaster recovery whitepaper describes four strategies in order of cost and complexity. It does not assign fixed RTO or RPO numbers, and neither should you without testing. The column that matters most after these incidents is how much each strategy depends on control plane actions at the moment of failover.
| Strategy | What runs in the recovery Region | Control plane work at failover | Cost and complexity |
|---|---|---|---|
| Backup and restore | Backups and infrastructure as code only | High: provision everything, restore data | Lowest |
| Pilot light | Data replication and core services; application servers off | High: start and scale application tier | Low |
| Warm standby | A scaled-down, fully working copy | Medium: scale up with Auto Scaling | Medium |
| Hot standby or active-active | Full capacity, ready or already serving | Low: shift traffic only | Highest |
For warm standby, the scale-up happens in the healthy Region, whose control plane should be available. The risk is capacity: during a large regional event, many customers may try to scale up in the same alternate Region at once. Pre-provision enough capacity to carry critical traffic while scaling continues. The cloud disaster recovery guide covers setting RTO and RPO targets and the cost of each option.
5. Reference failover architecture
The pattern below applies to a tier-1 workload with a primary Region and a hot or warm standby. It keeps every failover step on the data plane and treats the decision to fail over as a human, rehearsed action.
<!-- region-failover-flow -->Key components:
- Traffic control: Route 53 failover records and health checks created in advance, with ARC routing controls or ARC Region switch to make the switch deliberate. ARC safety rules can prevent both Regions from being active at once.
- Data: DynamoDB global tables, Aurora Global Database or S3 Replication, chosen per data store. In October 2025, AWS reported global table replicas were fully caught up at 2:32 AM PDT, minutes after DNS was restored.
- Compute: Enough pre-provisioned capacity in the standby Region to carry critical traffic before Auto Scaling catches up.
- Identity: Regional STS endpoints, multi-Region SAML trust and tested break-glass users.
- Delivery and observability: CI/CD, GitOps control planes, monitoring and runbooks reachable from outside the primary Region. The Argo CD vs Flux guide covers separating GitOps control planes by trust zone.
For new multi-Region setups, use ARC Region switch with its plan evaluation to check readiness. AWS states that ARC readiness check is no longer available to new customers.
6. Data: the decision that sets your RPO
Compute fails over in minutes; data is where multi-Region designs succeed or fail. Asynchronous replication gives a small but non-zero recovery point, and backups are still required because replication copies corruption and deletions as faithfully as good writes.
AWS describes three write patterns for active-active designs:
| Write pattern | How it works | Trade-off |
|---|---|---|
| Write global | All writes go to one Region; others serve reads | Simple consistency; writes need a failover step |
| Write local | Each Region writes locally and replicates | Lowest latency; requires conflict resolution |
| Write partitioned | Each record has a home Region based on a key | Avoids conflicts; routing must follow the key |
Decide per data store, document the expected recovery point, and test reconciliation after failback. Failback is often harder than failover and is rarely rehearsed.
7. Governance, decision rights and compliance
A failover is a business decision with technical execution. Write down who can declare it, on what evidence, and how the decision is reversed. Map the controls to your framework, for example the SOC 2 availability criteria (A1.2 and A1.3), ISO/IEC 27001:2022 controls 5.29 and 5.30, or the EU Digital Operational Resilience Act for financial entities.
| Activity | Accountable | Responsible | Consulted |
|---|---|---|---|
| Business impact analysis and workload tiering | Business owner | Architecture | Finance, risk |
| Declare a Regional failover | Incident commander | SRE on-call | Product and business owner |
| Execute and verify failover | SRE lead | Platform and service teams | Security |
| Customer and regulator communication | Communications lead | Support | Legal, compliance |
| Failback and data reconciliation | Service owner | Service and data teams | SRE |
| Quarterly DR test and evidence pack | Head of platform or SRE | Service teams | Internal audit |
Keep the evidence auditors expect: the tiering record, the last test date and result for each tier-1 service, measured recovery time against target, and the actions raised from each test. The AWS Well-Architected review guide shows how to rank reliability findings by business risk so the program gets funded.
8. Test it before it happens
A failover plan that has never run is a hypothesis. Build confidence in stages:
- Zonal: Practice zonal shift on a schedule and consider zonal autoshift for eligible resources. This exercises the common single-AZ case.
- Fault injection: AWS Fault Injection Service includes scenarios for Availability Zone power interruption and cross-Region connectivity disruption. Run them in non-production, then in production with guardrails.
- Regional game days: Fail a tier-1 service to its standby Region during business hours, measure detection, decision and recovery time, then fail back.
- Dependency drills: Block access to a dependency in a test environment and confirm the service degrades gracefully instead of failing.
Feed every test and real incident into the same blameless review process. The DevOps and SRE operating model guide covers incident learning and ownership.
Then score your current position:
<!-- region-readiness-scorecard -->9. Adoption roadmap
| Phase | Typical duration | Key activities | Exit criteria |
|---|---|---|---|
| Assess | 2–3 weeks | Business impact analysis, workload tiering, dependency mapping including global services | Approved tiering and target posture per workload |
| Stabilize the Region | 4–6 weeks | Multi-AZ static stability, Regional STS endpoints, out-of-region monitoring and break-glass access | Tier-1 services survive a simulated AZ loss without scaling actions |
| Build the standby | 6–12 weeks | Data replication, pre-provisioned capacity, ARC routing controls or Region switch plan | Failover runbook executes end to end in non-production |
| Prove it | 4 weeks | Production game day, failback, evidence pack | Measured recovery time within target; actions tracked |
| Operate | Ongoing | Quarterly tests, plan evaluation, cost review of standby capacity | Tests are routine and findings close on schedule |
10. Measuring resilience
| Metric | Definition | Direction to aim for |
|---|---|---|
| Measured RTO per tier | Time from failover decision to service restored, from the last test | Within target |
| Time to decide | Detection to failover declaration | Falling |
| Recovery point achieved | Data lost or reconciled in the last test | Within target |
| Tier-1 services tested this quarter | Share of tier-1 services with a passing failover test | 100% |
| Control plane steps in recovery runbooks | Recovery steps that create or change resources during an incident | Falling toward zero |
| Error budget burn during incidents | SLO impact per event | Falling |
The SLO guide explains how to connect these measures to error budgets your product teams already use.
11. Anti-patterns and how to avoid them
| Anti-pattern | Why it hurts | Do this instead |
|---|---|---|
| Failover runbook that edits Route 53 records | Record changes depend on a us-east-1 control plane | Pre-created failover records or ARC routing controls |
| Scaling from zero in the recovery Region | Capacity may be scarce when everyone fails over | Pre-provision capacity for critical traffic |
| Global STS endpoint in SDK configuration | Adds an avoidable dependency | Regional STS endpoints everywhere |
| Monitoring and paging only in the primary Region | You lose visibility when you need it most | Out-of-region observability and status tooling |
| Multi-Region for every workload | High cost and complexity with little benefit | Tier workloads; multi-AZ first |
| Replication treated as backup | Corruption and deletes replicate too | Point-in-time backups alongside replication |
| Failover tested once, at launch | Configuration drifts and the plan stops working | Quarterly tests with evidence |
| Automatic Regional failover on a single signal | Flapping and split brain | Human decision with clear criteria, automated execution |
Frequently asked questions
Should every workload be multi-Region after the us-east-1 outages?
No. Multi-Region adds cost and operational complexity. Tier workloads by business impact, make every workload statically stable across Availability Zones, and reserve multi-Region designs for the services where hours of downtime are unacceptable.
Is us-east-1 riskier than other AWS Regions?
us-east-1 is AWS's oldest Region and one of its largest, and it hosts the control planes of several global services, so its incidents are highly visible and can affect workloads elsewhere. The design lessons apply to every Region: avoid control plane dependencies in recovery and test failover.
Should Regional failover be automatic?
Automate the execution, but keep the decision with a person for most workloads. Automatic failover on a single health signal risks flapping and split-brain writes. Use clear declaration criteria and safety rules, such as ARC routing control safety rules, so the decision is fast and safe.
How often should we test Regional failover?
Test tier-1 services at least quarterly, and after any major architecture change. Practice zonal shift more frequently, since single-AZ events are more common.
Build Regional resilience with CloudForge
CloudForge helps engineering and risk leaders turn resilience goals into tested architecture. Engagements start from your workloads and recovery objectives, not a template.
- Resilience assessment: dependency mapping, workload tiering and a prioritized plan, delivered through an AWS Well-Architected Review or our Infrastructure Audit.
- Architecture and implementation: multi-AZ static stability, standby Regions and failover runbooks through SRE consulting and cloud migration consulting.
- Testing and evidence: game days, fault injection and audit-ready test records.
Want to know how your workloads would fare in the next us-east-1 event? Contact CloudForge to review your recovery path.
Sources and further reading
- AWS: Summary of the Amazon DynamoDB service disruption in the Northern Virginia (US-EAST-1) Region, October 2025
- AWS: Summary of the AWS service event in the Northern Virginia (US-EAST-1) Region, December 2021
- AWS: Summary of the AWS Lambda service event in the Northern Virginia (US-EAST-1) Region, June 2023
- DCD: AWS experiences power issues at Northern Virginia cloud region, May 2026
- AWS whitepaper: AWS fault isolation boundaries, global services
- Amazon Builders' Library: Static stability using Availability Zones
- AWS whitepaper: Disaster recovery options in the cloud
- Amazon Application Recovery Controller (ARC) developer guide
- ARC readiness check availability change
- AWS SDKs and Tools: STS Regionalized endpoints
Want this applied to your cloud environment?
Send CloudForge your requirements and the company will identify the highest-impact next step for your cost, delivery or reliability goals.
Contact CloudForge