CloudForge
Published 14 min read

Designing for a Region Outage: What Recent us-east-1 Incidents Teach

What the 2021, 2023 and 2025 AWS us-east-1 outages teach about control plane dependencies, static stability, multi-Region failover and resilience governance.

CloudForge field note
ResilienceDisaster RecoveryAWS

On October 20, 2025, a race condition in DynamoDB's DNS automation left the regional endpoint in us-east-1 with an empty DNS record. The DNS was restored in under three hours. EC2 instance launches in the Region did not return to normal for another eleven hours, and the event ended 14 hours and 32 minutes after it began.

That gap between the trigger and full recovery is the most useful thing an engineering leader can learn from recent us-east-1 incidents. This guide reviews what AWS's own post-event summaries say happened, draws the design lessons, and sets out the architecture, governance and testing an enterprise needs to keep serving customers when a Region, or a service inside it, is impaired.

Executive summary

  • Most Region "outages" are dependency failures. In October 2025 a DNS fault in one service cascaded into EC2 launches, Network Load Balancer health checks, Lambda, STS and container services.
  • Control planes fail first and recover last. In both the 2021 and 2025 events, running workloads largely kept working while the ability to launch, change or scale resources was impaired for hours.
  • Recovery that needs new resources is the riskiest recovery. Design for static stability: capacity, DNS records and credentials are in place before the incident, and failover is a data plane action.
  • Some global services have control planes in us-east-1. IAM, Organizations, Route 53 record changes and CloudFront configuration all depend on it, even for workloads that run elsewhere.
  • Multi-Region is a business decision, not a default. Tier workloads by business impact, and reserve active-active or warm standby for the services that justify the cost and complexity.
Workload profileRecommended postureWhy
Revenue or safety critical, minutes of downtime are materialMulti-Region active-active or hot standby, with pre-provisioned capacityFailover is a routing decision, not a build
Important, hours of downtime are tolerableWarm standby in a second Region, scaled down, with a rehearsed scale-upLower cost, but scale-up depends on the standby Region's control plane
Internal or batch workloadsMulti-AZ in one Region, plus pilot light or backup and restoreCovers the far more common single-AZ failure at low cost
Anything regulatedThe posture above, plus documented decision rights and test evidenceAuditors ask for proof, not architecture diagrams

1. What actually happened in us-east-1

Four recent events show the range of failures a Region can experience. All times below are as AWS or the cited report published them.

DateTriggerScopeDuration
December 7, 2021An automated scaling activity overwhelmed devices connecting AWS's internal network to the main networkRegional; control planes and several data pathsAbout 7 hours for network recovery; some services longer
June 13, 2023A latent defect in Lambda's frontend fleet when one cell crossed a new capacity thresholdRegional; Lambda and dependent services3 hours 48 minutes
October 19–20, 2025A race condition in DynamoDB's DNS management left the regional endpoint with an empty recordRegional; cascaded across many services14 hours 32 minutes
May 7–8, 2026Rising temperatures in one data center led to loss of power for affected hardwareOne Availability Zone (use1-az4); EC2 and EBSRecovery continued into the next day

October 2025: one DNS record, fourteen hours

AWS's post-event summary describes two independent DNS Enactors racing each other. A delayed Enactor applied an older plan after a newer one was already in place, and the newer plan's cleanup then deleted it, removing every IP address for dynamodb.us-east-1.amazonaws.com. Automation could not repair the inconsistent state, and operators had to intervene.

The DNS record was restored by 2:25 AM PDT. The damage was in what depended on DynamoDB. EC2's internal DropletWorkflow Manager lost its leases, and when DynamoDB returned, the backlog drove it into congestive collapse. Network configuration then queued behind it, new instances launched without connectivity, and Network Load Balancer health checks began removing healthy capacity. Each step had its own recovery time.

<!-- region-incident-timeline -->

December 2021: running workloads survived, changes did not

In 2021, AWS reported that running EC2 instances, existing load balancers, existing DNS answers and Lambda invocations were largely unaffected. What failed were the control planes: EC2 APIs, Route 53 record changes, console sign-in and STS. Teams that needed to launch capacity, edit DNS or sign in to the console to recover were blocked for hours. AWS also reported that network congestion prevented its Service Health Dashboard tooling from failing over to a standby Region.

June 2023: a cell boundary held, but not everywhere

A Lambda frontend cell crossed a capacity threshold it had never reached before and hit a latent defect. Other cells were unaffected, which shows cellular architecture working. Dependent services still felt it: STS errors throttled SAML federation sign-in, EventBridge delivery was delayed, and the us-east-1 console returned errors while consoles in other Regions were fine.

May 2026: the common case is still one Availability Zone

The most recent event was not regional. A thermal incident in one data center in use1-az4 impaired EC2 instances and EBS volumes on the affected hardware. Workloads spread across Availability Zones with spare capacity rode through it. Single-AZ failures remain far more frequent than Region-wide ones, so multi-AZ static stability is the first investment, before any multi-Region spend.

2. Five design lessons from the incidents

  1. Map dependencies, not just services. Your workload may not call DynamoDB directly, but in October 2025 EC2 launches, load balancer health and Lambda did. Ask what your recovery path depends on, two and three levels down.
  2. Assume control planes are unavailable during an incident. Launching instances, changing DNS records, creating IAM roles and editing CloudFront origins are all control plane actions. A runbook built on them may not run.
  3. Plan for recovery to take longer than the trigger. Backlogs, retries and throttles extend impact long after the root cause is fixed. Measure your RTO against the full tail, not the headline fix time.
  4. Know your hidden us-east-1 dependencies. AWS documents that IAM, Organizations and Account Management have control planes in us-east-1, and that Route 53 record, hosted zone and health check changes, CloudFront configuration and S3 bucket creation depend on it. A workload in eu-west-1 can still be affected.
  5. Keep your response tooling outside the blast radius. Monitoring, paging, runbooks, status pages and break-glass access should not live only in the Region they protect. AWS's own dashboard and support tooling were affected in 2021, 2023 and 2025.

3. Control plane, data plane and static stability

AWS's Builders' Library defines a statically stable system as one that keeps doing what it was doing when a dependency is impaired, even if it cannot receive updates. The control plane makes changes; the data plane keeps existing resources running. The data plane is simpler and built for higher availability, so recovery that only uses the data plane has fewer ways to fail.

Turn on the impairment below to see which common recovery actions still work.

<!-- region-plane-simulator -->

AWS's fault isolation guidance translates this into specific rules for recovery paths:

  • Credentials: Use Regional STS endpoints, not the global endpoint. Newer SDK major versions default to Regional endpoints; for older tools such as the AWS CLI v1 and the AWS SDK for Java 1.x, set sts_regional_endpoints = regional.
  • DNS: Do not create or edit Route 53 records, hosted zones or health checks to fail over. Pre-create failover records with health checks, or use Application Recovery Controller (ARC) routing controls, whose Regional cluster endpoints are on the data plane.
  • Edge: Do not change CloudFront origins during an incident. Configure origin failover groups in advance.
  • Capacity: Pre-provision load balancers, buckets and enough compute in the recovery location. If recovery needs control plane data, cache it in a data plane store such as Parameter Store, DynamoDB or S3.
  • Access: Allow SAML sign-in through more than one Regional endpoint, and keep tested break-glass access in case your identity provider is also affected.

4. Choosing a multi-Region strategy

AWS's disaster recovery whitepaper describes four strategies in order of cost and complexity. It does not assign fixed RTO or RPO numbers, and neither should you without testing. The column that matters most after these incidents is how much each strategy depends on control plane actions at the moment of failover.

StrategyWhat runs in the recovery RegionControl plane work at failoverCost and complexity
Backup and restoreBackups and infrastructure as code onlyHigh: provision everything, restore dataLowest
Pilot lightData replication and core services; application servers offHigh: start and scale application tierLow
Warm standbyA scaled-down, fully working copyMedium: scale up with Auto ScalingMedium
Hot standby or active-activeFull capacity, ready or already servingLow: shift traffic onlyHighest

For warm standby, the scale-up happens in the healthy Region, whose control plane should be available. The risk is capacity: during a large regional event, many customers may try to scale up in the same alternate Region at once. Pre-provision enough capacity to carry critical traffic while scaling continues. The cloud disaster recovery guide covers setting RTO and RPO targets and the cost of each option.

5. Reference failover architecture

The pattern below applies to a tier-1 workload with a primary Region and a hot or warm standby. It keeps every failover step on the data plane and treats the decision to fail over as a human, rehearsed action.

<!-- region-failover-flow -->

Key components:

  • Traffic control: Route 53 failover records and health checks created in advance, with ARC routing controls or ARC Region switch to make the switch deliberate. ARC safety rules can prevent both Regions from being active at once.
  • Data: DynamoDB global tables, Aurora Global Database or S3 Replication, chosen per data store. In October 2025, AWS reported global table replicas were fully caught up at 2:32 AM PDT, minutes after DNS was restored.
  • Compute: Enough pre-provisioned capacity in the standby Region to carry critical traffic before Auto Scaling catches up.
  • Identity: Regional STS endpoints, multi-Region SAML trust and tested break-glass users.
  • Delivery and observability: CI/CD, GitOps control planes, monitoring and runbooks reachable from outside the primary Region. The Argo CD vs Flux guide covers separating GitOps control planes by trust zone.

For new multi-Region setups, use ARC Region switch with its plan evaluation to check readiness. AWS states that ARC readiness check is no longer available to new customers.

6. Data: the decision that sets your RPO

Compute fails over in minutes; data is where multi-Region designs succeed or fail. Asynchronous replication gives a small but non-zero recovery point, and backups are still required because replication copies corruption and deletions as faithfully as good writes.

AWS describes three write patterns for active-active designs:

Write patternHow it worksTrade-off
Write globalAll writes go to one Region; others serve readsSimple consistency; writes need a failover step
Write localEach Region writes locally and replicatesLowest latency; requires conflict resolution
Write partitionedEach record has a home Region based on a keyAvoids conflicts; routing must follow the key

Decide per data store, document the expected recovery point, and test reconciliation after failback. Failback is often harder than failover and is rarely rehearsed.

7. Governance, decision rights and compliance

A failover is a business decision with technical execution. Write down who can declare it, on what evidence, and how the decision is reversed. Map the controls to your framework, for example the SOC 2 availability criteria (A1.2 and A1.3), ISO/IEC 27001:2022 controls 5.29 and 5.30, or the EU Digital Operational Resilience Act for financial entities.

ActivityAccountableResponsibleConsulted
Business impact analysis and workload tieringBusiness ownerArchitectureFinance, risk
Declare a Regional failoverIncident commanderSRE on-callProduct and business owner
Execute and verify failoverSRE leadPlatform and service teamsSecurity
Customer and regulator communicationCommunications leadSupportLegal, compliance
Failback and data reconciliationService ownerService and data teamsSRE
Quarterly DR test and evidence packHead of platform or SREService teamsInternal audit

Keep the evidence auditors expect: the tiering record, the last test date and result for each tier-1 service, measured recovery time against target, and the actions raised from each test. The AWS Well-Architected review guide shows how to rank reliability findings by business risk so the program gets funded.

8. Test it before it happens

A failover plan that has never run is a hypothesis. Build confidence in stages:

  • Zonal: Practice zonal shift on a schedule and consider zonal autoshift for eligible resources. This exercises the common single-AZ case.
  • Fault injection: AWS Fault Injection Service includes scenarios for Availability Zone power interruption and cross-Region connectivity disruption. Run them in non-production, then in production with guardrails.
  • Regional game days: Fail a tier-1 service to its standby Region during business hours, measure detection, decision and recovery time, then fail back.
  • Dependency drills: Block access to a dependency in a test environment and confirm the service degrades gracefully instead of failing.

Feed every test and real incident into the same blameless review process. The DevOps and SRE operating model guide covers incident learning and ownership.

Then score your current position:

<!-- region-readiness-scorecard -->

9. Adoption roadmap

PhaseTypical durationKey activitiesExit criteria
Assess2–3 weeksBusiness impact analysis, workload tiering, dependency mapping including global servicesApproved tiering and target posture per workload
Stabilize the Region4–6 weeksMulti-AZ static stability, Regional STS endpoints, out-of-region monitoring and break-glass accessTier-1 services survive a simulated AZ loss without scaling actions
Build the standby6–12 weeksData replication, pre-provisioned capacity, ARC routing controls or Region switch planFailover runbook executes end to end in non-production
Prove it4 weeksProduction game day, failback, evidence packMeasured recovery time within target; actions tracked
OperateOngoingQuarterly tests, plan evaluation, cost review of standby capacityTests are routine and findings close on schedule

10. Measuring resilience

MetricDefinitionDirection to aim for
Measured RTO per tierTime from failover decision to service restored, from the last testWithin target
Time to decideDetection to failover declarationFalling
Recovery point achievedData lost or reconciled in the last testWithin target
Tier-1 services tested this quarterShare of tier-1 services with a passing failover test100%
Control plane steps in recovery runbooksRecovery steps that create or change resources during an incidentFalling toward zero
Error budget burn during incidentsSLO impact per eventFalling

The SLO guide explains how to connect these measures to error budgets your product teams already use.

11. Anti-patterns and how to avoid them

Anti-patternWhy it hurtsDo this instead
Failover runbook that edits Route 53 recordsRecord changes depend on a us-east-1 control planePre-created failover records or ARC routing controls
Scaling from zero in the recovery RegionCapacity may be scarce when everyone fails overPre-provision capacity for critical traffic
Global STS endpoint in SDK configurationAdds an avoidable dependencyRegional STS endpoints everywhere
Monitoring and paging only in the primary RegionYou lose visibility when you need it mostOut-of-region observability and status tooling
Multi-Region for every workloadHigh cost and complexity with little benefitTier workloads; multi-AZ first
Replication treated as backupCorruption and deletes replicate tooPoint-in-time backups alongside replication
Failover tested once, at launchConfiguration drifts and the plan stops workingQuarterly tests with evidence
Automatic Regional failover on a single signalFlapping and split brainHuman decision with clear criteria, automated execution

Frequently asked questions

Should every workload be multi-Region after the us-east-1 outages?

No. Multi-Region adds cost and operational complexity. Tier workloads by business impact, make every workload statically stable across Availability Zones, and reserve multi-Region designs for the services where hours of downtime are unacceptable.

Is us-east-1 riskier than other AWS Regions?

us-east-1 is AWS's oldest Region and one of its largest, and it hosts the control planes of several global services, so its incidents are highly visible and can affect workloads elsewhere. The design lessons apply to every Region: avoid control plane dependencies in recovery and test failover.

Should Regional failover be automatic?

Automate the execution, but keep the decision with a person for most workloads. Automatic failover on a single health signal risks flapping and split-brain writes. Use clear declaration criteria and safety rules, such as ARC routing control safety rules, so the decision is fast and safe.

How often should we test Regional failover?

Test tier-1 services at least quarterly, and after any major architecture change. Practice zonal shift more frequently, since single-AZ events are more common.

Build Regional resilience with CloudForge

CloudForge helps engineering and risk leaders turn resilience goals into tested architecture. Engagements start from your workloads and recovery objectives, not a template.

Want to know how your workloads would fare in the next us-east-1 event? Contact CloudForge to review your recovery path.

Sources and further reading

Related expertise

Put this into practice

SRE consulting Cloud migration consulting AWS Well-Architected Review
Continue learning
Cloud Cost Anomaly Detection: An Operating Playbook for AWS, Azure and Google Cloud15 min read AWS EKS Cost Optimization: Karpenter, Spot, Graviton and Pod Rightsizing16 min read Google Cloud DevOps and SRE: Secure Delivery, GKE and Service Reliability18 min read

Want this applied to your cloud environment?

Send CloudForge your requirements and the company will identify the highest-impact next step for your cost, delivery or reliability goals.

Contact CloudForge