All projects
Cloud

Multi-Region Failover

A web application running across two AWS regions. Global Accelerator redirects traffic when the primary becomes unhealthy.

The multi-region demo application serving from Mumbai
Application screenshot

At a glance

  • 2 regions, 2 AZs each
  • Static anycast IPs, no DNS TTL wait
  • Regional traffic dials control the shift
  • Failover verified by stopping the primary

Global Accelerator · ALB · EC2 · VPC · Route 53

Architecture

Global Accelerator fronts two regional load balancers and shifts traffic to the standby region on health-check failureUsersAWS Global Acceleratorstatic anycast IPs · health-based routingAP-SOUTH-1 MUMBAI · PRIMARYApplication Load BalancerEC2 · 2 public subnets, 2 AZEU-CENTRAL-1 FRANKFURT · STANDBYApplication Load BalancerEC2 · 2 public subnets, 2 AZservingon health-check failureActive–standby: the app is stateless, the data layer is not.

Scroll the diagram horizontally to explore the full flow.

The problem

A single AWS region is still a single failure domain. DNS failover helps, but clients and recursive resolvers can retain cached records until their TTL expires.

The decision

AWS Global Accelerator provides static anycast addresses in front of regional Application Load Balancers. Health checks can remove an unhealthy regional endpoint without waiting for DNS caches to expire.

I chose active-standby rather than active-active. The application tier is stateless, but the data layer was not designed for multi-writer replication. One authoritative region and one warm standby made the system's consistency model honest.

Verification

I stopped the primary instance and observed traffic move to the standby region. Regional traffic dials also allowed a controlled shift without simulating a failure.

What broke

The first failover test marked a healthy endpoint as unavailable. The infrastructure looked complex enough to invite complex guesses, but the fault was a simple protocol mismatch between the configured health check and the endpoint.

Working upward through security groups, addressing, routing, network ACLs, health checks, and logs found the mismatch faster than starting with the accelerator.

Shape of the build

  • Two AWS regions
  • Two Availability Zones per region
  • Static anycast client entry points
  • Health-based regional failover
  • Active-standby traffic policy
  • Failure tested by stopping the primary
View Markdown on GitHubRead the write-up