All projects
Cloud

Zero-Downtime Deployment

An AWS deployment setup with two parallel environments. Health checks gate releases, and routing changes handle rollback.

The running blue environment from the deployment project
Deployment screenshot

At a glance

  • 2 Availability Zones per environment
  • Blue/green cutover + rolling Instance Refresh
  • No inbound SSH, Session Manager only
  • Unhealthy targets never receive traffic

EC2 · ALB · Auto Scaling · Launch Templates · Session Manager

Architecture

An Application Load Balancer listener shifts traffic between a live Blue environment and a validated Green environmentUsersApplication Load Balancerlistener rule = the cutoverBLUE · V1 LIVETarget group · passing health checksEC2 · AZ-aEC2 · AZ-bGREEN · V2 VALIDATEDTarget group · validated, no trafficEC2 · AZ-aEC2 · AZ-b100% of traffic0% until cutoverRollback is the same operation, pointed the other way.

Scroll the diagram horizontally to explore the full flow.

The problem

A release should not become a recovery project when the new version is unhealthy. I wanted deployment and rollback to be controlled routing changes instead of emergency repairs performed on running instances.

The decision

Blue continues serving the current version while Green comes up beside it. Green must pass Application Load Balancer health checks before the listener sends it production traffic. If a problem appears after cutover, rollback points the listener back to Blue because the previous environment is still intact.

Launch Templates define the instances, Auto Scaling maintains capacity across two Availability Zones, and Session Manager replaces inbound SSH access. The same foundation also supports rolling Instance Refresh when a full parallel environment is unnecessary.

What broke

Green never became healthy. A malformed instance-metadata URL in the startup script prevented the application from binding, so the load balancer correctly kept every new target out of rotation.

I corrected the Launch Template and allowed Auto Scaling to replace the failed instances. Fixing the template instead of hand-patching a host kept the repair reproducible and preserved the deployment model.

Shape of the build

  • Blue and Green target groups
  • Health-gated listener cutover
  • Two Availability Zones per environment
  • Auto Scaling and immutable Launch Templates
  • Session Manager with no inbound SSH
  • Reverse-routing rollback
View Markdown on GitHubRead the write-up