We Tested Our Disaster Recovery Plan Live. It Failed. Here Is What We Fixed.

A candid postmortem of a failed DR test: what broke, recovery time gaps we discovered, and the systematic fixes that brought our RTO from 4 hours to 38 minutes.

#aws#backup#disaster-recovery#testing#compliance
Cover image for the article: We Tested Our Disaster Recovery Plan Live. It Failed. Here Is What We Fixed.

In January 2026, we scheduled our first live disaster recovery test. We had backups configured for every critical system, runbooks written for every failure scenario, and confidence that our 2-hour RTO target was achievable. We were wrong.

The test took 4 hours and 12 minutes to reach full recovery. Three services never fully recovered during the exercise window. Two critical databases restored with data loss beyond our RPO commitment. This is the story of what went wrong, why our assumptions were dangerous, and the systematic changes that brought our actual RTO down to 38 minutes.

The Setup: What We Thought We Had

Our infrastructure spans three AWS regions with the primary workload in us-east-1. The backup strategy we designed looked solid on paper:

SystemBackup MethodFrequencyRetentionTarget RPOTarget RTO
Aurora PostgreSQLAutomated snapshotsEvery 1 hour35 days1 hour30 min
DynamoDBPoint-in-time recoveryContinuous35 days5 min15 min
S3 (user data)Cross-region replicationReal-timeIndefinite~0~0
ECS ServicesContainer redeploymentN/AN/AN/A10 min
Redis (ElastiCache)Daily snapshotEvery 24h7 days24 hours20 min
Secrets ManagerCross-region replicationReal-timeN/A~0~0

Total expected RTO: 45 minutes (longest chain). Total actual RTO: 4 hours 12 minutes.

The Test Protocol

We simulated a complete us-east-1 failure by:

  1. Disabling all Route 53 health checks to trigger failover
  2. Stopping all ECS services in the primary region
  3. Revoking the primary Aurora cluster endpoint
  4. Invalidating the primary ElastiCache cluster

The goal: bring all services online in us-west-2 using only backups and DR procedures.

Failure 1: Aurora Cross-Region Read Replica Promotion

Expected behavior: Promote the us-west-2 read replica to a standalone cluster in under 5 minutes.

Actual behavior: Promotion took 23 minutes, and the promoted cluster came up with 47 minutes of replication lag we did not know about.

The replication lag was invisible in our monitoring. We tracked AuroraReplicaLag in CloudWatch, but the metric only reports within the same region. Cross-region replication lag was not alarmed.

# What we added post-test: cross-region lag monitoring
import boto3

def check_cross_region_lag():
    source_client = boto3.client('rds', region_name='us-east-1')
    target_client = boto3.client('rds', region_name='us-west-2')
    
    source_cluster = source_client.describe_db_clusters(
        DBClusterIdentifier='prod-primary'
    )['DBClusters'][0]
    
    target_cluster = target_client.describe_db_clusters(
        DBClusterIdentifier='prod-replica-west'
    )['DBClusters'][0]
    
    source_latest = source_cluster['LatestRestorableTime']
    target_latest = target_cluster['LatestRestorableTime']
    
    lag_seconds = (source_latest - target_latest).total_seconds()
    
    if lag_seconds > 60:  # Alert if >60s lag
        publish_alarm(
            metric='CrossRegionReplicationLag',
            value=lag_seconds,
            threshold=60
        )
    
    return lag_seconds

Fix: We now monitor cross-region replication lag every 30 seconds and alert at 60 seconds. We also switched from asynchronous to global database topology, reducing typical lag from 30-50 seconds to under 2 seconds.

Failure 2: DynamoDB Restore Took 2.5 Hours

Expected behavior: Point-in-time recovery to a new table in 15 minutes.

Actual behavior: Our largest table (890 million items, 2.3 TB) took 2 hours and 34 minutes to restore.

The documentation says PITR restores create a new table, and restore time depends on table size. We had tested with a 10 GB table during setup. Nobody tested with production-scale data.

Table SizeItem CountRestore Time (actual)Expected
10 GB4.2M8 minutes15 min
180 GB78M34 minutes15 min
2.3 TB890M2h 34min15 min

DynamoDB restore time vs table size

Fix: We implemented DynamoDB Global Tables for all critical tables. Global Tables maintain a synchronized replica in us-west-2 with single-digit millisecond replication. The cost increase was 37%, but the RTO dropped from 2.5 hours to zero (the replica is already live).

# Before: Single-region table with PITR
resource "aws_dynamodb_table" "orders" {
  name         = "orders"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "order_id"
  
  point_in_time_recovery {
    enabled = true
  }
}

# After: Global Table with active replica
resource "aws_dynamodb_table" "orders" {
  name         = "orders"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "order_id"
  
  point_in_time_recovery {
    enabled = true
  }

  replica {
    region_name = "us-west-2"
    point_in_time_recovery = true
    
    tags = {
      Role = "disaster-recovery"
    }
  }
}

Failure 3: Secrets and Configuration Drift

Expected behavior: Secrets Manager cross-region replication ensures all secrets are available in us-west-2.

Actual behavior: 14 secrets existed in us-east-1 that were never configured for replication. These were added in the 6 months since our DR architecture was designed.

The root cause: no process enforcement ensuring new secrets get replication configured. Engineers created secrets for new services without updating the DR configuration.

Fix: We implemented a nightly compliance check:

def audit_secret_replication():
    east_client = boto3.client('secretsmanager', region_name='us-east-1')
    west_client = boto3.client('secretsmanager', region_name='us-west-2')
    
    east_secrets = set()
    paginator = east_client.get_paginator('list_secrets')
    for page in paginator.paginate():
        for secret in page['SecretList']:
            if not secret.get('DeletedDate'):
                east_secrets.add(secret['Name'])
    
    west_secrets = set()
    paginator = west_client.get_paginator('list_secrets')
    for page in paginator.paginate():
        for secret in page['SecretList']:
            if not secret.get('DeletedDate'):
                west_secrets.add(secret['Name'])
    
    unreplicated = east_secrets - west_secrets
    
    if unreplicated:
        alert_security_team(
            f"DR GAP: {len(unreplicated)} secrets not replicated to us-west-2",
            list(unreplicated)
        )
    
    return unreplicated

We also added a Terraform policy that requires replication configuration for any new aws_secretsmanager_secret resource.

Failure 4: DNS Failover Cascading Delays

Expected behavior: Route 53 health checks fail, DNS updates propagate, traffic shifts to us-west-2.

Actual behavior: DNS propagation took 72 seconds (acceptable), but 40% of clients cached the old DNS entry for 5+ additional minutes due to TTL settings.

Our health check evaluation period was set to 3 intervals of 30 seconds (90 seconds before failover triggers). Combined with DNS TTL of 60 seconds and client-side caching, the actual cutover window was 3-5 minutes.

Fix:

  • Reduced health check interval from 30s to 10s (fast health checks: additional $1/month per check)
  • Reduced failover evaluation from 3 to 2 intervals
  • Lowered DNS TTL from 60s to 10s for all failover records
  • Added client-side retry guidance in our SDK documentation

Failure 5: The Runbook Was Wrong

Perhaps the most embarrassing failure: our runbook referenced IAM role ARNs that had been renamed, used AWS CLI commands with deprecated flag syntax, and assumed manual steps that required VPN access to a bastion host — a bastion host that existed only in us-east-1.

# Runbook issues found during DR test

1. Step 4: "Assume role arn:aws:iam::111:role/DR-Executor"
   - Role was renamed to "InfraOps-DRExecution" 3 months ago
   - Nobody updated the runbook

2. Step 7: "SSH to bastion at 10.0.1.50"
   - Bastion only exists in us-east-1 (the failed region)
   - No bastion in us-west-2

3. Step 12: "Run aws rds promote-read-replica"
   - Correct command for RDS, wrong for Aurora
   - Aurora requires: aws rds failover-global-cluster

4. Step 15: "Verify application at staging.internal.company.com"
   - This DNS record does not exist in the DR region

Fix: Runbooks are now executable scripts, not documents. Every quarterly DR test runs the actual automation. If a step cannot be automated, it gets a pre-flight validation check.

The Rebuilt Architecture

After fixing all five failure categories, our DR architecture now achieves consistent sub-40-minute recovery:

SystemDR MethodActual RTO (tested)Actual RPO (tested)
Aurora PostgreSQLGlobal Database failover4 min<2 sec
DynamoDBGlobal Tables (active)0 min<1 sec
S3 (user data)Cross-region replication0 min~0
ECS ServicesMulti-region deployment6 minN/A
Redis (ElastiCache)Global Datastore12 min<1 sec
Secrets ManagerMulti-region + compliance audit0 min~0
DNS cutoverRoute 53 fast failover15 min (client propagation)N/A

Total tested RTO: 38 minutes (limited by DNS propagation to all clients)

Cost of Resilience

The DR improvements increased our monthly infrastructure cost by 34%:

ImprovementMonthly Cost Increase
Aurora Global Database+$420
DynamoDB Global Tables+$1,850
ElastiCache Global Datastore+$310
Fast health checks+$12
Multi-region ECS (warm standby)+$890
DR automation + monitoring+$145
Total monthly increase+$3,627

Against a 4-hour outage costing approximately $180,000 in lost revenue and SLA credits, the $43K annual DR investment pays for itself in a single avoided incident.

Key Takeaways

  1. Test with production-scale data. Restore times scale non-linearly. A 10 GB test tells you nothing about 2 TB behavior.
  2. Monitor replication lag across regions continuously. Cross-region lag is invisible unless you explicitly measure it.
  3. Audit DR coverage automatically. New resources added without DR configuration create silent gaps that only surface during actual failures.
  4. Make runbooks executable, not readable. If a human has to interpret instructions during a crisis, you have already failed.
  5. Budget 25-40% infrastructure cost increase for genuine multi-region DR. There is no cheap version of low RTO.

The most dangerous state for any DR program is untested confidence. We had backup configurations, runbooks, and architecture diagrams. What we did not have was proof that any of it worked at scale. Schedule the test. Accept that it will fail. Fix what breaks. That is the only path to recovery you can actually trust.

Comments

    No comments yet. Be the first to share your thoughts.