We Tested Our Disaster Recovery Plan Live. It Failed. Here Is What We Fixed.
A candid postmortem of a failed DR test: what broke, recovery time gaps we discovered, and the systematic fixes that brought our RTO from 4 hours to 38 minutes.

In January 2026, we scheduled our first live disaster recovery test. We had backups configured for every critical system, runbooks written for every failure scenario, and confidence that our 2-hour RTO target was achievable. We were wrong.
The test took 4 hours and 12 minutes to reach full recovery. Three services never fully recovered during the exercise window. Two critical databases restored with data loss beyond our RPO commitment. This is the story of what went wrong, why our assumptions were dangerous, and the systematic changes that brought our actual RTO down to 38 minutes.
The Setup: What We Thought We Had
Our infrastructure spans three AWS regions with the primary workload in us-east-1. The backup strategy we designed looked solid on paper:
| System | Backup Method | Frequency | Retention | Target RPO | Target RTO |
|---|---|---|---|---|---|
| Aurora PostgreSQL | Automated snapshots | Every 1 hour | 35 days | 1 hour | 30 min |
| DynamoDB | Point-in-time recovery | Continuous | 35 days | 5 min | 15 min |
| S3 (user data) | Cross-region replication | Real-time | Indefinite | ~0 | ~0 |
| ECS Services | Container redeployment | N/A | N/A | N/A | 10 min |
| Redis (ElastiCache) | Daily snapshot | Every 24h | 7 days | 24 hours | 20 min |
| Secrets Manager | Cross-region replication | Real-time | N/A | ~0 | ~0 |
Total expected RTO: 45 minutes (longest chain). Total actual RTO: 4 hours 12 minutes.
The Test Protocol
We simulated a complete us-east-1 failure by:
- Disabling all Route 53 health checks to trigger failover
- Stopping all ECS services in the primary region
- Revoking the primary Aurora cluster endpoint
- Invalidating the primary ElastiCache cluster
The goal: bring all services online in us-west-2 using only backups and DR procedures.
Failure 1: Aurora Cross-Region Read Replica Promotion
Expected behavior: Promote the us-west-2 read replica to a standalone cluster in under 5 minutes.
Actual behavior: Promotion took 23 minutes, and the promoted cluster came up with 47 minutes of replication lag we did not know about.
The replication lag was invisible in our monitoring. We tracked AuroraReplicaLag in CloudWatch, but the metric only reports within the same region. Cross-region replication lag was not alarmed.
# What we added post-test: cross-region lag monitoring
import boto3
def check_cross_region_lag():
source_client = boto3.client('rds', region_name='us-east-1')
target_client = boto3.client('rds', region_name='us-west-2')
source_cluster = source_client.describe_db_clusters(
DBClusterIdentifier='prod-primary'
)['DBClusters'][0]
target_cluster = target_client.describe_db_clusters(
DBClusterIdentifier='prod-replica-west'
)['DBClusters'][0]
source_latest = source_cluster['LatestRestorableTime']
target_latest = target_cluster['LatestRestorableTime']
lag_seconds = (source_latest - target_latest).total_seconds()
if lag_seconds > 60: # Alert if >60s lag
publish_alarm(
metric='CrossRegionReplicationLag',
value=lag_seconds,
threshold=60
)
return lag_seconds
Fix: We now monitor cross-region replication lag every 30 seconds and alert at 60 seconds. We also switched from asynchronous to global database topology, reducing typical lag from 30-50 seconds to under 2 seconds.
Failure 2: DynamoDB Restore Took 2.5 Hours
Expected behavior: Point-in-time recovery to a new table in 15 minutes.
Actual behavior: Our largest table (890 million items, 2.3 TB) took 2 hours and 34 minutes to restore.
The documentation says PITR restores create a new table, and restore time depends on table size. We had tested with a 10 GB table during setup. Nobody tested with production-scale data.
| Table Size | Item Count | Restore Time (actual) | Expected |
|---|---|---|---|
| 10 GB | 4.2M | 8 minutes | 15 min |
| 180 GB | 78M | 34 minutes | 15 min |
| 2.3 TB | 890M | 2h 34min | 15 min |
Fix: We implemented DynamoDB Global Tables for all critical tables. Global Tables maintain a synchronized replica in us-west-2 with single-digit millisecond replication. The cost increase was 37%, but the RTO dropped from 2.5 hours to zero (the replica is already live).
# Before: Single-region table with PITR
resource "aws_dynamodb_table" "orders" {
name = "orders"
billing_mode = "PAY_PER_REQUEST"
hash_key = "order_id"
point_in_time_recovery {
enabled = true
}
}
# After: Global Table with active replica
resource "aws_dynamodb_table" "orders" {
name = "orders"
billing_mode = "PAY_PER_REQUEST"
hash_key = "order_id"
point_in_time_recovery {
enabled = true
}
replica {
region_name = "us-west-2"
point_in_time_recovery = true
tags = {
Role = "disaster-recovery"
}
}
}
Failure 3: Secrets and Configuration Drift
Expected behavior: Secrets Manager cross-region replication ensures all secrets are available in us-west-2.
Actual behavior: 14 secrets existed in us-east-1 that were never configured for replication. These were added in the 6 months since our DR architecture was designed.
The root cause: no process enforcement ensuring new secrets get replication configured. Engineers created secrets for new services without updating the DR configuration.
Fix: We implemented a nightly compliance check:
def audit_secret_replication():
east_client = boto3.client('secretsmanager', region_name='us-east-1')
west_client = boto3.client('secretsmanager', region_name='us-west-2')
east_secrets = set()
paginator = east_client.get_paginator('list_secrets')
for page in paginator.paginate():
for secret in page['SecretList']:
if not secret.get('DeletedDate'):
east_secrets.add(secret['Name'])
west_secrets = set()
paginator = west_client.get_paginator('list_secrets')
for page in paginator.paginate():
for secret in page['SecretList']:
if not secret.get('DeletedDate'):
west_secrets.add(secret['Name'])
unreplicated = east_secrets - west_secrets
if unreplicated:
alert_security_team(
f"DR GAP: {len(unreplicated)} secrets not replicated to us-west-2",
list(unreplicated)
)
return unreplicated
We also added a Terraform policy that requires replication configuration for any new aws_secretsmanager_secret resource.
Failure 4: DNS Failover Cascading Delays
Expected behavior: Route 53 health checks fail, DNS updates propagate, traffic shifts to us-west-2.
Actual behavior: DNS propagation took 72 seconds (acceptable), but 40% of clients cached the old DNS entry for 5+ additional minutes due to TTL settings.
Our health check evaluation period was set to 3 intervals of 30 seconds (90 seconds before failover triggers). Combined with DNS TTL of 60 seconds and client-side caching, the actual cutover window was 3-5 minutes.
Fix:
- Reduced health check interval from 30s to 10s (fast health checks: additional $1/month per check)
- Reduced failover evaluation from 3 to 2 intervals
- Lowered DNS TTL from 60s to 10s for all failover records
- Added client-side retry guidance in our SDK documentation
Failure 5: The Runbook Was Wrong
Perhaps the most embarrassing failure: our runbook referenced IAM role ARNs that had been renamed, used AWS CLI commands with deprecated flag syntax, and assumed manual steps that required VPN access to a bastion host — a bastion host that existed only in us-east-1.
# Runbook issues found during DR test
1. Step 4: "Assume role arn:aws:iam::111:role/DR-Executor"
- Role was renamed to "InfraOps-DRExecution" 3 months ago
- Nobody updated the runbook
2. Step 7: "SSH to bastion at 10.0.1.50"
- Bastion only exists in us-east-1 (the failed region)
- No bastion in us-west-2
3. Step 12: "Run aws rds promote-read-replica"
- Correct command for RDS, wrong for Aurora
- Aurora requires: aws rds failover-global-cluster
4. Step 15: "Verify application at staging.internal.company.com"
- This DNS record does not exist in the DR region
Fix: Runbooks are now executable scripts, not documents. Every quarterly DR test runs the actual automation. If a step cannot be automated, it gets a pre-flight validation check.
The Rebuilt Architecture
After fixing all five failure categories, our DR architecture now achieves consistent sub-40-minute recovery:
| System | DR Method | Actual RTO (tested) | Actual RPO (tested) |
|---|---|---|---|
| Aurora PostgreSQL | Global Database failover | 4 min | <2 sec |
| DynamoDB | Global Tables (active) | 0 min | <1 sec |
| S3 (user data) | Cross-region replication | 0 min | ~0 |
| ECS Services | Multi-region deployment | 6 min | N/A |
| Redis (ElastiCache) | Global Datastore | 12 min | <1 sec |
| Secrets Manager | Multi-region + compliance audit | 0 min | ~0 |
| DNS cutover | Route 53 fast failover | 15 min (client propagation) | N/A |
Total tested RTO: 38 minutes (limited by DNS propagation to all clients)
Cost of Resilience
The DR improvements increased our monthly infrastructure cost by 34%:
| Improvement | Monthly Cost Increase |
|---|---|
| Aurora Global Database | +$420 |
| DynamoDB Global Tables | +$1,850 |
| ElastiCache Global Datastore | +$310 |
| Fast health checks | +$12 |
| Multi-region ECS (warm standby) | +$890 |
| DR automation + monitoring | +$145 |
| Total monthly increase | +$3,627 |
Against a 4-hour outage costing approximately $180,000 in lost revenue and SLA credits, the $43K annual DR investment pays for itself in a single avoided incident.
Key Takeaways
- Test with production-scale data. Restore times scale non-linearly. A 10 GB test tells you nothing about 2 TB behavior.
- Monitor replication lag across regions continuously. Cross-region lag is invisible unless you explicitly measure it.
- Audit DR coverage automatically. New resources added without DR configuration create silent gaps that only surface during actual failures.
- Make runbooks executable, not readable. If a human has to interpret instructions during a crisis, you have already failed.
- Budget 25-40% infrastructure cost increase for genuine multi-region DR. There is no cheap version of low RTO.
The most dangerous state for any DR program is untested confidence. We had backup configurations, runbooks, and architecture diagrams. What we did not have was proof that any of it worked at scale. Schedule the test. Accept that it will fail. Fix what breaks. That is the only path to recovery you can actually trust.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.