Multi-Cloud Disaster Recovery: Active-Active Across AWS and GCP with RTO < 5 Minutes
A production-tested architecture for active-active disaster recovery spanning AWS and GCP, achieving sub-5-minute RTO with automated failover and data consistency guarantees.

When your platform processes $2.3M in transactions per hour, a single-cloud disaster recovery strategy is a liability, not a plan. After experiencing a 47-minute regional outage that cost us $1.8M in lost revenue and immeasurable customer trust, we rebuilt our entire DR architecture to span both AWS and GCP in an active-active configuration. The result: sub-5-minute RTO, zero data loss during three subsequent incidents, and a 99.999% availability SLA we can actually defend.
This is not a theoretical exercise. This is the architecture we run in production serving 12M daily active users across financial services workloads where regulatory compliance demands geographic diversity and cloud provider independence.
The Problem with Single-Cloud DR
Most organizations treat disaster recovery as a checkbox exercise. They replicate within a single cloud provider, set up cross-region failover, and call it done. The fundamental flaw: your DR strategy shares a fate with your primary infrastructure.
AWS has experienced correlated multi-region failures. GCP has had global control plane outages. When your DR depends on the same provider's networking, IAM, and control plane as your primary workload, you have redundancy — not resilience.
| Failure Scenario | Single-Cloud DR | Multi-Cloud Active-Active |
|---|---|---|
| Single AZ failure | Recovered in ~2min | No impact (traffic rerouted) |
| Regional failure | RTO 15-30min | RTO < 5min |
| Provider control plane outage | RTO unknown | RTO < 5min |
| Provider-wide networking issue | Total outage | RTO < 5min |
| DNS/routing failure | Partial outage | Automatic failover |
Architecture Overview
Our active-active architecture operates across AWS us-east-1 and GCP us-central1, with each region handling approximately 50% of production traffic at steady state. Both regions maintain full read-write capability with conflict resolution at the data layer.
Core Components
Global Traffic Management — We use a combination of Cloudflare for DNS-level routing and custom health-check agents deployed in both clouds. The health checkers perform deep application-level validation, not just TCP port checks.
Data Replication Layer — CockroachDB spans both clouds with a custom replication topology. We chose CockroachDB over traditional cross-cloud replication because it provides serializable isolation across regions without custom conflict resolution logic.
State Synchronization — For session state and ephemeral data, we run a Redis Cluster with cross-cloud replication using a custom bridge built on Redis Streams.
Service Mesh — Istio with a shared control plane manages service-to-service communication within each cloud, with cross-cloud routing handled at the ingress layer.
Data Replication Strategy
The hardest problem in multi-cloud DR is not failover — it is data consistency. We evaluated three approaches before settling on our current architecture:
# CockroachDB topology configuration for multi-cloud
# Each cloud runs 3 nodes with region-aware replication
cluster_settings:
cluster.organization: "production"
kv.rangefeed.enabled: true
kv.closed_timestamp.target_duration: "200ms"
topology:
regions:
- name: aws-us-east-1
zones:
- aws-us-east-1a
- aws-us-east-1b
- aws-us-east-1c
nodes: 3
locality: "cloud=aws,region=us-east-1"
- name: gcp-us-central1
zones:
- gcp-us-central1-a
- gcp-us-central1-b
- gcp-us-central1-c
nodes: 3
locality: "cloud=gcp,region=us-central1"
replication:
num_replicas: 5
constraints:
- "+cloud=aws": 2
- "+cloud=gcp": 2
lease_preferences:
- constraints: ["+cloud=aws"]
- constraints: ["+cloud=gcp"]
Write Path Performance
With 5-replica replication across two clouds, write latency increases compared to single-region deployments. Our benchmarks show:
| Operation | Single Region | Multi-Cloud (p50) | Multi-Cloud (p99) |
|---|---|---|---|
| Single row INSERT | 2.1ms | 8.4ms | 24.6ms |
| Batch INSERT (100 rows) | 12ms | 34ms | 89ms |
| UPDATE with index | 3.2ms | 11.2ms | 31.4ms |
| Cross-region transaction | N/A | 22ms | 67ms |
The latency increase is acceptable for our workload because we batch writes aggressively and use the AOST (As Of System Time) follower reads for read-heavy paths.
Automated Failover Pipeline
Our failover is fully automated with human override capability. The decision engine runs independently in both clouds and uses a quorum-based approach to prevent split-brain scenarios.
# Simplified failover decision engine
# Runs as a stateless service in both clouds independently
import asyncio
from dataclasses import dataclass
from enum import Enum
from typing import List
class HealthStatus(Enum):
HEALTHY = "healthy"
DEGRADED = "degraded"
UNHEALTHY = "unhealthy"
@dataclass
class RegionHealth:
region: str
cloud: str
latency_p99_ms: float
error_rate_pct: float
saturation_pct: float
last_check: float
class FailoverDecisionEngine:
def __init__(self, quorum_size: int = 3):
self.quorum_size = quorum_size
self.health_history: List[RegionHealth] = []
self.failover_cooldown_seconds = 300
self.last_failover_time = 0
async def evaluate_health(self, checks: List[RegionHealth]) -> dict:
"""
Evaluate region health and determine if failover is needed.
Requires quorum agreement before triggering failover.
"""
unhealthy_signals = 0
for check in checks:
if self._is_unhealthy(check):
unhealthy_signals += 1
should_failover = unhealthy_signals >= self.quorum_size
if should_failover and self._cooldown_elapsed():
return {
"action": "failover",
"reason": f"{unhealthy_signals}/{len(checks)} checks unhealthy",
"target_region": self._select_target(checks),
"confidence": unhealthy_signals / len(checks)
}
return {"action": "none", "reason": "healthy or in cooldown"}
def _is_unhealthy(self, check: RegionHealth) -> bool:
return (
check.latency_p99_ms > 500
or check.error_rate_pct > 5.0
or check.saturation_pct > 95.0
)
def _cooldown_elapsed(self) -> bool:
import time
return (time.time() - self.last_failover_time) > self.failover_cooldown_seconds
def _select_target(self, checks: List[RegionHealth]) -> str:
healthy = [c for c in checks if not self._is_unhealthy(c)]
if healthy:
return min(healthy, key=lambda c: c.latency_p99_ms).region
return "manual_intervention_required"
Network Architecture
Cross-cloud connectivity uses dedicated interconnect rather than public internet. We provision:
- AWS Direct Connect to our colocation facility
- GCP Cloud Interconnect to the same facility
- Cross-connect within the colocation for sub-1ms cloud-to-cloud latency
This eliminates internet routing variability and provides consistent 4-6ms round-trip between our AWS and GCP regions.
Cost Analysis
Multi-cloud DR is not cheap. Here is our monthly cost breakdown for this architecture:
| Component | Monthly Cost | Notes |
|---|---|---|
| CockroachDB cluster (6 nodes) | $18,400 | 3 nodes per cloud, 32 vCPU each |
| Cross-cloud interconnect | $4,200 | 10Gbps dedicated |
| Data transfer (cross-cloud) | $6,800 | ~8TB/month replication traffic |
| Health check infrastructure | $1,200 | Redundant checkers in both clouds |
| DNS/traffic management | $800 | Cloudflare Enterprise |
| Total DR overhead | $31,400 | On top of standard compute costs |
Against our $2.3M/hour transaction volume, the DR infrastructure pays for itself in approximately 49 seconds of prevented downtime per month.
Lessons Learned After 18 Months in Production
Test failover weekly. We run automated failover drills every Wednesday at 2 AM UTC. In 18 months, we have caught 4 configuration drifts that would have prevented successful failover.
Data replication lag is your enemy. Monitor replication lag as a tier-0 metric. We alert at 100ms lag and page at 500ms. CockroachDB's closed timestamps give us precise lag measurement.
DNS TTL matters more than you think. Even with a 30-second TTL, some resolvers cache aggressively. We use Cloudflare's proxy mode with 0-second TTL for critical paths and client-side retry logic for the long tail.
Cross-cloud IAM is a maintenance burden. Maintaining parallel IAM configurations in AWS and GCP is operationally expensive. We use Terraform with a custom provider wrapper that generates equivalent policies for both clouds from a single policy definition.
Conclusion
Active-active multi-cloud DR is achievable, but it demands engineering discipline and ongoing investment. The architecture described here has sustained us through three cloud provider incidents with zero customer-facing impact. The key insight: treat both clouds as primary from day one. If you build with the assumption that either cloud can fail at any moment, your architecture naturally becomes resilient.
Start with data replication — it is the hardest problem and has the longest lead time. Layer in automated failover once you trust your data consistency. And test relentlessly: an untested DR plan is not a plan at all.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.