Multi-Cloud Disaster Recovery: Active-Active Across AWS and GCP with RTO < 5 Minutes

A production-tested architecture for active-active disaster recovery spanning AWS and GCP, achieving sub-5-minute RTO with automated failover and data consistency guarantees.

#multi-cloud#disaster-recovery#aws#gcp
Cover image for the article: Multi-Cloud Disaster Recovery: Active-Active Across AWS and GCP with RTO < 5 Minutes

When your platform processes $2.3M in transactions per hour, a single-cloud disaster recovery strategy is a liability, not a plan. After experiencing a 47-minute regional outage that cost us $1.8M in lost revenue and immeasurable customer trust, we rebuilt our entire DR architecture to span both AWS and GCP in an active-active configuration. The result: sub-5-minute RTO, zero data loss during three subsequent incidents, and a 99.999% availability SLA we can actually defend.

This is not a theoretical exercise. This is the architecture we run in production serving 12M daily active users across financial services workloads where regulatory compliance demands geographic diversity and cloud provider independence.

The Problem with Single-Cloud DR

Most organizations treat disaster recovery as a checkbox exercise. They replicate within a single cloud provider, set up cross-region failover, and call it done. The fundamental flaw: your DR strategy shares a fate with your primary infrastructure.

AWS has experienced correlated multi-region failures. GCP has had global control plane outages. When your DR depends on the same provider's networking, IAM, and control plane as your primary workload, you have redundancy — not resilience.

Failure ScenarioSingle-Cloud DRMulti-Cloud Active-Active
Single AZ failureRecovered in ~2minNo impact (traffic rerouted)
Regional failureRTO 15-30minRTO < 5min
Provider control plane outageRTO unknownRTO < 5min
Provider-wide networking issueTotal outageRTO < 5min
DNS/routing failurePartial outageAutomatic failover

Architecture Overview

Our active-active architecture operates across AWS us-east-1 and GCP us-central1, with each region handling approximately 50% of production traffic at steady state. Both regions maintain full read-write capability with conflict resolution at the data layer.

Multi-Cloud DR Architecture

Core Components

Global Traffic Management — We use a combination of Cloudflare for DNS-level routing and custom health-check agents deployed in both clouds. The health checkers perform deep application-level validation, not just TCP port checks.

Data Replication Layer — CockroachDB spans both clouds with a custom replication topology. We chose CockroachDB over traditional cross-cloud replication because it provides serializable isolation across regions without custom conflict resolution logic.

State Synchronization — For session state and ephemeral data, we run a Redis Cluster with cross-cloud replication using a custom bridge built on Redis Streams.

Service Mesh — Istio with a shared control plane manages service-to-service communication within each cloud, with cross-cloud routing handled at the ingress layer.

Data Replication Strategy

The hardest problem in multi-cloud DR is not failover — it is data consistency. We evaluated three approaches before settling on our current architecture:

# CockroachDB topology configuration for multi-cloud
# Each cloud runs 3 nodes with region-aware replication
cluster_settings:
  cluster.organization: "production"
  kv.rangefeed.enabled: true
  kv.closed_timestamp.target_duration: "200ms"
  
topology:
  regions:
    - name: aws-us-east-1
      zones:
        - aws-us-east-1a
        - aws-us-east-1b
        - aws-us-east-1c
      nodes: 3
      locality: "cloud=aws,region=us-east-1"
    - name: gcp-us-central1
      zones:
        - gcp-us-central1-a
        - gcp-us-central1-b
        - gcp-us-central1-c
      nodes: 3
      locality: "cloud=gcp,region=us-central1"

replication:
  num_replicas: 5
  constraints:
    - "+cloud=aws": 2
    - "+cloud=gcp": 2
  lease_preferences:
    - constraints: ["+cloud=aws"]
    - constraints: ["+cloud=gcp"]

Write Path Performance

With 5-replica replication across two clouds, write latency increases compared to single-region deployments. Our benchmarks show:

OperationSingle RegionMulti-Cloud (p50)Multi-Cloud (p99)
Single row INSERT2.1ms8.4ms24.6ms
Batch INSERT (100 rows)12ms34ms89ms
UPDATE with index3.2ms11.2ms31.4ms
Cross-region transactionN/A22ms67ms

The latency increase is acceptable for our workload because we batch writes aggressively and use the AOST (As Of System Time) follower reads for read-heavy paths.

Automated Failover Pipeline

Our failover is fully automated with human override capability. The decision engine runs independently in both clouds and uses a quorum-based approach to prevent split-brain scenarios.

# Simplified failover decision engine
# Runs as a stateless service in both clouds independently

import asyncio
from dataclasses import dataclass
from enum import Enum
from typing import List

class HealthStatus(Enum):
    HEALTHY = "healthy"
    DEGRADED = "degraded"
    UNHEALTHY = "unhealthy"

@dataclass
class RegionHealth:
    region: str
    cloud: str
    latency_p99_ms: float
    error_rate_pct: float
    saturation_pct: float
    last_check: float

class FailoverDecisionEngine:
    def __init__(self, quorum_size: int = 3):
        self.quorum_size = quorum_size
        self.health_history: List[RegionHealth] = []
        self.failover_cooldown_seconds = 300
        self.last_failover_time = 0

    async def evaluate_health(self, checks: List[RegionHealth]) -> dict:
        """
        Evaluate region health and determine if failover is needed.
        Requires quorum agreement before triggering failover.
        """
        unhealthy_signals = 0
        
        for check in checks:
            if self._is_unhealthy(check):
                unhealthy_signals += 1

        should_failover = unhealthy_signals >= self.quorum_size
        
        if should_failover and self._cooldown_elapsed():
            return {
                "action": "failover",
                "reason": f"{unhealthy_signals}/{len(checks)} checks unhealthy",
                "target_region": self._select_target(checks),
                "confidence": unhealthy_signals / len(checks)
            }
        
        return {"action": "none", "reason": "healthy or in cooldown"}

    def _is_unhealthy(self, check: RegionHealth) -> bool:
        return (
            check.latency_p99_ms > 500
            or check.error_rate_pct > 5.0
            or check.saturation_pct > 95.0
        )

    def _cooldown_elapsed(self) -> bool:
        import time
        return (time.time() - self.last_failover_time) > self.failover_cooldown_seconds

    def _select_target(self, checks: List[RegionHealth]) -> str:
        healthy = [c for c in checks if not self._is_unhealthy(c)]
        if healthy:
            return min(healthy, key=lambda c: c.latency_p99_ms).region
        return "manual_intervention_required"

Network Architecture

Cross-cloud connectivity uses dedicated interconnect rather than public internet. We provision:

  • AWS Direct Connect to our colocation facility
  • GCP Cloud Interconnect to the same facility
  • Cross-connect within the colocation for sub-1ms cloud-to-cloud latency

This eliminates internet routing variability and provides consistent 4-6ms round-trip between our AWS and GCP regions.

Cost Analysis

Multi-cloud DR is not cheap. Here is our monthly cost breakdown for this architecture:

ComponentMonthly CostNotes
CockroachDB cluster (6 nodes)$18,4003 nodes per cloud, 32 vCPU each
Cross-cloud interconnect$4,20010Gbps dedicated
Data transfer (cross-cloud)$6,800~8TB/month replication traffic
Health check infrastructure$1,200Redundant checkers in both clouds
DNS/traffic management$800Cloudflare Enterprise
Total DR overhead$31,400On top of standard compute costs

Against our $2.3M/hour transaction volume, the DR infrastructure pays for itself in approximately 49 seconds of prevented downtime per month.

Lessons Learned After 18 Months in Production

Test failover weekly. We run automated failover drills every Wednesday at 2 AM UTC. In 18 months, we have caught 4 configuration drifts that would have prevented successful failover.

Data replication lag is your enemy. Monitor replication lag as a tier-0 metric. We alert at 100ms lag and page at 500ms. CockroachDB's closed timestamps give us precise lag measurement.

DNS TTL matters more than you think. Even with a 30-second TTL, some resolvers cache aggressively. We use Cloudflare's proxy mode with 0-second TTL for critical paths and client-side retry logic for the long tail.

Cross-cloud IAM is a maintenance burden. Maintaining parallel IAM configurations in AWS and GCP is operationally expensive. We use Terraform with a custom provider wrapper that generates equivalent policies for both clouds from a single policy definition.

Conclusion

Active-active multi-cloud DR is achievable, but it demands engineering discipline and ongoing investment. The architecture described here has sustained us through three cloud provider incidents with zero customer-facing impact. The key insight: treat both clouds as primary from day one. If you build with the assumption that either cloud can fail at any moment, your architecture naturally becomes resilient.

Start with data replication — it is the hardest problem and has the longest lead time. Layer in automated failover once you trust your data consistency. And test relentlessly: an untested DR plan is not a plan at all.

Comments

    No comments yet. Be the first to share your thoughts.