Multi-Cloud Data Replication Patterns: Consistency Guarantees Across AWS and GCP
Cross-cloud replication architectures that maintain consistency guarantees while minimizing transfer costs and latency penalties.

Multi-cloud data replication is one of the hardest distributed systems problems you can choose to solve. You're fighting physics (latency between cloud provider networks), economics (egress costs that punish cross-cloud transfers), and consistency (CAP theorem doesn't care about your business requirements). After operating a multi-cloud platform across AWS and GCP for three years, here's what actually works. For the cost implications of cross-region traffic, see the AWS data transfer cost optimization guide.
Why Multi-Cloud Replication
The common arguments — vendor lock-in avoidance, best-of-breed services, regulatory compliance — all have merit. But in practice, the reason we replicate across clouds is disaster recovery at the provider level. When us-east-1 goes down (and it does), we need GCP to serve traffic within our 4-minute RTO.
Our replication requirements:
- RPO: 30 seconds for transactional data, 5 minutes for analytics
- RTO: 4 minutes for full traffic failover
- Consistency: Causal consistency for user-facing operations
- Cost ceiling: $8,000/month for replication infrastructure
Architecture Overview
We use a hub-and-spoke model where AWS is the primary write region and GCP serves as a hot standby. The replication layer consists of three components:
- Change Data Capture (CDC) from primary databases
- Cross-cloud message bus for ordered event delivery
- Conflict resolution engine for split-brain scenarios
Pattern 1: Event-Sourced Replication
Instead of replicating database state directly, we replicate events. This gives us ordering guarantees and makes conflict resolution deterministic.
import json
import hashlib
from dataclasses import dataclass, asdict
from datetime import datetime, timezone
from typing import Optional
from google.cloud import pubsub_v1
import boto3
@dataclass
class ReplicationEvent:
event_id: str
source_region: str
source_cloud: str
entity_type: str
entity_id: str
operation: str # INSERT, UPDATE, DELETE
payload: dict
vector_clock: dict
timestamp: str
checksum: str
@classmethod
def from_cdc_record(cls, record: dict, source: str) -> 'ReplicationEvent':
payload = record['dynamodb']['NewImage'] if 'NewImage' in record.get('dynamodb', {}) else record
payload_str = json.dumps(payload, sort_keys=True)
return cls(
event_id=record['eventID'],
source_region=record.get('awsRegion', 'unknown'),
source_cloud=source,
entity_type=record['eventSourceARN'].split('/')[1],
entity_id=record['dynamodb']['Keys']['pk']['S'],
operation=record['eventName'],
payload=payload,
vector_clock={source: int(datetime.now(timezone.utc).timestamp() * 1000)},
timestamp=datetime.now(timezone.utc).isoformat(),
checksum=hashlib.sha256(payload_str.encode()).hexdigest()
)
class CrossCloudReplicator:
"""Replicates events from AWS to GCP with ordering guarantees."""
def __init__(self, gcp_project: str, topic: str, aws_region: str):
self.publisher = pubsub_v1.PublisherClient()
self.topic_path = self.publisher.topic_path(gcp_project, topic)
self.kinesis = boto3.client('kinesis', region_name=aws_region)
self._sequence_counter = 0
def replicate_event(self, event: ReplicationEvent) -> str:
"""Publish event to GCP Pub/Sub with ordering key."""
message_data = json.dumps(asdict(event)).encode('utf-8')
# Use entity_id as ordering key for per-entity ordering
future = self.publisher.publish(
self.topic_path,
data=message_data,
ordering_key=f"{event.entity_type}:{event.entity_id}",
source_cloud=event.source_cloud,
event_type=event.operation,
checksum=event.checksum
)
return future.result()
def replicate_batch(self, events: list[ReplicationEvent]) -> dict:
"""Batch replicate with deduplication and ordering."""
results = {'success': 0, 'failed': 0, 'duplicates': 0}
seen_ids = set()
for event in sorted(events, key=lambda e: e.timestamp):
if event.event_id in seen_ids:
results['duplicates'] += 1
continue
seen_ids.add(event.event_id)
try:
self.replicate_event(event)
results['success'] += 1
except Exception as e:
results['failed'] += 1
self._handle_failure(event, e)
return results
def _handle_failure(self, event: ReplicationEvent, error: Exception) -> None:
"""Dead-letter failed events for retry."""
self.kinesis.put_record(
StreamName='replication-dlq',
Data=json.dumps(asdict(event)).encode(),
PartitionKey=event.entity_id
)
Pattern 2: Conflict Resolution with Vector Clocks
In a multi-cloud setup, split-brain scenarios are inevitable. When both clouds accept writes during a network partition, you need deterministic conflict resolution.
from typing import Any
class VectorClockResolver:
"""Resolve conflicts using vector clocks and last-writer-wins fallback."""
def __init__(self, local_node: str):
self.local_node = local_node
def compare_clocks(self, clock_a: dict, clock_b: dict) -> str:
"""Compare two vector clocks. Returns: 'a_wins', 'b_wins', or 'concurrent'."""
a_dominates = False
b_dominates = False
all_nodes = set(clock_a.keys()) | set(clock_b.keys())
for node in all_nodes:
val_a = clock_a.get(node, 0)
val_b = clock_b.get(node, 0)
if val_a > val_b:
a_dominates = True
elif val_b > val_a:
b_dominates = True
if a_dominates and not b_dominates:
return 'a_wins'
elif b_dominates and not a_dominates:
return 'b_wins'
else:
return 'concurrent'
def resolve(self, event_a: ReplicationEvent, event_b: ReplicationEvent) -> ReplicationEvent:
"""Resolve conflicting events."""
result = self.compare_clocks(event_a.vector_clock, event_b.vector_clock)
if result == 'a_wins':
return event_a
elif result == 'b_wins':
return event_b
else:
# Concurrent: use deterministic tiebreaker
# Higher checksum wins (arbitrary but consistent)
if event_a.checksum >= event_b.checksum:
return event_a
return event_b
Pattern 3: Bandwidth-Optimized Transfer
Cross-cloud egress costs $0.08-0.12/GB. At 50 TB/month of replication traffic, naive approaches cost $4,000-6,000/month. We optimize through:
- Delta compression: Only transfer changed fields, not full documents
- Batch windowing: Accumulate 5 seconds of changes, compress, send as single payload
- Deduplication: If an entity is updated 10 times in 5 seconds, only send the final state
Cost Analysis
| Approach | Monthly Transfer | Egress Cost | Infra Cost | Total |
|---|---|---|---|---|
| Naive full replication | 50 TB | $5,000 | $800 | $5,800 |
| Delta compression | 12 TB | $1,200 | $1,200 | $2,400 |
| Batched + compressed | 4.2 TB | $420 | $1,500 | $1,920 |
| Our approach (all combined) | 3.1 TB | $310 | $1,800 | $2,110 |
Consistency Guarantees
We provide causal consistency, not strong consistency. This means:
- If user A writes data and then reads it, they see their own write (read-your-writes)
- If user A writes data and tells user B, user B will eventually see the write (causal ordering)
- We do NOT guarantee that all users see all writes in the same order globally
This tradeoff reduces cross-cloud coordination from synchronous (100-200ms latency per write) to asynchronous (sub-second replication lag under normal conditions).
Monitoring Replication Health
Key metrics we track — see also the dedicated deep-dive on database replication lag monitoring for implementation patterns:
- Replication lag (p99): Target < 5 seconds, alert at > 30 seconds
- Conflict rate: Target < 0.01% of events, alert at > 0.1%
- Checksum mismatches: Target 0, alert at > 0 (indicates data corruption)
- Cross-cloud bandwidth: Track trends, alert on 20% deviation from baseline
Lessons Learned
- Start with eventual consistency. Trying to build strong consistency across clouds is a fool's errand. Accept eventual consistency and design your application to handle it.
- Egress costs dominate. The infrastructure to run replication is cheap. The bytes crossing cloud boundaries are expensive. Optimize transfer volume above all else.
- Test failover monthly. Replication that has never been failed over is replication that doesn't work. We run monthly chaos exercises cutting the cross-cloud link.
- Vector clocks beat timestamps. Clock skew between AWS and GCP makes timestamp-based ordering unreliable. Vector clocks give you deterministic ordering.
- Budget for the DLQ. 0.1% of events will fail to replicate on first attempt. Dead-letter queues with automated retry are not optional.
Multi-cloud replication is expensive and complex, but when your primary cloud provider has a bad day, it's the difference between a 4-minute failover and a 4-hour outage.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.