True Active-Active Architecture Across 3 Regions with Conflict Resolution
Building a true active-active multi-region architecture serving 890M daily requests with CRDTs, vector clocks, and automated conflict resolution achieving 99.999% availability.

Most architectures claiming to be "active-active" are actually active-passive with fast failover. True active-active means every region simultaneously accepts writes for the same data, resolves conflicts automatically, and maintains consistency guarantees your business logic can reason about. This is fundamentally harder than failover — it requires rethinking how your data model handles concurrent mutations.
After 18 months of operating a true active-active architecture across US-East, EU-West, and APAC-Southeast serving 890M daily requests, this article covers the conflict resolution strategies, data modeling patterns, and operational realities that make it work.
Why True Active-Active
The business case for active-active is not just availability — it is latency. With active-passive, users in Asia hit a 180ms round-trip to US-East for every write operation. With active-active, writes resolve locally in under 10ms with eventual convergence across regions.
| Architecture | Write Latency (local) | Write Latency (remote) | Availability | Data Consistency |
|---|---|---|---|---|
| Single region | 5ms | N/A | 99.95% | Strong |
| Active-passive | 5ms (primary), 180ms (secondary) | N/A | 99.99% | Strong |
| Active-active (ours) | 8ms | N/A (all local) | 99.999% | Eventual (bounded) |
The 99.999% availability target means less than 5.26 minutes of downtime per year. We achieved 99.9994% in the last 12 months — 3.15 minutes of total degradation, none of which was a complete outage.
Architecture Overview
Our active-active architecture spans three regions with independent write capability and asynchronous cross-region replication.
Region Topology
- US-East (us-east-1): Primary for North American users. 3 AZs, 4 database nodes.
- EU-West (eu-west-1): Primary for European users. GDPR-scoped data stays here. 3 AZs, 4 database nodes.
- APAC-Southeast (ap-southeast-1): Primary for Asia-Pacific users. 2 AZs, 3 database nodes.
Each region runs a fully independent application stack. No region depends on another for writes. Cross-region replication is asynchronous with a target lag of under 500ms.
Conflict Resolution Strategy
The core challenge of active-active is conflict resolution. When two regions accept concurrent writes to the same entity, which write wins? We employ three strategies depending on the data domain:
Strategy 1: CRDTs for Counters and Sets
For data that can be modeled as counters, sets, or registers, we use Conflict-free Replicated Data Types (CRDTs) that merge deterministically without coordination.
// G-Counter CRDT for distributed counting (e.g., view counts, inventory)
interface GCounter {
nodeId: string;
counts: Record<string, number>; // regionId -> count
}
class GCounterCRDT {
private state: GCounter;
constructor(nodeId: string) {
this.state = { nodeId, counts: {} };
}
increment(amount: number = 1): void {
const current = this.state.counts[this.state.nodeId] || 0;
this.state.counts[this.state.nodeId] = current + amount;
}
value(): number {
return Object.values(this.state.counts).reduce((sum, c) => sum + c, 0);
}
merge(remote: GCounter): void {
// Merge by taking max of each node's counter — guaranteed convergence
for (const [nodeId, count] of Object.entries(remote.counts)) {
this.state.counts[nodeId] = Math.max(
this.state.counts[nodeId] || 0,
count
);
}
}
}
// LWW-Register for single-value fields (last writer wins with vector clock)
interface VectorClock {
[regionId: string]: number;
}
interface LWWRegister<T> {
value: T;
timestamp: number;
vectorClock: VectorClock;
regionId: string;
}
class LWWRegisterCRDT<T> {
private state: LWWRegister<T>;
constructor(value: T, regionId: string) {
this.state = {
value,
timestamp: Date.now(),
vectorClock: { [regionId]: 1 },
regionId
};
}
set(value: T, regionId: string): void {
this.state = {
value,
timestamp: Date.now(),
vectorClock: {
...this.state.vectorClock,
[regionId]: (this.state.vectorClock[regionId] || 0) + 1
},
regionId
};
}
merge(remote: LWWRegister<T>): void {
// Compare vector clocks; if concurrent, use timestamp as tiebreaker
const comparison = this.compareVectorClocks(
this.state.vectorClock,
remote.vectorClock
);
if (comparison === 'before' ||
(comparison === 'concurrent' && remote.timestamp > this.state.timestamp)) {
this.state = remote;
}
}
private compareVectorClocks(a: VectorClock, b: VectorClock):
'before' | 'after' | 'concurrent' {
let aBeforeB = false;
let bBeforeA = false;
const allKeys = new Set([...Object.keys(a), ...Object.keys(b)]);
for (const key of allKeys) {
const aVal = a[key] || 0;
const bVal = b[key] || 0;
if (aVal < bVal) aBeforeB = true;
if (bVal < aVal) bBeforeA = true;
}
if (aBeforeB && !bBeforeA) return 'before';
if (bBeforeA && !aBeforeB) return 'after';
return 'concurrent';
}
}
Strategy 2: Domain-Specific Merge Functions
For complex business entities (orders, accounts), we define custom merge functions that encode domain logic.
// Domain-specific conflict resolution for order entities
interface OrderEvent {
orderId: string;
eventType: string;
regionId: string;
timestamp: number;
vectorClock: VectorClock;
payload: Record<string, any>;
}
class OrderConflictResolver {
resolve(local: OrderEvent[], remote: OrderEvent[]): OrderEvent[] {
// Merge event streams and apply domain rules
const merged = [...local, ...remote]
.sort((a, b) => this.causalOrder(a, b));
// Apply domain invariants
return this.applyInvariants(merged);
}
private applyInvariants(events: OrderEvent[]): OrderEvent[] {
const result: OrderEvent[] = [];
let orderState = { status: 'created', items: [], total: 0 };
for (const event of events) {
// Domain rule: cancelled orders cannot be modified
if (orderState.status === 'cancelled' &&
event.eventType !== 'order.refund_issued') {
continue; // Skip conflicting event
}
// Domain rule: shipped orders cannot be cancelled
if (orderState.status === 'shipped' &&
event.eventType === 'order.cancelled') {
// Convert to return request instead
result.push({
...event,
eventType: 'order.return_requested',
payload: { ...event.payload, reason: 'conflict_resolution' }
});
continue;
}
result.push(event);
orderState = this.applyEvent(orderState, event);
}
return result;
}
private causalOrder(a: OrderEvent, b: OrderEvent): number {
const vcCompare = this.compareVectorClocks(a.vectorClock, b.vectorClock);
if (vcCompare === 'before') return -1;
if (vcCompare === 'after') return 1;
// Concurrent: deterministic tiebreaker (regionId lexicographic)
return a.regionId.localeCompare(b.regionId);
}
}
Strategy 3: Reservation-Based Writes
For operations that cannot tolerate conflicts (financial transactions, inventory decrements), we use a reservation pattern that coordinates across regions before committing.
| Write Type | Strategy | Latency | Consistency |
|---|---|---|---|
| Profile updates | LWW Register | 8ms | Eventual |
| View/click counts | G-Counter CRDT | 5ms | Eventual |
| Order status changes | Domain merge | 8ms | Eventual (bounded) |
| Payment processing | Reservation | 85ms | Strong |
| Inventory decrement | Reservation | 62ms | Strong |
Replication Pipeline
Cross-region replication uses a custom change-data-capture pipeline built on Kafka with exactly-once delivery guarantees.
Replication Lag Monitoring
| Region Pair | Avg Lag (p50) | Lag (p99) | Lag (p99.9) |
|---|---|---|---|
| US-East → EU-West | 45ms | 180ms | 420ms |
| US-East → APAC-SE | 120ms | 340ms | 780ms |
| EU-West → APAC-SE | 95ms | 280ms | 640ms |
We alert at 1 second lag and page at 5 seconds. In 18 months, we have exceeded 5 seconds twice — both during cross-Atlantic cable degradation events affecting all cloud providers.
Request Routing and Affinity
Users are routed to their nearest region via latency-based DNS (Route 53 + Cloud DNS). Once routed, session affinity keeps a user in the same region for the duration of their session to prevent read-your-own-write violations.
# Envoy configuration for session-sticky routing with region affinity
clusters:
- name: api-cluster
type: EDS
lb_policy: RING_HASH
ring_hash_lb_config:
minimum_ring_size: 1024
common_lb_config:
healthy_panic_threshold:
value: 50
zone_aware_lb_config:
routing_enabled:
value: 100
min_cluster_size: 3
health_checks:
- timeout: 2s
interval: 5s
unhealthy_threshold: 3
healthy_threshold: 2
http_health_check:
path: /health/ready
expected_statuses:
- start: 200
end: 200
Operational Metrics
Availability Achievement
| Month | Availability | Downtime | Cause |
|---|---|---|---|
| Jan 2025 | 99.9998% | 5.2s | Config deployment race |
| Feb 2025 | 100% | 0 | — |
| Mar 2025 | 99.9991% | 23.4s | APAC DNS propagation delay |
| Apr 2025 | 100% | 0 | — |
| May 2025 | 100% | 0 | — |
| Jun 2025 | 99.9995% | 13.1s | EU certificate renewal overlap |
| Jul–Dec 2025 | 100% | 0 | — |
| Annual | 99.9994% | 41.7s |
Conflict Resolution Statistics
| Metric | Daily Average |
|---|---|
| Total write operations | 890M |
| Writes requiring conflict resolution | 2.1M (0.24%) |
| Auto-resolved by CRDTs | 1.8M (86% of conflicts) |
| Auto-resolved by domain merge | 280K (13%) |
| Requiring reservation coordination | 12K (0.6%) |
| Manual escalation needed | 0 |
The 0.24% conflict rate validates our region-affinity routing — most users stay in one region, so conflicts only arise from shared resources or users who travel across regions mid-session.
Lessons Learned
Start with CRDTs for everything possible. The more of your data model you can express as CRDTs, the less custom conflict resolution you need. We converted 73% of our entities to CRDT-compatible models, which dramatically reduced operational complexity.
Test with chaos engineering regularly. We inject region failures, network partitions, and replication lag weekly. Each test exercises the conflict resolution paths that normal traffic rarely triggers.
Monitor replication lag as a business metric. For us, replication lag directly translates to the window of potential conflicts. Lower lag means fewer conflicts, which means better user experience.
Accept that some operations need coordination. Trying to make everything eventually consistent leads to impossible domain logic. Financial transactions genuinely need cross-region coordination. Accept the latency cost for those specific paths.
Conclusion
True active-active multi-region architecture is achievable but demands a fundamentally different approach to data modeling. You must decide, for every field in every entity, how conflicts resolve. The reward is extraordinary availability (99.999%+) with local-region write latency for users worldwide.
The key is pragmatism: use CRDTs where possible, domain-specific merge where necessary, and coordination-based writes only where correctness absolutely requires it. This layered approach keeps 99.7% of writes fast and local while maintaining correctness guarantees for the critical 0.3%.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.