Kubernetes Multi-Cluster Federation: Achieving Global Availability Across 5 Regions
How we federated Kubernetes clusters across 5 regions for global high availability, handling 340K pods with unified observability and sub-second failover.

Running a single Kubernetes cluster is straightforward. Running five federated clusters across three continents serving 40M requests per minute while maintaining a unified developer experience — that requires an entirely different engineering discipline. After 14 months of building and operating our multi-cluster federation, we have reduced cross-region latency by 62%, eliminated single points of failure, and maintained a developer workflow that feels like deploying to a single cluster.
This article covers the architecture, tooling decisions, and operational patterns that make multi-cluster Kubernetes work at scale. Not the theory — the production reality.
Why Federation, Not a Single Large Cluster
We evaluated three approaches before committing to federation:
| Approach | Max Nodes | Blast Radius | Latency Optimization | Operational Complexity |
|---|---|---|---|---|
| Single mega-cluster | ~5,000 | Entire platform | None (single region) | Low |
| Independent clusters | Unlimited | Single region | Per-region | High (N clusters × M teams) |
| Federated clusters | Unlimited | Single region | Per-region with global routing | Medium (unified control plane) |
The single mega-cluster hits Kubernetes' scalability ceiling around 5,000 nodes. More critically, it creates a catastrophic blast radius — an etcd corruption or API server failure takes down everything. Independent clusters solve blast radius but create operational hell: each team must understand cluster-specific deployment, and cross-region service communication becomes an ad-hoc networking exercise.
Federation gives us the best of both worlds: regional blast radius containment with a unified API surface.
Architecture Overview
Our federation spans five clusters:
- us-east-1 (AWS EKS) — Primary for North American traffic
- eu-west-1 (AWS EKS) — Primary for European traffic, GDPR-compliant workloads
- ap-southeast-1 (GKE) — Primary for APAC traffic
- us-west-2 (AWS EKS) — DR for North America
- eu-central-1 (GKE) — DR for Europe, data sovereignty workloads
Control Plane Design
We use Admiralty as our federation control plane with custom extensions. Each cluster runs autonomously — if the federation control plane fails, clusters continue serving traffic independently. Federation adds coordination, not dependency.
# Admiralty source/target configuration for multi-cluster scheduling
apiVersion: multicluster.admiralty.io/v1alpha1
kind: ClusterSource
metadata:
name: us-east-1
namespace: admiralty-system
spec:
serviceAccount:
clusterName: us-east-1
namespace: admiralty-system
name: federation-agent
healthCheck:
interval: 5s
timeout: 3s
unhealthyThreshold: 3
---
apiVersion: multicluster.admiralty.io/v1alpha1
kind: ClusterTarget
metadata:
name: eu-west-1
namespace: admiralty-system
spec:
kubeconfigSecret:
name: eu-west-1-kubeconfig
namespace: admiralty-system
scheduling:
weight: 100
capacityThreshold: 0.85
constraints:
- key: "data-sovereignty"
operator: "In"
values: ["eu", "global"]
- key: "gpu-required"
operator: "DoesNotExist"
Service Mesh for Cross-Cluster Communication
We run Istio with a shared root CA across all clusters. Each cluster has its own Istiod control plane (no shared control plane risk), but services can communicate cross-cluster transparently using locality-aware routing.
# Istio DestinationRule for cross-cluster traffic management
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
name: payment-service
namespace: payments
spec:
host: payment-service.payments.svc.cluster.local
trafficPolicy:
connectionPool:
tcp:
maxConnections: 1000
connectTimeout: 250ms
http:
h2UpgradePolicy: UPGRADE
maxRequestsPerConnection: 100
outlierDetection:
consecutive5xxErrors: 3
interval: 10s
baseEjectionTime: 30s
maxEjectionPercent: 50
loadBalancer:
localityLbSetting:
enabled: true
failover:
- from: us-east-1
to: us-west-2
- from: eu-west-1
to: eu-central-1
- from: ap-southeast-1
to: us-west-2
failoverPriority:
- "topology.kubernetes.io/region"
- "topology.kubernetes.io/zone"
warmupDurationSecs: 60
GitOps with Multi-Cluster Awareness
Developers deploy to "the platform," not to individual clusters. Our GitOps pipeline (ArgoCD ApplicationSets) handles cluster targeting based on workload metadata.
# ArgoCD ApplicationSet for federated deployment
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: payment-service
namespace: argocd
spec:
generators:
- matrix:
generators:
- clusters:
selector:
matchLabels:
environment: production
matchExpressions:
- key: region-role
operator: In
values: ["primary", "dr"]
- git:
repoURL: https://github.com/org/payment-service
revision: HEAD
directories:
- path: "deploy/overlays/{{name}}"
template:
metadata:
name: "payment-service-{{name}}"
labels:
app: payment-service
cluster: "{{name}}"
spec:
project: payments
source:
repoURL: https://github.com/org/payment-service
targetRevision: HEAD
path: "deploy/overlays/{{name}}"
destination:
server: "{{server}}"
namespace: payments
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
- PrunePropagationPolicy=foreground
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
Observability Across Clusters
Unified observability is non-negotiable for multi-cluster operations. We run a centralized Thanos deployment that aggregates Prometheus metrics from all five clusters with cluster-level labels.
| Metric | Collection Method | Retention | Aggregation |
|---|---|---|---|
| Infrastructure metrics | Prometheus per cluster | 15 days local, 1 year central | Thanos sidecar + store |
| Application metrics | OpenTelemetry SDK | 30 days | Tempo with cluster labels |
| Traces | OTel Collector per cluster | 7 days (sampled) | Grafana Tempo |
| Logs | Vector agents | 30 days | Loki with cluster/region labels |
| Cost metrics | Kubecost per cluster | 90 days | Custom aggregation |
Cross-cluster request tracing was the hardest observability problem. We propagate a custom x-cluster-trace header through the mesh that encodes the originating cluster ID, allowing us to trace requests that span multiple clusters during failover events.
Capacity Management and Autoscaling
Each cluster runs Karpenter (on AWS) or GKE Autopilot with cluster-level capacity reservations. We maintain 20% headroom in each cluster to absorb failover traffic from a peer cluster.
Our capacity planning formula:
cluster_capacity = (steady_state_demand × 1.2) + (peer_cluster_demand × 0.5)
This ensures any single cluster can absorb 50% of its paired DR cluster's traffic immediately, with autoscaling handling the remaining 50% within 90 seconds.
Failure Modes and Recovery
In 14 months of production operation, we have experienced:
- 3 single-AZ failures — Handled automatically by Kubernetes node replacement. Zero customer impact.
- 1 regional network partition — Lasted 8 minutes. Federation routing redirected traffic to DR cluster in under 30 seconds. Data caught up via replication once connectivity restored.
- 1 etcd corruption — Single cluster etcd lost quorum. Cluster recovered from backup in 4 minutes. During recovery, federation routed all traffic to DR. Zero data loss.
- 2 certificate expiration near-misses — Caught by our automated cert-rotation monitoring 72 hours before expiration. This is why you monitor certs as a first-class resource.
Key Operational Patterns
Canary across clusters. We deploy canaries to a single cluster first (us-east-1), validate for 30 minutes, then roll out to remaining clusters. This catches region-specific issues early.
Cluster maintenance windows. We drain one cluster at a time for upgrades. Traffic shifts to remaining clusters automatically. Kubernetes version upgrades happen weekly with zero downtime.
Cost optimization through workload placement. Non-latency-sensitive workloads (batch processing, ML training) run in the cheapest region. Our scheduler places these based on spot pricing across clouds.
Conclusion
Multi-cluster federation is not a technology choice — it is an organizational capability. The technical implementation matters, but the operational discipline of testing failover, maintaining unified observability, and providing a simple developer interface is what makes it sustainable.
Start with two clusters if you are new to this. Get cross-cluster networking and observability right before adding more regions. Federation complexity scales faster than linearly — every new cluster adds N-1 new communication paths. Build the tooling to manage that complexity before it manages you.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.