Kubernetes Multi-Cluster Federation: Achieving Global Availability Across 5 Regions

How we federated Kubernetes clusters across 5 regions for global high availability, handling 340K pods with unified observability and sub-second failover.

#kubernetes#multi-cluster#federation#high-availability
Cover image for the article: Kubernetes Multi-Cluster Federation: Achieving Global Availability Across 5 Regions

Running a single Kubernetes cluster is straightforward. Running five federated clusters across three continents serving 40M requests per minute while maintaining a unified developer experience — that requires an entirely different engineering discipline. After 14 months of building and operating our multi-cluster federation, we have reduced cross-region latency by 62%, eliminated single points of failure, and maintained a developer workflow that feels like deploying to a single cluster.

This article covers the architecture, tooling decisions, and operational patterns that make multi-cluster Kubernetes work at scale. Not the theory — the production reality.

Why Federation, Not a Single Large Cluster

We evaluated three approaches before committing to federation:

ApproachMax NodesBlast RadiusLatency OptimizationOperational Complexity
Single mega-cluster~5,000Entire platformNone (single region)Low
Independent clustersUnlimitedSingle regionPer-regionHigh (N clusters × M teams)
Federated clustersUnlimitedSingle regionPer-region with global routingMedium (unified control plane)

The single mega-cluster hits Kubernetes' scalability ceiling around 5,000 nodes. More critically, it creates a catastrophic blast radius — an etcd corruption or API server failure takes down everything. Independent clusters solve blast radius but create operational hell: each team must understand cluster-specific deployment, and cross-region service communication becomes an ad-hoc networking exercise.

Federation gives us the best of both worlds: regional blast radius containment with a unified API surface.

Architecture Overview

Our federation spans five clusters:

  • us-east-1 (AWS EKS) — Primary for North American traffic
  • eu-west-1 (AWS EKS) — Primary for European traffic, GDPR-compliant workloads
  • ap-southeast-1 (GKE) — Primary for APAC traffic
  • us-west-2 (AWS EKS) — DR for North America
  • eu-central-1 (GKE) — DR for Europe, data sovereignty workloads

Kubernetes Federation Architecture

Control Plane Design

We use Admiralty as our federation control plane with custom extensions. Each cluster runs autonomously — if the federation control plane fails, clusters continue serving traffic independently. Federation adds coordination, not dependency.

# Admiralty source/target configuration for multi-cluster scheduling
apiVersion: multicluster.admiralty.io/v1alpha1
kind: ClusterSource
metadata:
  name: us-east-1
  namespace: admiralty-system
spec:
  serviceAccount:
    clusterName: us-east-1
    namespace: admiralty-system
    name: federation-agent
  healthCheck:
    interval: 5s
    timeout: 3s
    unhealthyThreshold: 3
---
apiVersion: multicluster.admiralty.io/v1alpha1
kind: ClusterTarget
metadata:
  name: eu-west-1
  namespace: admiralty-system
spec:
  kubeconfigSecret:
    name: eu-west-1-kubeconfig
    namespace: admiralty-system
  scheduling:
    weight: 100
    capacityThreshold: 0.85
  constraints:
    - key: "data-sovereignty"
      operator: "In"
      values: ["eu", "global"]
    - key: "gpu-required"
      operator: "DoesNotExist"

Service Mesh for Cross-Cluster Communication

We run Istio with a shared root CA across all clusters. Each cluster has its own Istiod control plane (no shared control plane risk), but services can communicate cross-cluster transparently using locality-aware routing.

# Istio DestinationRule for cross-cluster traffic management
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
  name: payment-service
  namespace: payments
spec:
  host: payment-service.payments.svc.cluster.local
  trafficPolicy:
    connectionPool:
      tcp:
        maxConnections: 1000
        connectTimeout: 250ms
      http:
        h2UpgradePolicy: UPGRADE
        maxRequestsPerConnection: 100
    outlierDetection:
      consecutive5xxErrors: 3
      interval: 10s
      baseEjectionTime: 30s
      maxEjectionPercent: 50
    loadBalancer:
      localityLbSetting:
        enabled: true
        failover:
          - from: us-east-1
            to: us-west-2
          - from: eu-west-1
            to: eu-central-1
          - from: ap-southeast-1
            to: us-west-2
        failoverPriority:
          - "topology.kubernetes.io/region"
          - "topology.kubernetes.io/zone"
      warmupDurationSecs: 60

GitOps with Multi-Cluster Awareness

Developers deploy to "the platform," not to individual clusters. Our GitOps pipeline (ArgoCD ApplicationSets) handles cluster targeting based on workload metadata.

# ArgoCD ApplicationSet for federated deployment
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: payment-service
  namespace: argocd
spec:
  generators:
    - matrix:
        generators:
          - clusters:
              selector:
                matchLabels:
                  environment: production
                matchExpressions:
                  - key: region-role
                    operator: In
                    values: ["primary", "dr"]
          - git:
              repoURL: https://github.com/org/payment-service
              revision: HEAD
              directories:
                - path: "deploy/overlays/{{name}}"
  template:
    metadata:
      name: "payment-service-{{name}}"
      labels:
        app: payment-service
        cluster: "{{name}}"
    spec:
      project: payments
      source:
        repoURL: https://github.com/org/payment-service
        targetRevision: HEAD
        path: "deploy/overlays/{{name}}"
      destination:
        server: "{{server}}"
        namespace: payments
      syncPolicy:
        automated:
          prune: true
          selfHeal: true
        syncOptions:
          - CreateNamespace=true
          - PrunePropagationPolicy=foreground
        retry:
          limit: 5
          backoff:
            duration: 5s
            factor: 2
            maxDuration: 3m

Observability Across Clusters

Unified observability is non-negotiable for multi-cluster operations. We run a centralized Thanos deployment that aggregates Prometheus metrics from all five clusters with cluster-level labels.

MetricCollection MethodRetentionAggregation
Infrastructure metricsPrometheus per cluster15 days local, 1 year centralThanos sidecar + store
Application metricsOpenTelemetry SDK30 daysTempo with cluster labels
TracesOTel Collector per cluster7 days (sampled)Grafana Tempo
LogsVector agents30 daysLoki with cluster/region labels
Cost metricsKubecost per cluster90 daysCustom aggregation

Cross-cluster request tracing was the hardest observability problem. We propagate a custom x-cluster-trace header through the mesh that encodes the originating cluster ID, allowing us to trace requests that span multiple clusters during failover events.

Capacity Management and Autoscaling

Each cluster runs Karpenter (on AWS) or GKE Autopilot with cluster-level capacity reservations. We maintain 20% headroom in each cluster to absorb failover traffic from a peer cluster.

Our capacity planning formula:

cluster_capacity = (steady_state_demand × 1.2) + (peer_cluster_demand × 0.5)

This ensures any single cluster can absorb 50% of its paired DR cluster's traffic immediately, with autoscaling handling the remaining 50% within 90 seconds.

Failure Modes and Recovery

In 14 months of production operation, we have experienced:

  • 3 single-AZ failures — Handled automatically by Kubernetes node replacement. Zero customer impact.
  • 1 regional network partition — Lasted 8 minutes. Federation routing redirected traffic to DR cluster in under 30 seconds. Data caught up via replication once connectivity restored.
  • 1 etcd corruption — Single cluster etcd lost quorum. Cluster recovered from backup in 4 minutes. During recovery, federation routed all traffic to DR. Zero data loss.
  • 2 certificate expiration near-misses — Caught by our automated cert-rotation monitoring 72 hours before expiration. This is why you monitor certs as a first-class resource.

Key Operational Patterns

Canary across clusters. We deploy canaries to a single cluster first (us-east-1), validate for 30 minutes, then roll out to remaining clusters. This catches region-specific issues early.

Cluster maintenance windows. We drain one cluster at a time for upgrades. Traffic shifts to remaining clusters automatically. Kubernetes version upgrades happen weekly with zero downtime.

Cost optimization through workload placement. Non-latency-sensitive workloads (batch processing, ML training) run in the cheapest region. Our scheduler places these based on spot pricing across clouds.

Conclusion

Multi-cluster federation is not a technology choice — it is an organizational capability. The technical implementation matters, but the operational discipline of testing failover, maintaining unified observability, and providing a simple developer interface is what makes it sustainable.

Start with two clusters if you are new to this. Get cross-cluster networking and observability right before adding more regions. Federation complexity scales faster than linearly — every new cluster adds N-1 new communication paths. Build the tooling to manage that complexity before it manages you.

Comments

    No comments yet. Be the first to share your thoughts.