GitOps with ArgoCD: Production Patterns for Multi-Cluster Deployments

Battle-tested ArgoCD patterns for managing multi-cluster Kubernetes deployments including app-of-apps, progressive sync waves, and disaster recovery.

#gitops#argocd#kubernetes#deployment
Cover image for the article: GitOps with ArgoCD: Production Patterns for Multi-Cluster Deployments

GitOps promises declarative infrastructure where Git is the single source of truth. ArgoCD makes that promise operational for Kubernetes. But running ArgoCD in production across multiple clusters requires patterns that go far beyond the getting-started tutorial. After managing 14 Kubernetes clusters across 3 regions with ArgoCD, here are the patterns that kept us shipping reliably.

The Problem: Multi-Cluster Complexity

Our infrastructure spans 14 clusters: 3 production (multi-region), 3 staging, 4 development, and 4 specialized (ML training, batch processing, edge). Managing deployments across this fleet introduced challenges:

  • Configuration drift between clusters: Same service running different configs in different regions
  • Deployment ordering: Services with dependencies deploying out of sequence
  • Secret management: Kubernetes secrets that cannot live in Git
  • Rollback coordination: Rolling back across clusters when one region fails
  • RBAC complexity: 12 teams with different access levels per cluster

Architecture: Hierarchical GitOps

We use a three-layer hierarchy: platform layer, team layer, and application layer. Each layer is managed by a different ArgoCD ApplicationSet or App-of-Apps pattern.

ArgoCD Multi-Cluster Architecture

Layer 1: Platform App-of-Apps

The top-level Application manages cluster-wide infrastructure components:

# platform/app-of-apps.yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: platform-components
  namespace: argocd
spec:
  project: platform
  source:
    repoURL: https://github.com/org/platform-gitops.git
    targetRevision: main
    path: clusters/{{ .Values.cluster.name }}/platform
  destination:
    server: https://kubernetes.default.svc
    namespace: argocd
  syncPolicy:
    automated:
      prune: true
      selfHeal: true
    syncOptions:
      - CreateNamespace=true
      - ServerSideApply=true
    retry:
      limit: 5
      backoff:
        duration: 5s
        factor: 2
        maxDuration: 3m

Layer 2: ApplicationSets for Team Services

Each team's services are generated from a matrix of clusters and services:

# teams/payments/applicationset.yaml
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: payments-team-services
  namespace: argocd
spec:
  generators:
    - matrix:
        generators:
          - git:
              repoURL: https://github.com/org/payments-gitops.git
              revision: main
              directories:
                - path: services/*
          - clusters:
              selector:
                matchLabels:
                  team: payments
                  env: production
  template:
    metadata:
      name: 'payments-{{path.basename}}-{{name}}'
      labels:
        team: payments
        service: '{{path.basename}}'
        cluster: '{{name}}'
    spec:
      project: payments
      source:
        repoURL: https://github.com/org/payments-gitops.git
        targetRevision: main
        path: 'services/{{path.basename}}/overlays/{{metadata.labels.region}}'
      destination:
        server: '{{server}}'
        namespace: 'payments'
      syncPolicy:
        automated:
          prune: true
          selfHeal: true

Layer 3: Sync Waves for Dependency Ordering

Services with dependencies use sync waves to ensure correct deployment order:

# services/payment-gateway/base/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: payment-gateway
  annotations:
    argocd.argoproj.io/sync-wave: "3"  # After databases (1) and caches (2)
spec:
  replicas: 3
  selector:
    matchLabels:
      app: payment-gateway
  template:
    spec:
      containers:
        - name: payment-gateway
          image: ecr.aws/org/payment-gateway:v2.14.3
          resources:
            requests:
              memory: "512Mi"
              cpu: "500m"
            limits:
              memory: "1Gi"
              cpu: "1000m"
---
# services/payment-gateway/base/database-migration-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: payment-gateway-migration
  annotations:
    argocd.argoproj.io/sync-wave: "1"   # Run before app deployment
    argocd.argoproj.io/hook: PreSync
    argocd.argoproj.io/hook-delete-policy: HookSucceeded
spec:
  template:
    spec:
      containers:
        - name: migrate
          image: ecr.aws/org/payment-gateway:v2.14.3
          command: ["npm", "run", "migrate"]
      restartPolicy: Never
  backoffLimit: 3

Secret Management: External Secrets Operator

Secrets never live in Git. We use External Secrets Operator to pull from AWS Secrets Manager:

# services/payment-gateway/base/external-secret.yaml
apiVersion: external-secrets.io/v1beta1
kind: ExternalSecret
metadata:
  name: payment-gateway-secrets
  annotations:
    argocd.argoproj.io/sync-wave: "0"  # Before everything else
spec:
  refreshInterval: 5m
  secretStoreRef:
    name: aws-secrets-manager
    kind: ClusterSecretStore
  target:
    name: payment-gateway-secrets
    creationPolicy: Owner
  data:
    - secretKey: DATABASE_URL
      remoteRef:
        key: /production/payment-gateway/database-url
    - secretKey: STRIPE_SECRET_KEY
      remoteRef:
        key: /production/payment-gateway/stripe-key
    - secretKey: JWT_SIGNING_KEY
      remoteRef:
        key: /production/shared/jwt-signing-key

Progressive Rollout Across Regions

Production deployments roll out region by region with validation gates between each:

# rollout/multi-region-strategy.yaml
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: payment-gateway-progressive
spec:
  generators:
    - list:
        elements:
          - cluster: prod-us-east-1
            order: "1"
            weight: "10"
          - cluster: prod-eu-west-1
            order: "2"
            weight: "40"
          - cluster: prod-ap-southeast-1
            order: "3"
            weight: "50"
  strategy:
    type: RollingSync
    rollingSync:
      steps:
        - matchExpressions:
            - key: order
              operator: In
              values: ["1"]
          maxUpdate: 1
        - matchExpressions:
            - key: order
              operator: In
              values: ["2"]
          maxUpdate: 1
        - matchExpressions:
            - key: order
              operator: In
              values: ["3"]
          maxUpdate: 1
  template:
    metadata:
      name: 'payment-gateway-{{cluster}}'
    spec:
      source:
        repoURL: https://github.com/org/payments-gitops.git
        path: services/payment-gateway/overlays/{{cluster}}
      destination:
        server: '{{url}}'

Drift Detection and Self-Healing

ArgoCD's self-heal feature automatically reverts manual changes. But we need visibility into what is being reverted:

# monitoring/argocd-alerts.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: argocd-drift-alerts
spec:
  groups:
    - name: argocd-drift
      rules:
        - alert: ArgoCD_SelfHealTriggered
          expr: |
            increase(argocd_app_reconcile_count{
              dest_server!~".*staging.*"
            }[5m]) > 3
          for: 2m
          labels:
            severity: warning
          annotations:
            summary: "ArgoCD self-healing triggered repeatedly for {{ $labels.name }}"
            description: "Application {{ $labels.name }} has been self-healed 3+ times in 5 minutes. Someone may be making manual changes."

Disaster Recovery: Cross-Region Failover

When a production region fails, ArgoCD facilitates rapid failover:

ArgoCD Disaster Recovery

ScenarioRecovery ActionRTO
Single service failureSelf-heal + alert30s
Node failureReschedule via sync2 min
Regional degradationTraffic shift + scale5 min
Full region failurePromote DR cluster8 min
ArgoCD control plane failureStandby ArgoCD takes over3 min

Performance at Scale

Managing 340+ Applications across 14 clusters required ArgoCD tuning:

# argocd-cm ConfigMap tuning
apiVersion: v1
kind: ConfigMap
metadata:
  name: argocd-cm
  namespace: argocd
data:
  # Reduce reconciliation pressure
  timeout.reconciliation: "180s"
  # Increase controller parallelism
  controller.status.processors: "50"
  controller.operation.processors: "25"
  # Resource exclusions to reduce watch load
  resource.exclusions: |
    - apiGroups: ["events.k8s.io"]
      kinds: ["Event"]
      clusters: ["*"]
    - apiGroups: ["metrics.k8s.io"]
      kinds: ["*"]
      clusters: ["*"]

Operational Results

MetricBefore ArgoCDAfter ArgoCDImprovement
Deployment lead time45 min8 min82% faster
Configuration drift incidents12/month0.3/month98% reduction
Cross-region deploy time2 hours25 min79% faster
Failed deployments8%1.2%85% reduction
MTTR (deployment issues)34 min4 min88% faster
Environments per service2 (staging, prod)52.5x more coverage

Key Takeaways

  1. Hierarchy matters: Platform, team, and application layers with clear boundaries prevent configuration sprawl and enable team autonomy.

  2. Sync waves enforce ordering: Database migrations before application deployments, secrets before both. Never rely on implicit ordering.

  3. Secrets belong outside Git: External Secrets Operator or Sealed Secrets. Pick one and standardize across the fleet.

  4. Progressive regional rollout: Never deploy to all regions simultaneously. Roll out region by region with automated validation between each step.

  5. Tune for scale: Default ArgoCD settings work for 20 applications. At 340+, you need aggressive resource exclusions, increased parallelism, and longer reconciliation intervals.

GitOps with ArgoCD gives you an auditable, reproducible, and recoverable deployment system. But it requires deliberate architecture to work at scale — the patterns above took us 8 months to refine.

Comments

    No comments yet. Be the first to share your thoughts.