GitOps with ArgoCD: Production Patterns for Multi-Cluster Deployments
Battle-tested ArgoCD patterns for managing multi-cluster Kubernetes deployments including app-of-apps, progressive sync waves, and disaster recovery.

GitOps promises declarative infrastructure where Git is the single source of truth. ArgoCD makes that promise operational for Kubernetes. But running ArgoCD in production across multiple clusters requires patterns that go far beyond the getting-started tutorial. After managing 14 Kubernetes clusters across 3 regions with ArgoCD, here are the patterns that kept us shipping reliably.
The Problem: Multi-Cluster Complexity
Our infrastructure spans 14 clusters: 3 production (multi-region), 3 staging, 4 development, and 4 specialized (ML training, batch processing, edge). Managing deployments across this fleet introduced challenges:
- Configuration drift between clusters: Same service running different configs in different regions
- Deployment ordering: Services with dependencies deploying out of sequence
- Secret management: Kubernetes secrets that cannot live in Git
- Rollback coordination: Rolling back across clusters when one region fails
- RBAC complexity: 12 teams with different access levels per cluster
Architecture: Hierarchical GitOps
We use a three-layer hierarchy: platform layer, team layer, and application layer. Each layer is managed by a different ArgoCD ApplicationSet or App-of-Apps pattern.
Layer 1: Platform App-of-Apps
The top-level Application manages cluster-wide infrastructure components:
# platform/app-of-apps.yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: platform-components
namespace: argocd
spec:
project: platform
source:
repoURL: https://github.com/org/platform-gitops.git
targetRevision: main
path: clusters/{{ .Values.cluster.name }}/platform
destination:
server: https://kubernetes.default.svc
namespace: argocd
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
- ServerSideApply=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
Layer 2: ApplicationSets for Team Services
Each team's services are generated from a matrix of clusters and services:
# teams/payments/applicationset.yaml
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: payments-team-services
namespace: argocd
spec:
generators:
- matrix:
generators:
- git:
repoURL: https://github.com/org/payments-gitops.git
revision: main
directories:
- path: services/*
- clusters:
selector:
matchLabels:
team: payments
env: production
template:
metadata:
name: 'payments-{{path.basename}}-{{name}}'
labels:
team: payments
service: '{{path.basename}}'
cluster: '{{name}}'
spec:
project: payments
source:
repoURL: https://github.com/org/payments-gitops.git
targetRevision: main
path: 'services/{{path.basename}}/overlays/{{metadata.labels.region}}'
destination:
server: '{{server}}'
namespace: 'payments'
syncPolicy:
automated:
prune: true
selfHeal: true
Layer 3: Sync Waves for Dependency Ordering
Services with dependencies use sync waves to ensure correct deployment order:
# services/payment-gateway/base/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-gateway
annotations:
argocd.argoproj.io/sync-wave: "3" # After databases (1) and caches (2)
spec:
replicas: 3
selector:
matchLabels:
app: payment-gateway
template:
spec:
containers:
- name: payment-gateway
image: ecr.aws/org/payment-gateway:v2.14.3
resources:
requests:
memory: "512Mi"
cpu: "500m"
limits:
memory: "1Gi"
cpu: "1000m"
---
# services/payment-gateway/base/database-migration-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
name: payment-gateway-migration
annotations:
argocd.argoproj.io/sync-wave: "1" # Run before app deployment
argocd.argoproj.io/hook: PreSync
argocd.argoproj.io/hook-delete-policy: HookSucceeded
spec:
template:
spec:
containers:
- name: migrate
image: ecr.aws/org/payment-gateway:v2.14.3
command: ["npm", "run", "migrate"]
restartPolicy: Never
backoffLimit: 3
Secret Management: External Secrets Operator
Secrets never live in Git. We use External Secrets Operator to pull from AWS Secrets Manager:
# services/payment-gateway/base/external-secret.yaml
apiVersion: external-secrets.io/v1beta1
kind: ExternalSecret
metadata:
name: payment-gateway-secrets
annotations:
argocd.argoproj.io/sync-wave: "0" # Before everything else
spec:
refreshInterval: 5m
secretStoreRef:
name: aws-secrets-manager
kind: ClusterSecretStore
target:
name: payment-gateway-secrets
creationPolicy: Owner
data:
- secretKey: DATABASE_URL
remoteRef:
key: /production/payment-gateway/database-url
- secretKey: STRIPE_SECRET_KEY
remoteRef:
key: /production/payment-gateway/stripe-key
- secretKey: JWT_SIGNING_KEY
remoteRef:
key: /production/shared/jwt-signing-key
Progressive Rollout Across Regions
Production deployments roll out region by region with validation gates between each:
# rollout/multi-region-strategy.yaml
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: payment-gateway-progressive
spec:
generators:
- list:
elements:
- cluster: prod-us-east-1
order: "1"
weight: "10"
- cluster: prod-eu-west-1
order: "2"
weight: "40"
- cluster: prod-ap-southeast-1
order: "3"
weight: "50"
strategy:
type: RollingSync
rollingSync:
steps:
- matchExpressions:
- key: order
operator: In
values: ["1"]
maxUpdate: 1
- matchExpressions:
- key: order
operator: In
values: ["2"]
maxUpdate: 1
- matchExpressions:
- key: order
operator: In
values: ["3"]
maxUpdate: 1
template:
metadata:
name: 'payment-gateway-{{cluster}}'
spec:
source:
repoURL: https://github.com/org/payments-gitops.git
path: services/payment-gateway/overlays/{{cluster}}
destination:
server: '{{url}}'
Drift Detection and Self-Healing
ArgoCD's self-heal feature automatically reverts manual changes. But we need visibility into what is being reverted:
# monitoring/argocd-alerts.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: argocd-drift-alerts
spec:
groups:
- name: argocd-drift
rules:
- alert: ArgoCD_SelfHealTriggered
expr: |
increase(argocd_app_reconcile_count{
dest_server!~".*staging.*"
}[5m]) > 3
for: 2m
labels:
severity: warning
annotations:
summary: "ArgoCD self-healing triggered repeatedly for {{ $labels.name }}"
description: "Application {{ $labels.name }} has been self-healed 3+ times in 5 minutes. Someone may be making manual changes."
Disaster Recovery: Cross-Region Failover
When a production region fails, ArgoCD facilitates rapid failover:
| Scenario | Recovery Action | RTO |
|---|---|---|
| Single service failure | Self-heal + alert | 30s |
| Node failure | Reschedule via sync | 2 min |
| Regional degradation | Traffic shift + scale | 5 min |
| Full region failure | Promote DR cluster | 8 min |
| ArgoCD control plane failure | Standby ArgoCD takes over | 3 min |
Performance at Scale
Managing 340+ Applications across 14 clusters required ArgoCD tuning:
# argocd-cm ConfigMap tuning
apiVersion: v1
kind: ConfigMap
metadata:
name: argocd-cm
namespace: argocd
data:
# Reduce reconciliation pressure
timeout.reconciliation: "180s"
# Increase controller parallelism
controller.status.processors: "50"
controller.operation.processors: "25"
# Resource exclusions to reduce watch load
resource.exclusions: |
- apiGroups: ["events.k8s.io"]
kinds: ["Event"]
clusters: ["*"]
- apiGroups: ["metrics.k8s.io"]
kinds: ["*"]
clusters: ["*"]
Operational Results
| Metric | Before ArgoCD | After ArgoCD | Improvement |
|---|---|---|---|
| Deployment lead time | 45 min | 8 min | 82% faster |
| Configuration drift incidents | 12/month | 0.3/month | 98% reduction |
| Cross-region deploy time | 2 hours | 25 min | 79% faster |
| Failed deployments | 8% | 1.2% | 85% reduction |
| MTTR (deployment issues) | 34 min | 4 min | 88% faster |
| Environments per service | 2 (staging, prod) | 5 | 2.5x more coverage |
Key Takeaways
-
Hierarchy matters: Platform, team, and application layers with clear boundaries prevent configuration sprawl and enable team autonomy.
-
Sync waves enforce ordering: Database migrations before application deployments, secrets before both. Never rely on implicit ordering.
-
Secrets belong outside Git: External Secrets Operator or Sealed Secrets. Pick one and standardize across the fleet.
-
Progressive regional rollout: Never deploy to all regions simultaneously. Roll out region by region with automated validation between each step.
-
Tune for scale: Default ArgoCD settings work for 20 applications. At 340+, you need aggressive resource exclusions, increased parallelism, and longer reconciliation intervals.
GitOps with ArgoCD gives you an auditable, reproducible, and recoverable deployment system. But it requires deliberate architecture to work at scale — the patterns above took us 8 months to refine.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.