AWS ECS vs EKS in Production: A 2-Year Retrospective with Real Metrics
Side-by-side production comparison of ECS and EKS after running both for 2 years across 340+ microservices with operational cost and complexity data.

We run both ECS and EKS in production. Not because we planned it that way, but because different teams made different choices at different times. After two years of operating 340+ microservices split across both platforms, I have hard data on where each excels and where each causes pain.
This is not a feature comparison matrix. This is operational reality.
The Setup
- ECS cluster: 180 services, Fargate + EC2 capacity providers, serving our B2B SaaS platform
- EKS cluster: 160 services, managed node groups + Karpenter, serving our data pipeline and ML inference workloads
- Team size: 12 platform engineers supporting both
- Region: us-east-1 primary, eu-west-1 DR
Cost Comparison: The Numbers Nobody Shares
Control Plane Costs
| Component | ECS | EKS | Notes |
|---|---|---|---|
| Control plane | $0 | $876/mo ($0.10/hr) | EKS charges for the cluster |
| Load balancers | $2,400/mo | $1,800/mo | ECS uses more ALBs by default |
| NAT Gateways | $1,200/mo | $1,200/mo | Same VPC topology |
| Service discovery | $180/mo (Cloud Map) | $0 (CoreDNS) | Built into EKS |
| Subtotal | $3,780/mo | $3,876/mo | Nearly identical |
Compute Costs (Normalized to Same Workload)
Here is where it gets interesting. For equivalent workloads (same CPU/memory requests):
| Metric | ECS (Fargate) | ECS (EC2) | EKS (Karpenter) |
|---|---|---|---|
| Monthly compute | $48,200 | $31,400 | $28,900 |
| Bin packing efficiency | 62% | 78% | 89% |
| Spot utilization | Limited | 35% | 68% |
| Waste (unused reserved) | 38% | 22% | 11% |
Karpenter on EKS achieved 89% bin packing efficiency versus 62% on Fargate. The difference: Karpenter provisions exactly the right instance types for pending pods, while Fargate rounds up to predefined vCPU/memory combinations.
# Karpenter NodePool that drives 89% efficiency
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: default
spec:
template:
spec:
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["5"]
nodeClassRef:
name: default
disruption:
consolidationPolicy: WhenUnderutilized
expireAfter: 720h
limits:
cpu: "2000"
memory: 4000Gi
Operational Complexity: The Hidden Tax
Deployment Pipeline Comparison
ECS deployment (via CodeDeploy Blue/Green):
{
"taskDefinition": "arn:aws:ecs:us-east-1:123456789012:task-definition/api-service:47",
"networkConfiguration": {
"awsvpcConfiguration": {
"subnets": ["subnet-abc123", "subnet-def456"],
"securityGroups": ["sg-789012"],
"assignPublicIp": "DISABLED"
}
},
"deploymentConfiguration": {
"deploymentCircuitBreaker": {
"enable": true,
"rollback": true
},
"maximumPercent": 200,
"minimumHealthyPercent": 100
}
}
EKS deployment (via ArgoCD + Helm):
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: api-service
namespace: argocd
spec:
project: production
source:
repoURL: https://github.com/org/helm-charts
targetRevision: HEAD
path: charts/api-service
helm:
values: |
replicaCount: 6
image:
tag: "v2.4.1"
resources:
requests:
cpu: 500m
memory: 512Mi
autoscaling:
enabled: true
minReplicas: 6
maxReplicas: 30
targetCPUUtilization: 65
destination:
server: https://kubernetes.default.svc
namespace: production
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
Operational Incident Data (12 months)
| Metric | ECS | EKS |
|---|---|---|
| Platform incidents (P1/P2) | 3 | 11 |
| Mean time to detect | 2.1 min | 1.8 min |
| Mean time to resolve | 18 min | 47 min |
| Upgrade-related incidents | 0 | 4 |
| Networking issues | 1 | 6 |
| Engineer hours on platform maintenance | 120 hrs/mo | 310 hrs/mo |
| On-call pages (platform) | 8/mo | 23/mo |
EKS required 2.6x more platform engineering hours. The primary cost centers: Kubernetes version upgrades (4 per year), CNI plugin issues, CoreDNS scaling, and etcd performance tuning on larger clusters.
Where Each Platform Excels
ECS Wins
- Simplicity for standard web services: CRUD APIs, web backends, and queue consumers deploy with 80% less configuration.
- Operational overhead: Zero cluster upgrades, zero etcd concerns, zero CNI debugging.
- AWS-native integration: Service Connect, App Mesh, and CloudMap integrate without CRDs or operators.
- Startup time: Fargate tasks start 15-30% faster than equivalent EKS pods due to no scheduler overhead.
EKS Wins
- Bin packing efficiency: Karpenter is unmatched. 89% efficiency versus 62-78% on ECS.
- ML/GPU workloads: Native GPU scheduling, NVIDIA device plugins, and Karpenter GPU node pools.
- Complex scheduling: Affinity rules, topology spread, and priority classes for multi-tenant workloads.
- Ecosystem: Service mesh (Istio/Linkerd), observability (Prometheus/Grafana), and policy engines (OPA/Kyverno).
- Portability: Teams that operate in multi-cloud or hybrid environments benefit from Kubernetes abstraction.
The Decision Framework We Use Today
After living with both, here is our rubric for new services:
| Factor | Choose ECS | Choose EKS |
|---|---|---|
| Team Kubernetes expertise | Low-Medium | High |
| Workload type | Standard web/API | ML, batch, GPU |
| Cost sensitivity | Moderate (Fargate ease > savings) | High (Karpenter ROI) |
| Operational budget | Limited platform team | Dedicated platform team |
| Multi-cloud requirement | No | Yes |
| Service count | <50 services | >100 services |
Key Takeaways
- ECS is not "EKS lite" — it is a different operational philosophy. Less flexibility, dramatically less complexity.
- EKS cost advantage is real but not free: You save 20-30% on compute but spend 2.5x more on platform engineering.
- Karpenter is the killer feature: If you need EKS, Karpenter alone justifies the operational overhead for cost-sensitive workloads.
- Avoid the hybrid tax: Running both platforms doubles your cognitive load and tooling investment. Consolidate if possible.
- Match the tool to the team: A team of 4 engineers should not operate EKS. A team of 40 with GPU workloads should not be on Fargate.
If I were starting from zero today with a team under 20 engineers and standard web workloads, I would choose ECS every time. The operational simplicity compounds over years. For data-intensive, ML-heavy, or multi-cloud workloads with a dedicated platform team — EKS with Karpenter is the superior choice.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.