AWS ECS vs EKS in Production: A 2-Year Retrospective with Real Metrics

Side-by-side production comparison of ECS and EKS after running both for 2 years across 340+ microservices with operational cost and complexity data.

#aws#ecs#eks#kubernetes#containers
Cover image for the article: AWS ECS vs EKS in Production: A 2-Year Retrospective with Real Metrics

We run both ECS and EKS in production. Not because we planned it that way, but because different teams made different choices at different times. After two years of operating 340+ microservices split across both platforms, I have hard data on where each excels and where each causes pain.

This is not a feature comparison matrix. This is operational reality.

The Setup

  • ECS cluster: 180 services, Fargate + EC2 capacity providers, serving our B2B SaaS platform
  • EKS cluster: 160 services, managed node groups + Karpenter, serving our data pipeline and ML inference workloads
  • Team size: 12 platform engineers supporting both
  • Region: us-east-1 primary, eu-west-1 DR

ECS vs EKS Architecture Overview

Cost Comparison: The Numbers Nobody Shares

Control Plane Costs

ComponentECSEKSNotes
Control plane$0$876/mo ($0.10/hr)EKS charges for the cluster
Load balancers$2,400/mo$1,800/moECS uses more ALBs by default
NAT Gateways$1,200/mo$1,200/moSame VPC topology
Service discovery$180/mo (Cloud Map)$0 (CoreDNS)Built into EKS
Subtotal$3,780/mo$3,876/moNearly identical

Compute Costs (Normalized to Same Workload)

Here is where it gets interesting. For equivalent workloads (same CPU/memory requests):

MetricECS (Fargate)ECS (EC2)EKS (Karpenter)
Monthly compute$48,200$31,400$28,900
Bin packing efficiency62%78%89%
Spot utilizationLimited35%68%
Waste (unused reserved)38%22%11%

Karpenter on EKS achieved 89% bin packing efficiency versus 62% on Fargate. The difference: Karpenter provisions exactly the right instance types for pending pods, while Fargate rounds up to predefined vCPU/memory combinations.

# Karpenter NodePool that drives 89% efficiency
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
  name: default
spec:
  template:
    spec:
      requirements:
        - key: kubernetes.io/arch
          operator: In
          values: ["amd64", "arm64"]
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"]
        - key: karpenter.k8s.aws/instance-category
          operator: In
          values: ["c", "m", "r"]
        - key: karpenter.k8s.aws/instance-generation
          operator: Gt
          values: ["5"]
      nodeClassRef:
        name: default
  disruption:
    consolidationPolicy: WhenUnderutilized
    expireAfter: 720h
  limits:
    cpu: "2000"
    memory: 4000Gi

Operational Complexity: The Hidden Tax

Deployment Pipeline Comparison

ECS deployment (via CodeDeploy Blue/Green):

{
  "taskDefinition": "arn:aws:ecs:us-east-1:123456789012:task-definition/api-service:47",
  "networkConfiguration": {
    "awsvpcConfiguration": {
      "subnets": ["subnet-abc123", "subnet-def456"],
      "securityGroups": ["sg-789012"],
      "assignPublicIp": "DISABLED"
    }
  },
  "deploymentConfiguration": {
    "deploymentCircuitBreaker": {
      "enable": true,
      "rollback": true
    },
    "maximumPercent": 200,
    "minimumHealthyPercent": 100
  }
}

EKS deployment (via ArgoCD + Helm):

apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: api-service
  namespace: argocd
spec:
  project: production
  source:
    repoURL: https://github.com/org/helm-charts
    targetRevision: HEAD
    path: charts/api-service
    helm:
      values: |
        replicaCount: 6
        image:
          tag: "v2.4.1"
        resources:
          requests:
            cpu: 500m
            memory: 512Mi
        autoscaling:
          enabled: true
          minReplicas: 6
          maxReplicas: 30
          targetCPUUtilization: 65
  destination:
    server: https://kubernetes.default.svc
    namespace: production
  syncPolicy:
    automated:
      prune: true
      selfHeal: true
    syncOptions:
      - CreateNamespace=true

Operational Incident Data (12 months)

MetricECSEKS
Platform incidents (P1/P2)311
Mean time to detect2.1 min1.8 min
Mean time to resolve18 min47 min
Upgrade-related incidents04
Networking issues16
Engineer hours on platform maintenance120 hrs/mo310 hrs/mo
On-call pages (platform)8/mo23/mo

EKS required 2.6x more platform engineering hours. The primary cost centers: Kubernetes version upgrades (4 per year), CNI plugin issues, CoreDNS scaling, and etcd performance tuning on larger clusters.

Operational Incidents Comparison

Where Each Platform Excels

ECS Wins

  1. Simplicity for standard web services: CRUD APIs, web backends, and queue consumers deploy with 80% less configuration.
  2. Operational overhead: Zero cluster upgrades, zero etcd concerns, zero CNI debugging.
  3. AWS-native integration: Service Connect, App Mesh, and CloudMap integrate without CRDs or operators.
  4. Startup time: Fargate tasks start 15-30% faster than equivalent EKS pods due to no scheduler overhead.

EKS Wins

  1. Bin packing efficiency: Karpenter is unmatched. 89% efficiency versus 62-78% on ECS.
  2. ML/GPU workloads: Native GPU scheduling, NVIDIA device plugins, and Karpenter GPU node pools.
  3. Complex scheduling: Affinity rules, topology spread, and priority classes for multi-tenant workloads.
  4. Ecosystem: Service mesh (Istio/Linkerd), observability (Prometheus/Grafana), and policy engines (OPA/Kyverno).
  5. Portability: Teams that operate in multi-cloud or hybrid environments benefit from Kubernetes abstraction.

The Decision Framework We Use Today

After living with both, here is our rubric for new services:

FactorChoose ECSChoose EKS
Team Kubernetes expertiseLow-MediumHigh
Workload typeStandard web/APIML, batch, GPU
Cost sensitivityModerate (Fargate ease > savings)High (Karpenter ROI)
Operational budgetLimited platform teamDedicated platform team
Multi-cloud requirementNoYes
Service count<50 services>100 services

Key Takeaways

  1. ECS is not "EKS lite" — it is a different operational philosophy. Less flexibility, dramatically less complexity.
  2. EKS cost advantage is real but not free: You save 20-30% on compute but spend 2.5x more on platform engineering.
  3. Karpenter is the killer feature: If you need EKS, Karpenter alone justifies the operational overhead for cost-sensitive workloads.
  4. Avoid the hybrid tax: Running both platforms doubles your cognitive load and tooling investment. Consolidate if possible.
  5. Match the tool to the team: A team of 4 engineers should not operate EKS. A team of 40 with GPU workloads should not be on Fargate.

If I were starting from zero today with a team under 20 engineers and standard web workloads, I would choose ECS every time. The operational simplicity compounds over years. For data-intensive, ML-heavy, or multi-cloud workloads with a dedicated platform team — EKS with Karpenter is the superior choice.

Comments

    No comments yet. Be the first to share your thoughts.