AWS Graviton3 Cost-Performance Analysis: 40% Savings Across 12 Workload Types
Price-performance comparison of Graviton3 vs x86 instances across web APIs, batch processing, ML inference, and databases with migration benchmarks.

Graviton3 is not just "cheaper ARM instances." After migrating 78% of our fleet from x86 (Intel/AMD) to Graviton3 over nine months, we reduced compute spend by 40% while improving performance on most workloads. But the migration was not without surprises — some workloads performed worse, and the compatibility testing required significant engineering investment.
Here is the full analysis across 12 production workload types.
The Business Case
Our starting position:
- Fleet size: 420 EC2 instances + 180 Fargate tasks
- Monthly compute spend: $186,000
- Instance families: c5, m5, r5 (Intel), c5a, m5a (AMD)
- Workloads: Web APIs, batch processing, ML inference, databases, CI/CD
The Graviton3 pricing advantage is consistent across instance sizes:
| Instance Size | x86 (c5.xlarge) | Graviton3 (c7g.xlarge) | Savings |
|---|---|---|---|
| xlarge | $0.170/hr | $0.1088/hr | 36% |
| 2xlarge | $0.340/hr | $0.2176/hr | 36% |
| 4xlarge | $0.680/hr | $0.4352/hr | 36% |
| 8xlarge | $1.360/hr | $0.8704/hr | 36% |
But pricing tells only half the story. The performance differential determines actual cost-per-transaction.
Benchmark Methodology
We ran each workload type for 7 days on equivalent instance sizes (same vCPU/RAM) and compared:
- Throughput (requests/sec, jobs/hour, or ops/sec)
- Latency (p50, p95, p99)
- Cost per unit of work ($/1M requests, $/1000 jobs)
All benchmarks used Amazon Linux 2023 with identical configurations except architecture-specific compiler flags.
Results by Workload Type
Web APIs (Node.js 20)
# Load test configuration
k6 run --vus 200 --duration 7d api-benchmark.js
| Metric | c5.2xlarge (x86) | c7g.2xlarge (Graviton3) | Delta |
|---|---|---|---|
| Requests/sec | 12,400 | 14,200 | +14.5% |
| Latency p50 | 8.2ms | 7.1ms | -13.4% |
| Latency p99 | 42ms | 34ms | -19.0% |
| Cost/1M requests | $0.0076 | $0.0043 | -43.4% |
Node.js performs exceptionally well on Graviton3 due to V8's mature ARM64 JIT compiler.
Web APIs (Go 1.22)
| Metric | c5.2xlarge (x86) | c7g.2xlarge (Graviton3) | Delta |
|---|---|---|---|
| Requests/sec | 28,600 | 34,100 | +19.2% |
| Latency p50 | 3.4ms | 2.8ms | -17.6% |
| Latency p99 | 14ms | 11ms | -21.4% |
| Cost/1M requests | $0.0033 | $0.0018 | -45.5% |
Go's native ARM64 compilation produces excellent results. This was our best-performing migration.
Batch Processing (Python 3.12)
| Metric | m5.4xlarge (x86) | m7g.4xlarge (Graviton3) | Delta |
|---|---|---|---|
| Jobs/hour | 2,400 | 2,180 | -9.2% |
| CPU utilization | 82% | 88% | +7.3% |
| Cost/1000 jobs | $0.284 | $0.199 | -29.9% |
Python CPU-intensive workloads showed a 9% performance regression on Graviton3 due to some numpy/scipy operations not having fully optimized ARM NEON paths. However, the 36% price reduction still yielded 30% cost savings.
ML Inference (PyTorch 2.1)
| Metric | c5.4xlarge (x86) | c7g.4xlarge (Graviton3) | Delta |
|---|---|---|---|
| Inferences/sec | 340 | 420 | +23.5% |
| Latency p50 | 12ms | 9.4ms | -21.7% |
| Cost/1M inferences | $0.556 | $0.293 | -47.3% |
PyTorch 2.1 with ARM-optimized oneDNN kernels showed dramatic improvements for transformer-based models.
# Ensure ARM-optimized inference
import torch
torch.set_num_threads(16) # Match vCPU count
# Use channels-last memory format for ARM optimization
model = model.to(memory_format=torch.channels_last)
# Enable BF16 for Graviton3 (has native BF16 support)
with torch.autocast(device_type='cpu', dtype=torch.bfloat16):
output = model(input_tensor)
PostgreSQL (Aurora on Graviton3)
| Metric | db.r5.2xlarge | db.r7g.2xlarge | Delta |
|---|---|---|---|
| TPS (pgbench) | 14,200 | 16,800 | +18.3% |
| Read latency p50 | 1.2ms | 0.9ms | -25% |
| Write latency p50 | 2.8ms | 2.3ms | -17.9% |
| Monthly cost | $1,640 | $1,050 | -36% |
| Cost/1M transactions | $0.081 | $0.044 | -45.7% |
Redis (ElastiCache on Graviton3)
| Metric | cache.r5.xlarge | cache.r7g.xlarge | Delta |
|---|---|---|---|
| GET ops/sec | 245,000 | 298,000 | +21.6% |
| SET ops/sec | 218,000 | 261,000 | +19.7% |
| Latency p99 | 0.8ms | 0.6ms | -25% |
| Monthly cost | $486 | $311 | -36% |
CI/CD Builds (Docker Multi-Architecture)
This was our most complex migration. Building ARM64 images requires either native ARM builders or QEMU emulation:
# GitHub Actions with Graviton3 runners
name: Build and Deploy
on: [push]
jobs:
build:
runs-on: ubuntu-24.04-arm64 # Graviton3-backed runner
steps:
- uses: actions/checkout@v4
- name: Build Docker image
run: |
docker buildx build \
--platform linux/arm64 \
--tag $ECR_REPO:$GITHUB_SHA \
--push .
| Metric | x86 Builder | Graviton3 Builder | Delta |
|---|---|---|---|
| Build time (avg) | 4.2 min | 3.8 min | -9.5% |
| Build cost/run | $0.028 | $0.018 | -35.7% |
| Monthly CI cost | $2,400 | $1,540 | -35.8% |
Migration Strategy: The Three-Wave Approach
Wave 1: Stateless Services (Weeks 1-4)
Start with stateless web APIs and workers. Lowest risk, highest reward.
# Multi-architecture Dockerfile
FROM --platform=$TARGETPLATFORM node:20-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci --only=production
COPY . .
RUN npm run build
FROM --platform=$TARGETPLATFORM node:20-alpine
WORKDIR /app
COPY --from=builder /app/dist ./dist
COPY --from=builder /app/node_modules ./node_modules
CMD ["node", "dist/server.js"]
# Build for both architectures
docker buildx build \
--platform linux/amd64,linux/arm64 \
--tag myapp:latest \
--push .
Wave 2: Databases and Caches (Weeks 5-8)
Aurora and ElastiCache support in-place Graviton migration with minimal downtime:
# Aurora: modify instance class (triggers failover ~30s)
aws rds modify-db-instance \
--db-instance-identifier prod-writer \
--db-instance-class db.r7g.2xlarge \
--apply-immediately
# ElastiCache: modify node type
aws elasticache modify-replication-group \
--replication-group-id prod-redis \
--cache-node-type cache.r7g.xlarge \
--apply-immediately
Wave 3: Problematic Workloads (Weeks 9-12)
Some workloads required code changes:
// Native addon compatibility check
// Some npm packages with native C++ addons need ARM64 builds
// Check with: npm rebuild --arch=arm64
// Problematic: sharp < 0.33 (fixed in 0.33+)
// Problematic: bcrypt < 5.1 (use bcryptjs as pure JS alternative)
// Problematic: canvas (required apt-get install build-essential)
Workloads We Kept on x86
Not everything migrated:
- Legacy .NET Framework services: No ARM64 support for .NET Framework (only .NET 6+)
- Specific Python ML models: Custom CUDA kernels without ARM equivalents
- Third-party vendor agents: Monitoring agents without ARM64 builds (DataDog resolved this in 2025, but one niche agent did not)
These represent 22% of our fleet and remain on c5/m5 instances.
Total Impact After 9 Months
| Metric | Before | After | Impact |
|---|---|---|---|
| Fleet on Graviton3 | 0% | 78% | - |
| Monthly compute spend | $186,000 | $112,000 | -$74,000 (-40%) |
| Average request latency | 12ms | 9.2ms | -23% |
| Annual savings | - | $888,000 | - |
| Migration engineering cost | - | ~$120,000 (one-time) | - |
| Payback period | - | 1.6 months | - |
Key Takeaways
- Start with Go and Node.js workloads: These show the best Graviton3 performance gains (15-20%) on top of the 36% price reduction.
- Build multi-arch from day one: Always build
linux/amd64,linux/arm64Docker images. This gives you the option to migrate without rebuilding. - Test Python workloads carefully: CPU-intensive Python with numpy/scipy may regress. Still cost-effective, but validate first.
- Database migration is nearly free performance: Aurora and ElastiCache Graviton3 instances deliver 18-22% better performance at 36% lower cost.
- The payback period is measured in weeks, not months: A 40% compute reduction with 1.6-month payback is one of the best infrastructure investments available.
Graviton3 is the rare AWS offering where the marketing claims understate the reality. The 40% price-performance improvement we measured exceeded AWS's stated 25% improvement for most workload types.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.