Aurora Serverless v2 Benchmarks: When Auto-Scaling Meets Production Reality
Performance benchmarks of Aurora Serverless v2 under varying load patterns including cold scaling, connection storms, and cost comparison with provisioned instances.

Aurora Serverless v2 promises automatic scaling without capacity planning. After running it in production for 14 months across three workloads — a SaaS platform, batch analytics, and a microservices backend — I have benchmark data that reveals where it delivers on that promise and where it falls short.
The short version: it is excellent for variable workloads with predictable baseline demand. It is expensive for steady-state workloads and problematic for sudden spike scenarios.
Test Environment
- Engine: Aurora PostgreSQL 15.4
- ACU range: 2-64 (configurable min/max)
- Region: us-east-1
- Client: pgbench and custom Go load generator on EC2 (c6i.4xlarge)
- Comparison baseline: db.r6g.2xlarge provisioned (8 vCPU, 64GB RAM)
An ACU (Aurora Capacity Unit) equals approximately 2GB of RAM with corresponding CPU. The provisioned baseline (8 vCPU, 64GB) maps to roughly 32 ACU.
Benchmark 1: Steady-State OLTP Performance
Running pgbench with 100 concurrent connections at sustained load for 1 hour:
pgbench -h aurora-cluster.cluster-xyz.us-east-1.rds.amazonaws.com \
-U benchuser -d benchdb \
-c 100 -j 16 -T 3600 \
--protocol=prepared \
-f custom_oltp.sql
Results:
| Metric | Provisioned (r6g.2xl) | Serverless v2 (2-64 ACU) | Delta |
|---|---|---|---|
| TPS (transactions/sec) | 14,200 | 13,800 | -2.8% |
| Latency p50 | 6.8ms | 7.1ms | +4.4% |
| Latency p99 | 18.2ms | 19.8ms | +8.8% |
| Settled ACU | - | 34 ACU | - |
| Monthly cost | $1,640 | $2,448 | +49.3% |
At steady state, Serverless v2 performs within 5-10% of provisioned but costs 49% more. The ACU settled at 34 (just above our 32 ACU baseline equivalent) and stayed there — demonstrating that for steady workloads, you pay a premium for the scaling capability you are not using.
Benchmark 2: Scaling Response Time
The critical question: how fast does Serverless v2 scale when load increases suddenly?
Test: ramp from 10 concurrent connections to 500 over 60 seconds.
// load-generator.go — Ramp test
func rampTest(ctx context.Context, pool *pgxpool.Pool) {
startConnections := 10
endConnections := 500
rampDuration := 60 * time.Second
step := (endConnections - startConnections) / 60
for i := 0; i < 60; i++ {
currentConns := startConnections + (step * i)
for j := 0; j < step; j++ {
go func() {
for {
select {
case <-ctx.Done():
return
default:
executeOLTPTransaction(pool)
}
}
}()
}
time.Sleep(1 * time.Second)
log.Printf("Second %d: %d connections, ACU: %.1f", i, currentConns, getCurrentACU())
}
}
Scaling behavior observed:
| Time (seconds) | Connections | ACU | Latency p99 | Notes |
|---|---|---|---|---|
| 0 | 10 | 4.0 | 5ms | Baseline |
| 15 | 130 | 8.5 | 12ms | Scaling begins |
| 30 | 255 | 18.0 | 28ms | Mid-ramp |
| 45 | 380 | 32.0 | 45ms | Scaling accelerating |
| 60 | 500 | 48.0 | 62ms | Near-peak |
| 90 | 500 | 56.0 | 24ms | Fully scaled |
| 120 | 500 | 58.0 | 19ms | Optimized |
Key finding: Serverless v2 takes 90-120 seconds to fully scale to meet sudden demand spikes. During this window, latency increases 3-5x above steady-state levels. This is acceptable for most web workloads but problematic for latency-sensitive financial applications.
Benchmark 3: Connection Storm Resilience
What happens when 500 connections arrive simultaneously (not ramped)?
# Simulate connection storm
pgbench -h aurora-cluster.cluster-xyz.us-east-1.rds.amazonaws.com \
-U benchuser -d benchdb \
-c 500 -j 50 -T 300 \
-C # Reconnect mode — creates new connections continuously
| Metric | Provisioned | Serverless v2 |
|---|---|---|
| Connection establishment p50 | 3.2ms | 8.4ms |
| Connection establishment p99 | 12ms | 340ms |
| Failed connections (5 min) | 0 | 23 |
| Latency spike duration | None | 45 seconds |
| Error rate during spike | 0% | 0.4% |
Serverless v2 rejected 23 connections during the initial spike while scaling ACU from 4 to 48. The provisioned instance handled it cleanly because capacity was pre-allocated.
Mitigation: Set a minimum ACU floor based on your expected baseline load. We found minCapacity: 8 eliminated connection failures for our SaaS workload.
Benchmark 4: Scale-to-Zero and Cold Start
With minCapacity: 0 (pause after inactivity):
| Metric | Value |
|---|---|
| Time to pause (no connections) | ~5 minutes |
| Cold resume time (first query) | 12-25 seconds |
| Warm resume (recent pause) | 3-8 seconds |
| Resume failure rate | 0.2% |
Cold resume takes 12-25 seconds — unacceptable for production web apps. However, for batch analytics that run on a schedule, this eliminates 100% of idle costs.
-- First query after cold resume takes 12-25 seconds
-- Subsequent queries are normal latency
SELECT count(*) FROM orders WHERE created_at > now() - interval '1 day';
-- Time: 14,234ms (first), 3ms (subsequent)
Cost Modeling: When Serverless v2 Wins
We modeled costs across different traffic patterns:
| Workload Pattern | Provisioned Cost | Serverless v2 Cost | Winner |
|---|---|---|---|
| Steady 24/7 high load | $1,640/mo | $2,448/mo | Provisioned (-33%) |
| Business hours only (10hr/day) | $1,640/mo | $1,020/mo | Serverless (-38%) |
| Variable (3x daily spikes) | $3,280/mo (sized for peak) | $1,860/mo | Serverless (-43%) |
| Dev/staging environments | $820/mo | $180/mo | Serverless (-78%) |
| Weekend batch processing | $1,640/mo | $290/mo | Serverless (-82%) |
The formula: Serverless v2 saves money when utilization is below 65% of peak capacity on average.
// Cost estimator
function estimateMonthlyCost(params: {
peakACU: number;
averageACU: number;
hoursPerDay: number;
daysPerMonth: number;
}): { provisioned: number; serverless: number } {
const acuHourCost = 0.12; // us-east-1, PostgreSQL
const provisionedHourCost = params.peakACU * acuHourCost;
// Provisioned: pay for peak 24/7
const provisioned = provisionedHourCost * 730; // hours/month
// Serverless: pay for actual usage
const serverless = params.averageACU * acuHourCost * params.hoursPerDay * params.daysPerMonth;
return { provisioned, serverless };
}
// Example: SaaS with business-hours traffic
const result = estimateMonthlyCost({
peakACU: 32,
averageACU: 14,
hoursPerDay: 12,
daysPerMonth: 22,
});
// { provisioned: $2,803, serverless: $443 }
Production Configuration Recommendations
Based on 14 months of operation, here is our recommended configuration:
resource "aws_rds_cluster" "production" {
cluster_identifier = "production-api"
engine = "aurora-postgresql"
engine_mode = "provisioned"
engine_version = "15.4"
serverlessv2_scaling_configuration {
min_capacity = 8 # Never below baseline — prevents cold scaling issues
max_capacity = 64 # 2x expected peak — safety margin
}
# Performance Insights for ACU right-sizing
performance_insights_enabled = true
performance_insights_retention_period = 731 # 2 years
# Enhanced monitoring for scaling correlation
monitoring_interval = 5
monitoring_role_arn = aws_iam_role.rds_monitoring.arn
}
resource "aws_rds_cluster_instance" "writer" {
cluster_identifier = aws_rds_cluster.production.id
instance_class = "db.serverless"
engine = aws_rds_cluster.production.engine
engine_version = aws_rds_cluster.production.engine_version
}
resource "aws_rds_cluster_instance" "reader" {
count = 2
cluster_identifier = aws_rds_cluster.production.id
instance_class = "db.serverless"
engine = aws_rds_cluster.production.engine
engine_version = aws_rds_cluster.production.engine_version
}
Key Takeaways
- Serverless v2 is not cheaper at steady state: For constant high-load workloads, provisioned instances cost 30-50% less. Serverless shines for variable workloads.
- Set a minimum ACU floor: Never set
minCapacity: 0in production web apps. The 12-25s cold resume is unacceptable for user-facing traffic. - Scale-up takes 90-120 seconds: Budget for latency spikes during scaling events. If your SLA cannot tolerate this, use provisioned.
- Connection storms expose scaling lag: Use RDS Proxy or application-level connection pooling (PgBouncer) to buffer connection bursts.
- The sweet spot is variable workloads at 30-65% average utilization: Below 30%, consider scale-to-zero. Above 65%, consider provisioned.
Aurora Serverless v2 is a capacity planning tool, not a cost reduction tool. It eliminates over-provisioning waste for unpredictable workloads. If your workload is predictable, stick with provisioned instances and save the premium.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.