Aurora Serverless v2 Benchmarks: When Auto-Scaling Meets Production Reality

Performance benchmarks of Aurora Serverless v2 under varying load patterns including cold scaling, connection storms, and cost comparison with provisioned instances.

#aws#aurora#database#serverless
Cover image for the article: Aurora Serverless v2 Benchmarks: When Auto-Scaling Meets Production Reality

Aurora Serverless v2 promises automatic scaling without capacity planning. After running it in production for 14 months across three workloads — a SaaS platform, batch analytics, and a microservices backend — I have benchmark data that reveals where it delivers on that promise and where it falls short.

The short version: it is excellent for variable workloads with predictable baseline demand. It is expensive for steady-state workloads and problematic for sudden spike scenarios.

Test Environment

  • Engine: Aurora PostgreSQL 15.4
  • ACU range: 2-64 (configurable min/max)
  • Region: us-east-1
  • Client: pgbench and custom Go load generator on EC2 (c6i.4xlarge)
  • Comparison baseline: db.r6g.2xlarge provisioned (8 vCPU, 64GB RAM)

An ACU (Aurora Capacity Unit) equals approximately 2GB of RAM with corresponding CPU. The provisioned baseline (8 vCPU, 64GB) maps to roughly 32 ACU.

Benchmark 1: Steady-State OLTP Performance

Running pgbench with 100 concurrent connections at sustained load for 1 hour:

pgbench -h aurora-cluster.cluster-xyz.us-east-1.rds.amazonaws.com \
  -U benchuser -d benchdb \
  -c 100 -j 16 -T 3600 \
  --protocol=prepared \
  -f custom_oltp.sql

Results:

MetricProvisioned (r6g.2xl)Serverless v2 (2-64 ACU)Delta
TPS (transactions/sec)14,20013,800-2.8%
Latency p506.8ms7.1ms+4.4%
Latency p9918.2ms19.8ms+8.8%
Settled ACU-34 ACU-
Monthly cost$1,640$2,448+49.3%

At steady state, Serverless v2 performs within 5-10% of provisioned but costs 49% more. The ACU settled at 34 (just above our 32 ACU baseline equivalent) and stayed there — demonstrating that for steady workloads, you pay a premium for the scaling capability you are not using.

Aurora Steady State Performance Comparison

Benchmark 2: Scaling Response Time

The critical question: how fast does Serverless v2 scale when load increases suddenly?

Test: ramp from 10 concurrent connections to 500 over 60 seconds.

// load-generator.go — Ramp test
func rampTest(ctx context.Context, pool *pgxpool.Pool) {
    startConnections := 10
    endConnections := 500
    rampDuration := 60 * time.Second
    step := (endConnections - startConnections) / 60

    for i := 0; i < 60; i++ {
        currentConns := startConnections + (step * i)
        for j := 0; j < step; j++ {
            go func() {
                for {
                    select {
                    case <-ctx.Done():
                        return
                    default:
                        executeOLTPTransaction(pool)
                    }
                }
            }()
        }
        time.Sleep(1 * time.Second)
        log.Printf("Second %d: %d connections, ACU: %.1f", i, currentConns, getCurrentACU())
    }
}

Scaling behavior observed:

Time (seconds)ConnectionsACULatency p99Notes
0104.05msBaseline
151308.512msScaling begins
3025518.028msMid-ramp
4538032.045msScaling accelerating
6050048.062msNear-peak
9050056.024msFully scaled
12050058.019msOptimized

Key finding: Serverless v2 takes 90-120 seconds to fully scale to meet sudden demand spikes. During this window, latency increases 3-5x above steady-state levels. This is acceptable for most web workloads but problematic for latency-sensitive financial applications.

Aurora Serverless v2 Scaling Response

Benchmark 3: Connection Storm Resilience

What happens when 500 connections arrive simultaneously (not ramped)?

# Simulate connection storm
pgbench -h aurora-cluster.cluster-xyz.us-east-1.rds.amazonaws.com \
  -U benchuser -d benchdb \
  -c 500 -j 50 -T 300 \
  -C  # Reconnect mode — creates new connections continuously
MetricProvisionedServerless v2
Connection establishment p503.2ms8.4ms
Connection establishment p9912ms340ms
Failed connections (5 min)023
Latency spike durationNone45 seconds
Error rate during spike0%0.4%

Serverless v2 rejected 23 connections during the initial spike while scaling ACU from 4 to 48. The provisioned instance handled it cleanly because capacity was pre-allocated.

Mitigation: Set a minimum ACU floor based on your expected baseline load. We found minCapacity: 8 eliminated connection failures for our SaaS workload.

Benchmark 4: Scale-to-Zero and Cold Start

With minCapacity: 0 (pause after inactivity):

MetricValue
Time to pause (no connections)~5 minutes
Cold resume time (first query)12-25 seconds
Warm resume (recent pause)3-8 seconds
Resume failure rate0.2%

Cold resume takes 12-25 seconds — unacceptable for production web apps. However, for batch analytics that run on a schedule, this eliminates 100% of idle costs.

-- First query after cold resume takes 12-25 seconds
-- Subsequent queries are normal latency
SELECT count(*) FROM orders WHERE created_at > now() - interval '1 day';
-- Time: 14,234ms (first), 3ms (subsequent)

Cost Modeling: When Serverless v2 Wins

We modeled costs across different traffic patterns:

Workload PatternProvisioned CostServerless v2 CostWinner
Steady 24/7 high load$1,640/mo$2,448/moProvisioned (-33%)
Business hours only (10hr/day)$1,640/mo$1,020/moServerless (-38%)
Variable (3x daily spikes)$3,280/mo (sized for peak)$1,860/moServerless (-43%)
Dev/staging environments$820/mo$180/moServerless (-78%)
Weekend batch processing$1,640/mo$290/moServerless (-82%)

The formula: Serverless v2 saves money when utilization is below 65% of peak capacity on average.

// Cost estimator
function estimateMonthlyCost(params: {
  peakACU: number;
  averageACU: number;
  hoursPerDay: number;
  daysPerMonth: number;
}): { provisioned: number; serverless: number } {
  const acuHourCost = 0.12; // us-east-1, PostgreSQL
  const provisionedHourCost = params.peakACU * acuHourCost;

  // Provisioned: pay for peak 24/7
  const provisioned = provisionedHourCost * 730; // hours/month

  // Serverless: pay for actual usage
  const serverless = params.averageACU * acuHourCost * params.hoursPerDay * params.daysPerMonth;

  return { provisioned, serverless };
}

// Example: SaaS with business-hours traffic
const result = estimateMonthlyCost({
  peakACU: 32,
  averageACU: 14,
  hoursPerDay: 12,
  daysPerMonth: 22,
});
// { provisioned: $2,803, serverless: $443 }

Production Configuration Recommendations

Based on 14 months of operation, here is our recommended configuration:

resource "aws_rds_cluster" "production" {
  cluster_identifier = "production-api"
  engine             = "aurora-postgresql"
  engine_mode        = "provisioned"
  engine_version     = "15.4"

  serverlessv2_scaling_configuration {
    min_capacity = 8    # Never below baseline — prevents cold scaling issues
    max_capacity = 64   # 2x expected peak — safety margin
  }

  # Performance Insights for ACU right-sizing
  performance_insights_enabled    = true
  performance_insights_retention_period = 731  # 2 years

  # Enhanced monitoring for scaling correlation
  monitoring_interval = 5
  monitoring_role_arn = aws_iam_role.rds_monitoring.arn
}

resource "aws_rds_cluster_instance" "writer" {
  cluster_identifier = aws_rds_cluster.production.id
  instance_class     = "db.serverless"
  engine             = aws_rds_cluster.production.engine
  engine_version     = aws_rds_cluster.production.engine_version
}

resource "aws_rds_cluster_instance" "reader" {
  count              = 2
  cluster_identifier = aws_rds_cluster.production.id
  instance_class     = "db.serverless"
  engine             = aws_rds_cluster.production.engine
  engine_version     = aws_rds_cluster.production.engine_version
}

Key Takeaways

  1. Serverless v2 is not cheaper at steady state: For constant high-load workloads, provisioned instances cost 30-50% less. Serverless shines for variable workloads.
  2. Set a minimum ACU floor: Never set minCapacity: 0 in production web apps. The 12-25s cold resume is unacceptable for user-facing traffic.
  3. Scale-up takes 90-120 seconds: Budget for latency spikes during scaling events. If your SLA cannot tolerate this, use provisioned.
  4. Connection storms expose scaling lag: Use RDS Proxy or application-level connection pooling (PgBouncer) to buffer connection bursts.
  5. The sweet spot is variable workloads at 30-65% average utilization: Below 30%, consider scale-to-zero. Above 65%, consider provisioned.

Aurora Serverless v2 is a capacity planning tool, not a cost reduction tool. It eliminates over-provisioning waste for unpredictable workloads. If your workload is predictable, stick with provisioned instances and save the premium.

Comments

    No comments yet. Be the first to share your thoughts.