AWS RDS Proxy Connection Pooling: Surviving Lambda Concurrency Spikes

How RDS Proxy eliminated connection exhaustion during Lambda bursts, reducing database errors by 99.7% and cutting connection setup latency from 45ms to 3ms.

#aws#rds#database#connection-pooling
Cover image for the article: AWS RDS Proxy Connection Pooling: Surviving Lambda Concurrency Spikes

At 09:14 on a Tuesday, our order processing system started throwing too many connections errors. CloudWatch showed Lambda concurrency spike from 50 to 1,200 concurrent executions in 90 seconds as a marketing email landed in 200,000 inboxes simultaneously. Each Lambda created a fresh PostgreSQL connection, and our RDS instance's max_connections of 500 was overwhelmed instantly.

After deploying RDS Proxy, we handle 3,000+ concurrent Lambda executions against a single RDS instance without connection exhaustion. Here is the architecture, configuration, and the performance data that surprised us.

The Problem: Lambda's Connection Model Breaks Traditional Databases

Lambda functions are stateless. Each invocation creates a new database connection, uses it for 50-200ms, and discards it. Under steady load, this works because connections are reused across warm invocations. But during cold start bursts:

  • 1,000 simultaneous cold starts = 1,000 simultaneous connection attempts
  • PostgreSQL connection setup takes 30-50ms (TLS handshake, auth, session initialization)
  • Each connection consumes ~10MB of database memory
  • max_connections is typically 500-2000 depending on instance size

The math simply does not work. You either over-provision your database (expensive) or accept connection failures during spikes (unacceptable).

Lambda Connection Storm

Architecture: RDS Proxy as Connection Multiplexer

RDS Proxy sits between Lambda and RDS, maintaining a warm pool of database connections. Multiple Lambda invocations share connections from this pool through multiplexing:

  • Lambda requests a connection from RDS Proxy (3ms)
  • RDS Proxy assigns an existing warm connection from its pool
  • Lambda executes queries
  • Lambda releases the connection back to the pool (not closed)
  • RDS Proxy keeps the connection warm for the next Lambda

This transforms 3,000 ephemeral Lambda connections into 100-200 persistent database connections.

Configuration: Tuning for Lambda Workloads

The default RDS Proxy settings are optimized for traditional applications, not Lambda. We tuned several parameters:

import * as cdk from 'aws-cdk-lib';
import * as rds from 'aws-cdk-lib/aws-rds';
import * as ec2 from 'aws-cdk-lib/aws-ec2';

const proxy = new rds.DatabaseProxy(this, 'OrderDbProxy', {
  proxyTarget: rds.ProxyTarget.fromInstance(dbInstance),
  vpc,
  secrets: [dbInstance.secret!],
  dbProxyName: 'order-db-proxy',
  requireTLS: true,
  idleClientTimeout: cdk.Duration.minutes(3),
  maxConnectionsPercent: 90,
  maxIdleConnectionsPercent: 20,
  borrowTimeout: cdk.Duration.seconds(30),
  vpcSubnets: { subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS },
  securityGroups: [proxySecurityGroup],
});

// Grant Lambda access
proxy.grantConnect(orderProcessorLambda, 'app_user');

Key tuning decisions:

  • maxConnectionsPercent: 90 — Use up to 90% of RDS max_connections. Reserve 10% for admin access and monitoring.
  • maxIdleConnectionsPercent: 20 — Keep 20% of connections idle-warm for burst absorption.
  • idleClientTimeout: 3 minutes — Lambda functions holding connections without activity get disconnected after 3 minutes. Catches hanging invocations.
  • borrowTimeout: 30 seconds — If no connection is available, wait up to 30 seconds before failing. Smooths short spikes.

Lambda-Side Connection Handling

The Lambda code must cooperate with RDS Proxy for optimal performance:

import { Client } from 'pg';
import { Signer } from '@aws-sdk/rds-signer';

let cachedClient: Client | null = null;

async function getConnection(): Promise<Client> {
  if (cachedClient && !cachedClient.ended) {
    // Validate cached connection is still alive
    try {
      await cachedClient.query('SELECT 1');
      return cachedClient;
    } catch {
      cachedClient = null;
    }
  }

  const signer = new Signer({
    hostname: process.env.PROXY_ENDPOINT!,
    port: 5432,
    username: 'app_user',
    region: 'us-east-1',
  });

  const token = await signer.getAuthToken();

  cachedClient = new Client({
    host: process.env.PROXY_ENDPOINT,
    port: 5432,
    user: 'app_user',
    password: token,
    database: 'orders',
    ssl: { rejectUnauthorized: true },
    connectionTimeoutMillis: 5000,
    query_timeout: 10000,
  });

  await cachedClient.connect();
  return cachedClient;
}

export async function handler(event: OrderEvent): Promise<OrderResult> {
  const client = await getConnection();

  try {
    await client.query('BEGIN');
    const result = await client.query(
      'INSERT INTO orders (customer_id, amount, status) VALUES ($1, $2, $3) RETURNING id',
      [event.customerId, event.amount, 'pending']
    );
    await client.query('COMMIT');

    return { orderId: result.rows[0].id, status: 'created' };
  } catch (error) {
    await client.query('ROLLBACK');
    throw error;
  }
  // Note: we do NOT close the connection — reuse across invocations
}

Critical patterns here:

  • Connection caching between invocations (reuse in warm Lambda)
  • IAM authentication instead of password (tokens rotate automatically, no secrets in env vars)
  • Connection validation before reuse (detect stale connections)
  • Never close the connection at end of invocation (let RDS Proxy manage lifecycle)

Performance Benchmarks

We ran a controlled test simulating our Tuesday morning incident: 0 to 3,000 concurrent Lambda executions in 60 seconds, sustained for 10 minutes:

MetricDirect RDSWith RDS ProxyImprovement
Max concurrent Lambdas (no errors)4803,200+6.7x
Connection setup time (P50)45ms3ms93.3% faster
Connection setup time (P99)180ms12ms93.3% faster
Database errors during burst2,340799.7% reduction
Active DB connections at peak500 (max)18762.6% fewer
Total query latency (P99)890ms124ms86.1% faster
Monthly cost delta$0$432—

The P99 query latency improvement comes from eliminating connection setup from the critical path. When Lambda hits RDS Proxy with a cached connection, the first query executes immediately instead of waiting 45-180ms for TCP+TLS+auth.

RDS Proxy Performance Comparison

Connection Pinning: The Hidden Gotcha

RDS Proxy multiplexes connections by detaching a database connection from a client session when the session is idle. But certain operations "pin" a connection, preventing multiplexing:

  • SET statements (SET timezone, SET search_path)
  • Prepared statements with server-side state
  • Temporary tables
  • Advisory locks

Pinned connections cannot be shared, defeating the purpose of the proxy. We reduced pinning by 94% through three changes:

  1. Moved timezone handling to the query level (AT TIME ZONE 'UTC') instead of SET timezone
  2. Used parameterized queries instead of named prepared statements
  3. Eliminated temporary tables in favor of CTEs

Monitor DatabaseConnectionsCurrentlySessionPinned in CloudWatch. If pinning exceeds 20% of your pool, investigate.

Failover Behavior

RDS Proxy automatically detects RDS failover and routes connections to the new primary. During our last failover test:

  • Failover initiated: T+0s
  • RDS Proxy detects new primary: T+6s
  • New connections routed to new primary: T+8s
  • Application errors during transition: 3 queries (auto-retried)
  • Total visible impact: 0 (retry logic absorbed the gap)

Without RDS Proxy, failover causes all existing connections to drop. Every Lambda invocation during the 15-30 second failover window fails with a connection reset error. RDS Proxy reduces visible failover impact from 30 seconds to effectively zero with proper retry logic.

Lessons Learned

IAM auth is mandatory for Lambda. Database passwords in environment variables are a security liability and operational burden (rotation). IAM auth tokens are generated per-invocation, rotate automatically, and integrate with Lambda's execution role.

Monitor ClientConnectionsNoTLS. If this metric is non-zero, something is connecting without encryption. We set a CloudWatch alarm on this metric — it caught a misconfigured health check within hours.

Right-size the proxy for your burst pattern. RDS Proxy itself has connection limits based on instance class. An rds.proxy.small supports 200 max connections. If your RDS instance allows 500 connections, ensure your proxy can utilize them.

Keep transactions short. A connection held for 5 seconds during a long transaction cannot be multiplexed. We set a hard rule: no transaction exceeds 2 seconds. Long operations are broken into multiple short transactions.

Conclusion

RDS Proxy transforms the Lambda-to-database integration from a liability into a strength. For $432/month, we eliminated connection exhaustion entirely, reduced connection latency by 93%, and gained transparent failover handling. The key is understanding that RDS Proxy is not just a connection pool — it is a connection multiplexer that requires cooperation from both your application code (avoid pinning) and your database configuration (right-size max_connections). For any Lambda-based application making more than 100 concurrent database calls, RDS Proxy is not optional — it is required infrastructure.

Comments

    No comments yet. Be the first to share your thoughts.