Beyond Provisioned Concurrency: 5 Techniques That Eliminated Cold Starts Entirely

Provisioned concurrency is expensive. Here are 5 alternative techniques we used to eliminate Lambda cold starts for latency-sensitive APIs at 40% lower cost.

#serverless#cold-start#performance#aws#optimization
Cover image for the article: Beyond Provisioned Concurrency: 5 Techniques That Eliminated Cold Starts Entirely

Cold starts are the Achilles' heel of serverless architectures. At our scale — 2,000+ Lambda invocations per second across 45 functions — cold starts were causing P99 latency spikes of 3-5 seconds on critical user-facing APIs. The obvious solution is AWS Provisioned Concurrency, but at $0.0000041667 per GB-second, keeping 200 instances warm 24/7 would cost us $14,400/month. We found better ways.

After six months of experimentation, we eliminated cold starts from our critical path using five techniques that cost us $8,600/month — a 40% savings over provisioned concurrency while achieving better P99 performance.

Understanding What Actually Causes Cold Starts

Before optimizing, we instrumented every cold start to understand the breakdown:

PhaseDuration (Node.js)Duration (Java)What Happens
Container init200-400ms200-400msAWS allocates compute, downloads code
Runtime init50-100ms800-2000msLanguage runtime boots
Dependency loading100-500ms500-3000msImport/require of packages
Handler init50-200ms200-1000msYour initialization code runs
Total cold start400-1200ms1700-6400msUser waits this long

The key insight: 80% of our cold start time was dependency loading and handler initialization, not container or runtime boot. This meant optimizations targeting those phases had the highest impact.

Technique 1: Tiered Warming with Scheduled Scaling

Instead of keeping all functions warm all the time, we built a warming system that matches our actual traffic patterns:

// warming-scheduler.ts - Predictive warming based on traffic patterns
import { LambdaClient, InvokeCommand } from '@aws-sdk/client-lambda';

interface WarmingConfig {
  functionName: string;
  baselineConcurrency: number;
  peakConcurrency: number;
  peakHoursUtc: [number, number]; // [startHour, endHour]
  rampMinutes: number; // Minutes before peak to start warming
}

const warmingConfigs: WarmingConfig[] = [
  {
    functionName: 'api-gateway-handler',
    baselineConcurrency: 20,
    peakConcurrency: 150,
    peakHoursUtc: [8, 22],    // 8 AM - 10 PM UTC (covers US + EU business)
    rampMinutes: 15,
  },
  {
    functionName: 'payment-processor',
    baselineConcurrency: 10,
    peakConcurrency: 80,
    peakHoursUtc: [13, 5],    // 1 PM UTC - 5 AM UTC (US business hours)
    rampMinutes: 10,
  },
];

async function executeWarming(config: WarmingConfig): Promise<void> {
  const currentHour = new Date().getUTCHours();
  const targetConcurrency = isInPeakHours(currentHour, config)
    ? config.peakConcurrency
    : config.baselineConcurrency;

  // Invoke N concurrent warming requests
  const warmingPayload = JSON.stringify({ _warming: true });
  const promises = Array.from({ length: targetConcurrency }, (_, i) =>
    lambda.send(new InvokeCommand({
      FunctionName: config.functionName,
      InvocationType: 'RequestResponse',
      Payload: Buffer.from(warmingPayload),
    }))
  );

  await Promise.allSettled(promises);
}

The warming function runs every 5 minutes via EventBridge Scheduler. By matching our actual traffic curve, we keep exactly the right number of instances warm:

Warming vs Traffic Pattern

Cost: $2,800/month (vs. $14,400 for 24/7 provisioned concurrency) Cold start rate: Reduced from 4.2% to 0.3% of invocations

Technique 2: Lazy Initialization with Module-Level Caching

Most Lambda functions eagerly load everything at import time. We restructured to lazy-load expensive dependencies only when specific code paths need them:

// Before: Everything loads at cold start
import { DynamoDBClient } from '@aws-sdk/client-dynamodb';
import { S3Client } from '@aws-sdk/client-s3';
import { SESClient } from '@aws-sdk/client-ses';
import Stripe from 'stripe';
import { createClient } from 'redis';

// After: Lazy initialization with singleton caching
let _dynamodb: DynamoDBClient | null = null;
let _s3: S3Client | null = null;
let _stripe: Stripe | null = null;
let _redis: ReturnType<typeof createClient> | null = null;

function getDynamoDB(): DynamoDBClient {
  if (!_dynamodb) {
    _dynamodb = new DynamoDBClient({
      region: process.env.AWS_REGION,
      maxAttempts: 3,
    });
  }
  return _dynamodb;
}

function getStripe(): Stripe {
  if (!_stripe) {
    // Only import Stripe when payment code path is hit
    const StripeModule = require('stripe');
    _stripe = new StripeModule(process.env.STRIPE_SECRET_KEY);
  }
  return _stripe;
}

// Handler only initializes what this specific request needs
export async function handler(event: APIGatewayEvent) {
  if (event.path.startsWith('/payments')) {
    // Stripe loads ONLY for payment routes
    const stripe = getStripe();
    return processPayment(event, stripe);
  }

  // Most requests only need DynamoDB
  const db = getDynamoDB();
  return processRequest(event, db);
}

Impact: Cold start time reduced by 45% (from 1200ms to 660ms average) because only essential modules load during initialization.

Technique 3: SnapStart for Java Functions (and the Node.js Equivalent)

AWS SnapStart snapshots the initialized state of a Java function and restores from that snapshot instead of cold-booting. For Java functions, this reduced cold starts from 5+ seconds to under 200ms:

// Java handler with SnapStart optimization
@SnapStart
public class ApiHandler implements RequestHandler<APIGatewayProxyRequestEvent, APIGatewayProxyResponseEvent> {

    // These initialize during the snapshot phase (once)
    private final DynamoDbClient dynamoDb = DynamoDbClient.create();
    private final ObjectMapper mapper = new ObjectMapper();

    // CRaC hook for resource restoration
    @Override
    public void beforeCheckpoint(Context<? extends Resource> context) {
        // Close non-restorable connections before snapshot
        // DynamoDB client handles this automatically
    }

    @Override
    public void afterRestore(Context<? extends Resource> context) {
        // Re-establish connections after restore
        dynamoDb.describeEndpoints(DescribeEndpointsRequest.builder().build());
    }

    @Override
    public APIGatewayProxyResponseEvent handleRequest(APIGatewayProxyRequestEvent event, Context context) {
        // Handler executes from restored snapshot - no cold start
        return processEvent(event);
    }
}

For Node.js (where SnapStart isn't available), we achieved a similar effect with our custom init-snapshot pattern using Lambda Extensions:

// Extension that pre-warms expensive initialization during INIT phase
// lambda-extension/prewarmer.ts
process.on('SIGTERM', () => process.exit(0));

async function prewarmInit(): Promise<void> {
  // Pre-resolve DNS for all known endpoints
  const endpoints = [
    'dynamodb.us-east-1.amazonaws.com',
    'sqs.us-east-1.amazonaws.com',
    's3.us-east-1.amazonaws.com',
  ];

  await Promise.all(endpoints.map(host =>
    new Promise<void>((resolve) => {
      require('dns').resolve4(host, () => resolve());
    })
  ));

  // Pre-establish TCP connections
  // These get reused across invocations
  await getDynamoDB().send(new DescribeEndpointsCommand({}));
}

Impact: Java cold starts reduced from 5.2s to 180ms. Node.js DNS pre-resolution saved 80-120ms per cold start.

Technique 4: Bundle Size Optimization with Tree Shaking

Lambda deployment package size directly correlates with cold start duration. We measured:

Package SizeCold Start (Node.js)Cold Start (Python)
5 MB350ms400ms
25 MB800ms900ms
50 MB (max zip)1400ms1600ms
250 MB (layer)2200ms2500ms

We used esbuild with aggressive tree shaking:

// build.config.ts - esbuild configuration for minimal bundles
import { build } from 'esbuild';

await build({
  entryPoints: ['src/handlers/api.ts'],
  bundle: true,
  minify: true,
  treeShaking: true,
  platform: 'node',
  target: 'node20',
  outfile: 'dist/api.js',
  external: [
    '@aws-sdk/*', // Use Lambda runtime's built-in SDK
  ],
  metafile: true, // Analyze bundle composition
  define: {
    'process.env.NODE_ENV': '"production"',
  },
});

Key decisions:

  • External AWS SDK: Lambda runtime includes AWS SDK v3 — don't bundle it
  • Single-file bundles: One file loads faster than node_modules traversal
  • Minification: Reduces parse time, not just file size

Impact: Average bundle size reduced from 28MB to 3.2MB. Cold starts improved by 400ms.

Technique 5: Response Streaming for Perceived Performance

For requests where total processing time exceeds 1 second regardless of cold starts, we used Lambda Response Streaming to send the first byte immediately:

// Streaming response - user sees data before function finishes
import { streamifyResponse, ResponseStream } from 'lambda-stream';

export const handler = streamifyResponse(
  async (event: APIGatewayEvent, responseStream: ResponseStream) => {
    // Send headers immediately (even during cold start)
    responseStream.setContentType('application/json');

    // Start streaming partial results
    responseStream.write('{"status":"processing","results":[');

    const items = await fetchItems(event.queryStringParameters);

    for (let i = 0; i < items.length; i++) {
      const processed = await enrichItem(items[i]);
      responseStream.write(JSON.stringify(processed));
      if (i < items.length - 1) responseStream.write(',');
    }

    responseStream.write(']}');
    responseStream.end();
  }
);

This doesn't eliminate cold starts technically, but reduces perceived latency because users receive the first byte within 200ms even if the function took 1.5s to fully initialize.

Combined Results

After implementing all five techniques across our 45 production Lambda functions:

MetricBeforeAfterImprovement
Cold start rate4.2%0.08%98% reduction
P99 latency4,800ms420ms91% improvement
P95 latency1,200ms180ms85% improvement
Monthly cost (warming)$0 (no warming)$8,600N/A
Monthly cost (provisioned concurrency equivalent)$14,400N/A40% savings

Cold Start Elimination Results

When to Still Use Provisioned Concurrency

Despite our optimizations, provisioned concurrency remains the right choice for:

  • Sub-50ms P99 requirements: When even 200ms warm starts are too slow
  • Regulatory compliance: Some financial services require guaranteed response times
  • Predictable, flat traffic: If your traffic doesn't have peaks and valleys, the economics change
  • Java/heavy runtimes: Where cold starts exceed 5s and SnapStart isn't available for your configuration

Key Takeaways

  1. Provisioned concurrency is a blunt instrument — predictive warming with traffic-pattern matching gives the same benefit at 40% lower cost.

  2. Dependency loading dominates cold start time — lazy initialization and tree shaking have more impact than runtime choice.

  3. Bundle size matters linearly — every 10MB reduction saves ~200ms of cold start time. Use the built-in AWS SDK.

  4. SnapStart transforms Java serverless — if you're running Java Lambdas, this is the single highest-impact optimization available.

  5. Measure from the user's perspective — response streaming can eliminate perceived cold starts even when they still occur technically.

  6. Combine techniques for compounding gains — no single technique eliminates cold starts. The combination of all five brought us from 4.2% to 0.08% cold start rate.

Cold starts are solvable without throwing money at provisioned concurrency. The techniques here require more engineering effort upfront but pay for themselves within two months through reduced infrastructure spend.

Comments

    No comments yet. Be the first to share your thoughts.