Beyond Provisioned Concurrency: 5 Techniques That Eliminated Cold Starts Entirely
Provisioned concurrency is expensive. Here are 5 alternative techniques we used to eliminate Lambda cold starts for latency-sensitive APIs at 40% lower cost.

Cold starts are the Achilles' heel of serverless architectures. At our scale — 2,000+ Lambda invocations per second across 45 functions — cold starts were causing P99 latency spikes of 3-5 seconds on critical user-facing APIs. The obvious solution is AWS Provisioned Concurrency, but at $0.0000041667 per GB-second, keeping 200 instances warm 24/7 would cost us $14,400/month. We found better ways.
After six months of experimentation, we eliminated cold starts from our critical path using five techniques that cost us $8,600/month — a 40% savings over provisioned concurrency while achieving better P99 performance.
Understanding What Actually Causes Cold Starts
Before optimizing, we instrumented every cold start to understand the breakdown:
| Phase | Duration (Node.js) | Duration (Java) | What Happens |
|---|---|---|---|
| Container init | 200-400ms | 200-400ms | AWS allocates compute, downloads code |
| Runtime init | 50-100ms | 800-2000ms | Language runtime boots |
| Dependency loading | 100-500ms | 500-3000ms | Import/require of packages |
| Handler init | 50-200ms | 200-1000ms | Your initialization code runs |
| Total cold start | 400-1200ms | 1700-6400ms | User waits this long |
The key insight: 80% of our cold start time was dependency loading and handler initialization, not container or runtime boot. This meant optimizations targeting those phases had the highest impact.
Technique 1: Tiered Warming with Scheduled Scaling
Instead of keeping all functions warm all the time, we built a warming system that matches our actual traffic patterns:
// warming-scheduler.ts - Predictive warming based on traffic patterns
import { LambdaClient, InvokeCommand } from '@aws-sdk/client-lambda';
interface WarmingConfig {
functionName: string;
baselineConcurrency: number;
peakConcurrency: number;
peakHoursUtc: [number, number]; // [startHour, endHour]
rampMinutes: number; // Minutes before peak to start warming
}
const warmingConfigs: WarmingConfig[] = [
{
functionName: 'api-gateway-handler',
baselineConcurrency: 20,
peakConcurrency: 150,
peakHoursUtc: [8, 22], // 8 AM - 10 PM UTC (covers US + EU business)
rampMinutes: 15,
},
{
functionName: 'payment-processor',
baselineConcurrency: 10,
peakConcurrency: 80,
peakHoursUtc: [13, 5], // 1 PM UTC - 5 AM UTC (US business hours)
rampMinutes: 10,
},
];
async function executeWarming(config: WarmingConfig): Promise<void> {
const currentHour = new Date().getUTCHours();
const targetConcurrency = isInPeakHours(currentHour, config)
? config.peakConcurrency
: config.baselineConcurrency;
// Invoke N concurrent warming requests
const warmingPayload = JSON.stringify({ _warming: true });
const promises = Array.from({ length: targetConcurrency }, (_, i) =>
lambda.send(new InvokeCommand({
FunctionName: config.functionName,
InvocationType: 'RequestResponse',
Payload: Buffer.from(warmingPayload),
}))
);
await Promise.allSettled(promises);
}
The warming function runs every 5 minutes via EventBridge Scheduler. By matching our actual traffic curve, we keep exactly the right number of instances warm:
Cost: $2,800/month (vs. $14,400 for 24/7 provisioned concurrency) Cold start rate: Reduced from 4.2% to 0.3% of invocations
Technique 2: Lazy Initialization with Module-Level Caching
Most Lambda functions eagerly load everything at import time. We restructured to lazy-load expensive dependencies only when specific code paths need them:
// Before: Everything loads at cold start
import { DynamoDBClient } from '@aws-sdk/client-dynamodb';
import { S3Client } from '@aws-sdk/client-s3';
import { SESClient } from '@aws-sdk/client-ses';
import Stripe from 'stripe';
import { createClient } from 'redis';
// After: Lazy initialization with singleton caching
let _dynamodb: DynamoDBClient | null = null;
let _s3: S3Client | null = null;
let _stripe: Stripe | null = null;
let _redis: ReturnType<typeof createClient> | null = null;
function getDynamoDB(): DynamoDBClient {
if (!_dynamodb) {
_dynamodb = new DynamoDBClient({
region: process.env.AWS_REGION,
maxAttempts: 3,
});
}
return _dynamodb;
}
function getStripe(): Stripe {
if (!_stripe) {
// Only import Stripe when payment code path is hit
const StripeModule = require('stripe');
_stripe = new StripeModule(process.env.STRIPE_SECRET_KEY);
}
return _stripe;
}
// Handler only initializes what this specific request needs
export async function handler(event: APIGatewayEvent) {
if (event.path.startsWith('/payments')) {
// Stripe loads ONLY for payment routes
const stripe = getStripe();
return processPayment(event, stripe);
}
// Most requests only need DynamoDB
const db = getDynamoDB();
return processRequest(event, db);
}
Impact: Cold start time reduced by 45% (from 1200ms to 660ms average) because only essential modules load during initialization.
Technique 3: SnapStart for Java Functions (and the Node.js Equivalent)
AWS SnapStart snapshots the initialized state of a Java function and restores from that snapshot instead of cold-booting. For Java functions, this reduced cold starts from 5+ seconds to under 200ms:
// Java handler with SnapStart optimization
@SnapStart
public class ApiHandler implements RequestHandler<APIGatewayProxyRequestEvent, APIGatewayProxyResponseEvent> {
// These initialize during the snapshot phase (once)
private final DynamoDbClient dynamoDb = DynamoDbClient.create();
private final ObjectMapper mapper = new ObjectMapper();
// CRaC hook for resource restoration
@Override
public void beforeCheckpoint(Context<? extends Resource> context) {
// Close non-restorable connections before snapshot
// DynamoDB client handles this automatically
}
@Override
public void afterRestore(Context<? extends Resource> context) {
// Re-establish connections after restore
dynamoDb.describeEndpoints(DescribeEndpointsRequest.builder().build());
}
@Override
public APIGatewayProxyResponseEvent handleRequest(APIGatewayProxyRequestEvent event, Context context) {
// Handler executes from restored snapshot - no cold start
return processEvent(event);
}
}
For Node.js (where SnapStart isn't available), we achieved a similar effect with our custom init-snapshot pattern using Lambda Extensions:
// Extension that pre-warms expensive initialization during INIT phase
// lambda-extension/prewarmer.ts
process.on('SIGTERM', () => process.exit(0));
async function prewarmInit(): Promise<void> {
// Pre-resolve DNS for all known endpoints
const endpoints = [
'dynamodb.us-east-1.amazonaws.com',
'sqs.us-east-1.amazonaws.com',
's3.us-east-1.amazonaws.com',
];
await Promise.all(endpoints.map(host =>
new Promise<void>((resolve) => {
require('dns').resolve4(host, () => resolve());
})
));
// Pre-establish TCP connections
// These get reused across invocations
await getDynamoDB().send(new DescribeEndpointsCommand({}));
}
Impact: Java cold starts reduced from 5.2s to 180ms. Node.js DNS pre-resolution saved 80-120ms per cold start.
Technique 4: Bundle Size Optimization with Tree Shaking
Lambda deployment package size directly correlates with cold start duration. We measured:
| Package Size | Cold Start (Node.js) | Cold Start (Python) |
|---|---|---|
| 5 MB | 350ms | 400ms |
| 25 MB | 800ms | 900ms |
| 50 MB (max zip) | 1400ms | 1600ms |
| 250 MB (layer) | 2200ms | 2500ms |
We used esbuild with aggressive tree shaking:
// build.config.ts - esbuild configuration for minimal bundles
import { build } from 'esbuild';
await build({
entryPoints: ['src/handlers/api.ts'],
bundle: true,
minify: true,
treeShaking: true,
platform: 'node',
target: 'node20',
outfile: 'dist/api.js',
external: [
'@aws-sdk/*', // Use Lambda runtime's built-in SDK
],
metafile: true, // Analyze bundle composition
define: {
'process.env.NODE_ENV': '"production"',
},
});
Key decisions:
- External AWS SDK: Lambda runtime includes AWS SDK v3 — don't bundle it
- Single-file bundles: One file loads faster than node_modules traversal
- Minification: Reduces parse time, not just file size
Impact: Average bundle size reduced from 28MB to 3.2MB. Cold starts improved by 400ms.
Technique 5: Response Streaming for Perceived Performance
For requests where total processing time exceeds 1 second regardless of cold starts, we used Lambda Response Streaming to send the first byte immediately:
// Streaming response - user sees data before function finishes
import { streamifyResponse, ResponseStream } from 'lambda-stream';
export const handler = streamifyResponse(
async (event: APIGatewayEvent, responseStream: ResponseStream) => {
// Send headers immediately (even during cold start)
responseStream.setContentType('application/json');
// Start streaming partial results
responseStream.write('{"status":"processing","results":[');
const items = await fetchItems(event.queryStringParameters);
for (let i = 0; i < items.length; i++) {
const processed = await enrichItem(items[i]);
responseStream.write(JSON.stringify(processed));
if (i < items.length - 1) responseStream.write(',');
}
responseStream.write(']}');
responseStream.end();
}
);
This doesn't eliminate cold starts technically, but reduces perceived latency because users receive the first byte within 200ms even if the function took 1.5s to fully initialize.
Combined Results
After implementing all five techniques across our 45 production Lambda functions:
| Metric | Before | After | Improvement |
|---|---|---|---|
| Cold start rate | 4.2% | 0.08% | 98% reduction |
| P99 latency | 4,800ms | 420ms | 91% improvement |
| P95 latency | 1,200ms | 180ms | 85% improvement |
| Monthly cost (warming) | $0 (no warming) | $8,600 | N/A |
| Monthly cost (provisioned concurrency equivalent) | $14,400 | N/A | 40% savings |
When to Still Use Provisioned Concurrency
Despite our optimizations, provisioned concurrency remains the right choice for:
- Sub-50ms P99 requirements: When even 200ms warm starts are too slow
- Regulatory compliance: Some financial services require guaranteed response times
- Predictable, flat traffic: If your traffic doesn't have peaks and valleys, the economics change
- Java/heavy runtimes: Where cold starts exceed 5s and SnapStart isn't available for your configuration
Key Takeaways
-
Provisioned concurrency is a blunt instrument — predictive warming with traffic-pattern matching gives the same benefit at 40% lower cost.
-
Dependency loading dominates cold start time — lazy initialization and tree shaking have more impact than runtime choice.
-
Bundle size matters linearly — every 10MB reduction saves ~200ms of cold start time. Use the built-in AWS SDK.
-
SnapStart transforms Java serverless — if you're running Java Lambdas, this is the single highest-impact optimization available.
-
Measure from the user's perspective — response streaming can eliminate perceived cold starts even when they still occur technically.
-
Combine techniques for compounding gains — no single technique eliminates cold starts. The combination of all five brought us from 4.2% to 0.08% cold start rate.
Cold starts are solvable without throwing money at provisioned concurrency. The techniques here require more engineering effort upfront but pay for themselves within two months through reduced infrastructure spend.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.