WebSocket APIs at Scale: Handling 100K Concurrent Connections on API Gateway
Architecture patterns for scaling API Gateway WebSocket APIs to 100K+ concurrent connections — connection management, fan-out optimization, and cost analysis.

The Real-Time Challenge
Every SaaS product eventually needs real-time: live dashboards, collaborative editing, notifications, chat, or streaming data feeds. The architectural question isn't whether to add WebSockets — it's how to scale them without rebuilding your infrastructure.
API Gateway WebSocket APIs provide managed WebSocket infrastructure with automatic scaling, but the documentation undersells the architectural decisions required at scale. After operating a WebSocket service handling 100K concurrent connections for a real-time analytics dashboard, here's what we learned.
Architecture: Connection Management
API Gateway handles the WebSocket lifecycle (connect, message, disconnect) by routing each event to a Lambda function. The critical design decision is your connection store — how you track which connections subscribe to which data.
We use DynamoDB with a single-table design optimized for two access patterns:
- By connection ID — For disconnect cleanup and targeted messages
- By topic/channel — For fan-out to all subscribers of a data feed
// Connection store schema (single-table design)
// PK: CONNECTION#{connectionId} | SK: META
// PK: TOPIC#{topicId} | SK: CONNECTION#{connectionId}
// GSI1PK: CONNECTION#{connectionId} | GSI1SK: TOPIC#{topicId}
import { DynamoDBDocumentClient, PutCommand, QueryCommand, BatchWriteCommand } from '@aws-sdk/lib-dynamodb';
interface ConnectionRecord {
PK: string;
SK: string;
GSI1PK?: string;
GSI1SK?: string;
connectionId: string;
topicId?: string;
userId: string;
connectedAt: string;
ttl: number; // Auto-cleanup stale connections
}
export class ConnectionStore {
constructor(private readonly ddb: DynamoDBDocumentClient) {}
async subscribe(connectionId: string, topicId: string, userId: string): Promise<void> {
const ttl = Math.floor(Date.now() / 1000) + 86400; // 24h TTL
await this.ddb.send(new PutCommand({
TableName: process.env.CONNECTIONS_TABLE!,
Item: {
PK: `TOPIC#${topicId}`,
SK: `CONNECTION#${connectionId}`,
GSI1PK: `CONNECTION#${connectionId}`,
GSI1SK: `TOPIC#${topicId}`,
connectionId,
topicId,
userId,
connectedAt: new Date().toISOString(),
ttl,
},
}));
}
async getTopicConnections(topicId: string): Promise<string[]> {
const result = await this.ddb.send(new QueryCommand({
TableName: process.env.CONNECTIONS_TABLE!,
KeyConditionExpression: 'PK = :pk',
ExpressionAttributeValues: { ':pk': `TOPIC#${topicId}` },
ProjectionExpression: 'connectionId',
}));
return result.Items?.map(i => i.connectionId) || [];
}
}
The TTL field is critical: WebSocket disconnects aren't always clean. Browsers crash, networks drop, mobile apps get killed. The TTL ensures stale connections are automatically purged without a background cleanup process.
Fan-Out: The Scaling Bottleneck
Sending a message to 100K connections means 100K PostToConnection API calls. A single Lambda function making sequential HTTP calls would take minutes. Our fan-out architecture uses parallel Lambda invocations:
// Fan-out Lambda: sends messages to batches of connections
import { ApiGatewayManagementApiClient, PostToConnectionCommand } from '@aws-sdk/client-apigatewaymanagementapi';
const apigw = new ApiGatewayManagementApiClient({
endpoint: process.env.WEBSOCKET_ENDPOINT,
});
export async function handler(event: FanOutEvent) {
const { connectionIds, payload } = event;
const message = JSON.stringify(payload);
// Process connections in parallel batches of 25
const batchSize = 25;
const staleConnections: string[] = [];
for (let i = 0; i < connectionIds.length; i += batchSize) {
const batch = connectionIds.slice(i, i + batchSize);
const results = await Promise.allSettled(
batch.map(async (connectionId) => {
try {
await apigw.send(new PostToConnectionCommand({
ConnectionId: connectionId,
Data: Buffer.from(message),
}));
} catch (error: any) {
if (error.statusCode === 410) {
// Connection is gone — mark for cleanup
staleConnections.push(connectionId);
} else {
throw error;
}
}
})
);
}
// Batch delete stale connections
if (staleConnections.length > 0) {
await cleanupConnections(staleConnections);
}
return { sent: connectionIds.length - staleConnections.length, stale: staleConnections.length };
}
For 100K connections, we split the fan-out across 100 Lambda invocations of 1,000 connections each, triggered in parallel via a Step Functions Map state. Total fan-out time: 2-4 seconds.
Benchmarks: Scaling Characteristics
We load-tested the WebSocket API from 1K to 150K concurrent connections:
| Concurrent Connections | Fan-Out Latency (p50) | Fan-Out Latency (p99) | Monthly Cost | PostToConnection Throttling |
|---|---|---|---|---|
| 1,000 | 120ms | 340ms | $89 | None |
| 10,000 | 450ms | 1.2s | $670 | Occasional |
| 50,000 | 1.8s | 4.2s | $3,100 | Frequent |
| 100,000 | 3.1s | 7.8s | $5,800 | Managed via batching |
| 150,000 | 4.5s | 12s | $8,200 | Requires limit increase |
Key finding: PostToConnection has an account-level TPS limit of 10,000 calls/second (default). At 100K connections, you need a service limit increase or connection-level message batching.
Cost Breakdown at 100K Connections
The cost model for API Gateway WebSocket APIs has three components:
Connection minutes: 100,000 connections × 720 hours × $0.25/million = $1,080/month
Messages (inbound): 5 msg/min/connection × 100K × 720h = 21.6B messages × $1.00/million = $21,600/month
Messages (outbound): 2 broadcasts/min × 100K recipients = 8.64B messages × $1.00/million = $8,640/month
Total at naive pricing: ~$31,320/month
This is expensive. Our optimization strategies:
- Message batching — Instead of sending each data point individually, batch updates every 5 seconds. Reduces message count by 80%.
- Delta compression — Send only changed fields, not full state. Reduces payload size by 60%.
- Topic segmentation — Users subscribe to specific topics, not all data. Reduces fan-out by 70%.
After optimizations: $5,800/month for 100K concurrent connections with sub-5s fan-out latency.
Connection Authentication
WebSocket $connect routes receive the full HTTP upgrade request including headers and query parameters. We authenticate at connection time and store the user context:
// $connect handler — validates token and stores connection
export async function connectHandler(event: APIGatewayProxyWebSocketEvent) {
const token = event.queryStringParameters?.token;
if (!token) {
return { statusCode: 401, body: 'Missing token' };
}
try {
const user = await verifyToken(token);
const connectionId = event.requestContext.connectionId!;
await connectionStore.register(connectionId, {
userId: user.sub,
tenantId: user.tenantId,
permissions: user.permissions,
connectedAt: new Date().toISOString(),
});
return { statusCode: 200, body: 'Connected' };
} catch (error) {
return { statusCode: 403, body: 'Invalid token' };
}
}
Critical security pattern: never trust messages from connected clients without re-validating permissions. The connection token proves identity at connect time, but topic subscriptions must be authorized against the stored user context.
Operational Patterns
- Heartbeat mechanism — API Gateway disconnects idle connections after 10 minutes. Send a server-side ping every 5 minutes to keep connections alive.
- Graceful degradation — If fan-out latency exceeds your SLA, switch to polling fallback for new connections until backpressure clears.
- Connection limits per user — Enforce max 5 connections per user to prevent resource exhaustion from misbehaving clients.
- Stale connection cleanup — Handle 410 (Gone) responses from
PostToConnectionimmediately. Don't wait for TTL expiry. - Regional deployment — Deploy WebSocket APIs in regions closest to your users. Connection latency directly impacts user experience.
When to Choose Alternatives
API Gateway WebSocket is right for:
- 1K–150K concurrent connections
- Message patterns where server broadcasts dominate
- Teams already invested in AWS serverless
Consider alternatives when:
- >500K connections: Use Elixir/Phoenix on ECS or a dedicated WebSocket service (Ably, Pusher)
- Sub-100ms fan-out required: API Gateway adds overhead. Consider AppSync subscriptions or self-managed MQTT
- Peer-to-peer messaging dominates: WebRTC or a message broker pattern may be more efficient
Conclusion
API Gateway WebSocket APIs scale to 100K+ concurrent connections with careful architecture — but the naive approach (single Lambda fan-out, no batching, no topic segmentation) hits cost and latency walls quickly.
The keys to successful scaling: DynamoDB-backed connection stores with TTL cleanup, parallel fan-out via Step Functions or batched Lambda invocations, message batching and delta compression to control costs, and strict connection authentication at the $connect route.
At $5,800/month for 100K connections with sub-5s fan-out, the total cost of ownership is competitive with self-hosted alternatives when you factor in the operational burden of managing WebSocket infrastructure yourself. The managed model wins on operational simplicity; self-hosted wins on per-message latency and cost at extreme scale.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.