WebSocket APIs at Scale: Handling 100K Concurrent Connections on API Gateway

Architecture patterns for scaling API Gateway WebSocket APIs to 100K+ concurrent connections — connection management, fan-out optimization, and cost analysis.

#aws#api-gateway#websocket#real-time
Cover image for the article: WebSocket APIs at Scale: Handling 100K Concurrent Connections on API Gateway

The Real-Time Challenge

Every SaaS product eventually needs real-time: live dashboards, collaborative editing, notifications, chat, or streaming data feeds. The architectural question isn't whether to add WebSockets — it's how to scale them without rebuilding your infrastructure.

API Gateway WebSocket APIs provide managed WebSocket infrastructure with automatic scaling, but the documentation undersells the architectural decisions required at scale. After operating a WebSocket service handling 100K concurrent connections for a real-time analytics dashboard, here's what we learned.

WebSocket Architecture at Scale

Architecture: Connection Management

API Gateway handles the WebSocket lifecycle (connect, message, disconnect) by routing each event to a Lambda function. The critical design decision is your connection store — how you track which connections subscribe to which data.

We use DynamoDB with a single-table design optimized for two access patterns:

  1. By connection ID — For disconnect cleanup and targeted messages
  2. By topic/channel — For fan-out to all subscribers of a data feed
// Connection store schema (single-table design)
// PK: CONNECTION#{connectionId} | SK: META
// PK: TOPIC#{topicId}          | SK: CONNECTION#{connectionId}
// GSI1PK: CONNECTION#{connectionId} | GSI1SK: TOPIC#{topicId}

import { DynamoDBDocumentClient, PutCommand, QueryCommand, BatchWriteCommand } from '@aws-sdk/lib-dynamodb';

interface ConnectionRecord {
  PK: string;
  SK: string;
  GSI1PK?: string;
  GSI1SK?: string;
  connectionId: string;
  topicId?: string;
  userId: string;
  connectedAt: string;
  ttl: number; // Auto-cleanup stale connections
}

export class ConnectionStore {
  constructor(private readonly ddb: DynamoDBDocumentClient) {}

  async subscribe(connectionId: string, topicId: string, userId: string): Promise<void> {
    const ttl = Math.floor(Date.now() / 1000) + 86400; // 24h TTL

    await this.ddb.send(new PutCommand({
      TableName: process.env.CONNECTIONS_TABLE!,
      Item: {
        PK: `TOPIC#${topicId}`,
        SK: `CONNECTION#${connectionId}`,
        GSI1PK: `CONNECTION#${connectionId}`,
        GSI1SK: `TOPIC#${topicId}`,
        connectionId,
        topicId,
        userId,
        connectedAt: new Date().toISOString(),
        ttl,
      },
    }));
  }

  async getTopicConnections(topicId: string): Promise<string[]> {
    const result = await this.ddb.send(new QueryCommand({
      TableName: process.env.CONNECTIONS_TABLE!,
      KeyConditionExpression: 'PK = :pk',
      ExpressionAttributeValues: { ':pk': `TOPIC#${topicId}` },
      ProjectionExpression: 'connectionId',
    }));

    return result.Items?.map(i => i.connectionId) || [];
  }
}

The TTL field is critical: WebSocket disconnects aren't always clean. Browsers crash, networks drop, mobile apps get killed. The TTL ensures stale connections are automatically purged without a background cleanup process.

Fan-Out: The Scaling Bottleneck

Sending a message to 100K connections means 100K PostToConnection API calls. A single Lambda function making sequential HTTP calls would take minutes. Our fan-out architecture uses parallel Lambda invocations:

// Fan-out Lambda: sends messages to batches of connections
import { ApiGatewayManagementApiClient, PostToConnectionCommand } from '@aws-sdk/client-apigatewaymanagementapi';

const apigw = new ApiGatewayManagementApiClient({
  endpoint: process.env.WEBSOCKET_ENDPOINT,
});

export async function handler(event: FanOutEvent) {
  const { connectionIds, payload } = event;
  const message = JSON.stringify(payload);

  // Process connections in parallel batches of 25
  const batchSize = 25;
  const staleConnections: string[] = [];

  for (let i = 0; i < connectionIds.length; i += batchSize) {
    const batch = connectionIds.slice(i, i + batchSize);
    const results = await Promise.allSettled(
      batch.map(async (connectionId) => {
        try {
          await apigw.send(new PostToConnectionCommand({
            ConnectionId: connectionId,
            Data: Buffer.from(message),
          }));
        } catch (error: any) {
          if (error.statusCode === 410) {
            // Connection is gone — mark for cleanup
            staleConnections.push(connectionId);
          } else {
            throw error;
          }
        }
      })
    );
  }

  // Batch delete stale connections
  if (staleConnections.length > 0) {
    await cleanupConnections(staleConnections);
  }

  return { sent: connectionIds.length - staleConnections.length, stale: staleConnections.length };
}

For 100K connections, we split the fan-out across 100 Lambda invocations of 1,000 connections each, triggered in parallel via a Step Functions Map state. Total fan-out time: 2-4 seconds.

Benchmarks: Scaling Characteristics

We load-tested the WebSocket API from 1K to 150K concurrent connections:

Concurrent ConnectionsFan-Out Latency (p50)Fan-Out Latency (p99)Monthly CostPostToConnection Throttling
1,000120ms340ms$89None
10,000450ms1.2s$670Occasional
50,0001.8s4.2s$3,100Frequent
100,0003.1s7.8s$5,800Managed via batching
150,0004.5s12s$8,200Requires limit increase

Key finding: PostToConnection has an account-level TPS limit of 10,000 calls/second (default). At 100K connections, you need a service limit increase or connection-level message batching.

Cost Breakdown at 100K Connections

The cost model for API Gateway WebSocket APIs has three components:

Connection minutes: 100,000 connections × 720 hours × $0.25/million = $1,080/month
Messages (inbound): 5 msg/min/connection × 100K × 720h = 21.6B messages × $1.00/million = $21,600/month
Messages (outbound): 2 broadcasts/min × 100K recipients = 8.64B messages × $1.00/million = $8,640/month

Total at naive pricing: ~$31,320/month

This is expensive. Our optimization strategies:

  1. Message batching — Instead of sending each data point individually, batch updates every 5 seconds. Reduces message count by 80%.
  2. Delta compression — Send only changed fields, not full state. Reduces payload size by 60%.
  3. Topic segmentation — Users subscribe to specific topics, not all data. Reduces fan-out by 70%.

After optimizations: $5,800/month for 100K concurrent connections with sub-5s fan-out latency.

Connection Authentication

WebSocket $connect routes receive the full HTTP upgrade request including headers and query parameters. We authenticate at connection time and store the user context:

// $connect handler — validates token and stores connection
export async function connectHandler(event: APIGatewayProxyWebSocketEvent) {
  const token = event.queryStringParameters?.token;

  if (!token) {
    return { statusCode: 401, body: 'Missing token' };
  }

  try {
    const user = await verifyToken(token);
    const connectionId = event.requestContext.connectionId!;

    await connectionStore.register(connectionId, {
      userId: user.sub,
      tenantId: user.tenantId,
      permissions: user.permissions,
      connectedAt: new Date().toISOString(),
    });

    return { statusCode: 200, body: 'Connected' };
  } catch (error) {
    return { statusCode: 403, body: 'Invalid token' };
  }
}

Critical security pattern: never trust messages from connected clients without re-validating permissions. The connection token proves identity at connect time, but topic subscriptions must be authorized against the stored user context.

Operational Patterns

  1. Heartbeat mechanism — API Gateway disconnects idle connections after 10 minutes. Send a server-side ping every 5 minutes to keep connections alive.
  2. Graceful degradation — If fan-out latency exceeds your SLA, switch to polling fallback for new connections until backpressure clears.
  3. Connection limits per user — Enforce max 5 connections per user to prevent resource exhaustion from misbehaving clients.
  4. Stale connection cleanup — Handle 410 (Gone) responses from PostToConnection immediately. Don't wait for TTL expiry.
  5. Regional deployment — Deploy WebSocket APIs in regions closest to your users. Connection latency directly impacts user experience.

When to Choose Alternatives

API Gateway WebSocket is right for:

  • 1K–150K concurrent connections
  • Message patterns where server broadcasts dominate
  • Teams already invested in AWS serverless

Consider alternatives when:

  • >500K connections: Use Elixir/Phoenix on ECS or a dedicated WebSocket service (Ably, Pusher)
  • Sub-100ms fan-out required: API Gateway adds overhead. Consider AppSync subscriptions or self-managed MQTT
  • Peer-to-peer messaging dominates: WebRTC or a message broker pattern may be more efficient

Conclusion

API Gateway WebSocket APIs scale to 100K+ concurrent connections with careful architecture — but the naive approach (single Lambda fan-out, no batching, no topic segmentation) hits cost and latency walls quickly.

The keys to successful scaling: DynamoDB-backed connection stores with TTL cleanup, parallel fan-out via Step Functions or batched Lambda invocations, message batching and delta compression to control costs, and strict connection authentication at the $connect route.

At $5,800/month for 100K connections with sub-5s fan-out, the total cost of ownership is competitive with self-hosted alternatives when you factor in the operational burden of managing WebSocket infrastructure yourself. The managed model wins on operational simplicity; self-hosted wins on per-message latency and cost at extreme scale.

Comments

    No comments yet. Be the first to share your thoughts.