EventBridge Event-Driven Architecture: Patterns That Handle 2M Events/Hour

Event routing patterns, schema governance, and throughput analysis from building an event-driven platform processing 2M events/hour across 23 microservices.

#aws#eventbridge#event-driven#architecture
Cover image for the article: EventBridge Event-Driven Architecture: Patterns That Handle 2M Events/Hour

Event-driven architecture is easy to start and hard to operate at scale. After building an EventBridge-based platform that routes 2M events/hour across 23 microservices, I have learned that the routing patterns, schema governance, and dead-letter strategies matter far more than the event bus itself.

This is the architecture we built, the patterns that work, and the operational gotchas that cost us weekends.

Architecture Overview

Our platform processes order lifecycle events: creation, payment, fulfillment, shipping, delivery, and returns. Each event triggers downstream reactions across inventory, notifications, analytics, and fraud detection services.

  • Event volume: 2.1M events/hour at peak (Friday evenings)
  • Event bus: Custom EventBridge bus with 23 rules
  • Sources: 8 producer services
  • Targets: 15 consumer services
  • Average end-to-end latency: 340ms (event published to consumer processed)
  • Monthly EventBridge cost: $4,200

EventBridge Architecture Overview

Pattern 1: Content-Based Routing with Rule Composition

EventBridge rules use content-based filtering. We structure events to maximize routing efficiency:

{
  "source": "com.platform.orders",
  "detail-type": "OrderStateChanged",
  "detail": {
    "metadata": {
      "version": "2.1",
      "correlationId": "uuid-123",
      "timestamp": "2025-12-01T10:30:00Z",
      "environment": "production"
    },
    "data": {
      "orderId": "ORD-456789",
      "customerId": "CUST-123",
      "previousState": "PAYMENT_PENDING",
      "currentState": "PAYMENT_CONFIRMED",
      "orderValue": 15499,
      "currency": "USD",
      "fulfillmentCenter": "fc-east-1",
      "isHighValue": true,
      "channel": "web"
    }
  }
}

Routing rules target specific consumers based on event content:

{
  "Name": "HighValueOrderToFraud",
  "EventPattern": {
    "source": ["com.platform.orders"],
    "detail-type": ["OrderStateChanged"],
    "detail": {
      "data": {
        "currentState": ["PAYMENT_CONFIRMED"],
        "isHighValue": [true],
        "orderValue": [{"numeric": [">", 10000]}]
      }
    }
  },
  "Targets": [{
    "Arn": "arn:aws:sqs:us-east-1:123456789012:fraud-detection-queue",
    "Id": "fraud-target",
    "InputTransformer": {
      "InputPathsMap": {
        "orderId": "$.detail.data.orderId",
        "amount": "$.detail.data.orderValue",
        "customer": "$.detail.data.customerId"
      },
      "InputTemplate": "{\"orderId\": <orderId>, \"amount\": <amount>, \"customerId\": <customer>, \"checkType\": \"high-value\"}"
    }
  }]
}

Key insight: use InputTransformer to send consumers only the data they need. This reduces message size by 60-80% and decouples consumers from the full event schema.

Pattern 2: Schema Registry and Evolution

Schema governance prevents the #1 cause of event-driven failures: producers changing event shapes without consumer awareness.

// event-schema.ts — Strongly typed event definitions
import { z } from 'zod';

export const OrderStateChangedV2 = z.object({
  source: z.literal('com.platform.orders'),
  'detail-type': z.literal('OrderStateChanged'),
  detail: z.object({
    metadata: z.object({
      version: z.string(),
      correlationId: z.string().uuid(),
      timestamp: z.string().datetime(),
      environment: z.enum(['production', 'staging', 'development']),
    }),
    data: z.object({
      orderId: z.string(),
      customerId: z.string(),
      previousState: z.enum([
        'CREATED', 'PAYMENT_PENDING', 'PAYMENT_CONFIRMED',
        'FULFILLING', 'SHIPPED', 'DELIVERED', 'RETURNED', 'CANCELLED'
      ]),
      currentState: z.enum([
        'CREATED', 'PAYMENT_PENDING', 'PAYMENT_CONFIRMED',
        'FULFILLING', 'SHIPPED', 'DELIVERED', 'RETURNED', 'CANCELLED'
      ]),
      orderValue: z.number().positive(),
      currency: z.string().length(3),
      fulfillmentCenter: z.string(),
      isHighValue: z.boolean(),
      channel: z.enum(['web', 'mobile', 'api', 'pos']),
    }),
  }),
});

// Publish with validation
export async function publishOrderEvent(
  client: EventBridgeClient,
  event: z.infer<typeof OrderStateChangedV2>
): Promise<void> {
  // Validate before publishing — fail fast
  OrderStateChangedV2.parse(event);

  await client.send(new PutEventsCommand({
    Entries: [{
      Source: event.source,
      DetailType: event['detail-type'],
      Detail: JSON.stringify(event.detail),
      EventBusName: 'platform-events',
    }],
  }));
}

We enforce schema validation at publish time — not consume time. A malformed event caught at the producer is 100x cheaper to fix than one that corrupts downstream state.

Pattern 3: Dead-Letter Queues and Replay

Every EventBridge rule targets an SQS dead-letter queue for failed deliveries:

# CDK Infrastructure
const orderEventRule = new events.Rule(this, 'OrderToInventory', {
  eventBus: platformBus,
  eventPattern: {
    source: ['com.platform.orders'],
    detailType: ['OrderStateChanged'],
    detail: { data: { currentState: ['PAYMENT_CONFIRMED'] } },
  },
});

const dlq = new sqs.Queue(this, 'InventoryDLQ', {
  retentionPeriod: Duration.days(14),
  visibilityTimeout: Duration.minutes(5),
});

orderEventRule.addTarget(new targets.LambdaFunction(inventoryHandler, {
  deadLetterQueue: dlq,
  retryAttempts: 2,
  maxEventAge: Duration.hours(1),
}));

The replay mechanism for DLQ processing:

// dlq-replayer.ts
import { SQSClient, ReceiveMessageCommand, DeleteMessageCommand } from '@aws-sdk/client-sqs';
import { EventBridgeClient, PutEventsCommand } from '@aws-sdk/client-eventbridge';

async function replayDeadLetters(queueUrl: string, batchSize: number = 10): Promise<number> {
  const sqs = new SQSClient({});
  const eb = new EventBridgeClient({});
  let replayed = 0;

  while (true) {
    const response = await sqs.send(new ReceiveMessageCommand({
      QueueUrl: queueUrl,
      MaxNumberOfMessages: batchSize,
      WaitTimeSeconds: 5,
    }));

    if (!response.Messages || response.Messages.length === 0) break;

    for (const message of response.Messages) {
      const event = JSON.parse(message.Body!);

      // Re-publish to event bus with replay metadata
      await eb.send(new PutEventsCommand({
        Entries: [{
          Source: event.source,
          DetailType: event['detail-type'],
          Detail: JSON.stringify({
            ...event.detail,
            metadata: {
              ...event.detail.metadata,
              isReplay: true,
              originalTimestamp: event.detail.metadata.timestamp,
              replayTimestamp: new Date().toISOString(),
            },
          }),
          EventBusName: 'platform-events',
        }],
      }));

      await sqs.send(new DeleteMessageCommand({
        QueueUrl: queueUrl,
        ReceiptHandle: message.ReceiptHandle!,
      }));

      replayed++;
    }
  }

  return replayed;
}

Pattern 4: Event Archive and Replay for Disaster Recovery

EventBridge Archive stores events for replay — critical for disaster recovery and building new consumers:

// Archive configuration
const archive = new events.Archive(this, 'PlatformArchive', {
  sourceEventBus: platformBus,
  eventPattern: {
    source: [{ prefix: 'com.platform' }],
  },
  archiveName: 'platform-events-archive',
  retention: Duration.days(365),
});

// Replay events for a new consumer (backfill)
async function replayForNewConsumer(
  startTime: Date,
  endTime: Date,
  eventPattern: object
): Promise<string> {
  const eb = new EventBridgeClient({});

  const response = await eb.send(new StartReplayCommand({
    ReplayName: `backfill-${Date.now()}`,
    EventSourceArn: 'arn:aws:events:us-east-1:123456789012:archive/platform-events-archive',
    Destination: {
      Arn: 'arn:aws:events:us-east-1:123456789012:event-bus/platform-events',
      FilterArns: [],
    },
    EventStartTime: startTime,
    EventEndTime: endTime,
  }));

  return response.ReplayArn!;
}

We used archive replay when launching a new analytics consumer — replaying 90 days of events to populate dashboards without modifying any producer.

Throughput Analysis and Limits

EventBridge has soft limits that become real constraints at scale:

LimitDefaultOur SettingImpact at 2M events/hr
PutEvents TPS10,00050,000 (raised)Hit during flash sales
Rules per bus300300Using 23/300
Targets per rule55Necessitates fan-out patterns
Event size256KB256KBMust trim payloads
Invocations per second (Lambda target)1,0005,000 (raised)Critical for burst handling

The 5-target-per-rule limit is the most constraining. When 8 services need the same event, we use a fan-out pattern:

EventBridge Rule → SNS Topic → 8 SQS Queues → 8 Lambda Functions

This adds ~50ms latency but eliminates the target limit constraint.

EventBridge Throughput Under Load

Operational Metrics (12 Months)

MetricValue
Total events processed14.2 billion
Delivery success rate99.97%
Events to DLQ42,000 (0.0003%)
Average delivery latency68ms
P99 delivery latency340ms
Schema validation failures1,247 (caught at publish)
Archive replay operations8 (new consumer backfills)
Monthly cost breakdownEvents: $2,800 / Archive: $940 / Schema Registry: $460

Key Takeaways

  1. Schema governance prevents cascading failures: Validate at publish time. A producer breaking schema should fail loudly, not silently corrupt consumers.
  2. Every rule needs a DLQ: EventBridge delivery failures are silent by default. DLQs make them visible and replayable.
  3. InputTransformer decouples consumers from producers: Send consumers only what they need. This makes schema evolution manageable.
  4. Archive everything: The ability to replay 90 days of events for a new consumer without touching producers is transformational for team velocity.
  5. Request limit increases proactively: Default PutEvents TPS (10,000) is surprisingly easy to hit during traffic spikes. Monitor and raise limits before production incidents.

Event-driven architecture with EventBridge is not about the bus — it is about the governance, observability, and failure handling you build around it.

Comments

    No comments yet. Be the first to share your thoughts.