EventBridge Event-Driven Architecture: Patterns That Handle 2M Events/Hour
Event routing patterns, schema governance, and throughput analysis from building an event-driven platform processing 2M events/hour across 23 microservices.

Event-driven architecture is easy to start and hard to operate at scale. After building an EventBridge-based platform that routes 2M events/hour across 23 microservices, I have learned that the routing patterns, schema governance, and dead-letter strategies matter far more than the event bus itself.
This is the architecture we built, the patterns that work, and the operational gotchas that cost us weekends.
Architecture Overview
Our platform processes order lifecycle events: creation, payment, fulfillment, shipping, delivery, and returns. Each event triggers downstream reactions across inventory, notifications, analytics, and fraud detection services.
- Event volume: 2.1M events/hour at peak (Friday evenings)
- Event bus: Custom EventBridge bus with 23 rules
- Sources: 8 producer services
- Targets: 15 consumer services
- Average end-to-end latency: 340ms (event published to consumer processed)
- Monthly EventBridge cost: $4,200
Pattern 1: Content-Based Routing with Rule Composition
EventBridge rules use content-based filtering. We structure events to maximize routing efficiency:
{
"source": "com.platform.orders",
"detail-type": "OrderStateChanged",
"detail": {
"metadata": {
"version": "2.1",
"correlationId": "uuid-123",
"timestamp": "2025-12-01T10:30:00Z",
"environment": "production"
},
"data": {
"orderId": "ORD-456789",
"customerId": "CUST-123",
"previousState": "PAYMENT_PENDING",
"currentState": "PAYMENT_CONFIRMED",
"orderValue": 15499,
"currency": "USD",
"fulfillmentCenter": "fc-east-1",
"isHighValue": true,
"channel": "web"
}
}
}
Routing rules target specific consumers based on event content:
{
"Name": "HighValueOrderToFraud",
"EventPattern": {
"source": ["com.platform.orders"],
"detail-type": ["OrderStateChanged"],
"detail": {
"data": {
"currentState": ["PAYMENT_CONFIRMED"],
"isHighValue": [true],
"orderValue": [{"numeric": [">", 10000]}]
}
}
},
"Targets": [{
"Arn": "arn:aws:sqs:us-east-1:123456789012:fraud-detection-queue",
"Id": "fraud-target",
"InputTransformer": {
"InputPathsMap": {
"orderId": "$.detail.data.orderId",
"amount": "$.detail.data.orderValue",
"customer": "$.detail.data.customerId"
},
"InputTemplate": "{\"orderId\": <orderId>, \"amount\": <amount>, \"customerId\": <customer>, \"checkType\": \"high-value\"}"
}
}]
}
Key insight: use InputTransformer to send consumers only the data they need. This reduces message size by 60-80% and decouples consumers from the full event schema.
Pattern 2: Schema Registry and Evolution
Schema governance prevents the #1 cause of event-driven failures: producers changing event shapes without consumer awareness.
// event-schema.ts — Strongly typed event definitions
import { z } from 'zod';
export const OrderStateChangedV2 = z.object({
source: z.literal('com.platform.orders'),
'detail-type': z.literal('OrderStateChanged'),
detail: z.object({
metadata: z.object({
version: z.string(),
correlationId: z.string().uuid(),
timestamp: z.string().datetime(),
environment: z.enum(['production', 'staging', 'development']),
}),
data: z.object({
orderId: z.string(),
customerId: z.string(),
previousState: z.enum([
'CREATED', 'PAYMENT_PENDING', 'PAYMENT_CONFIRMED',
'FULFILLING', 'SHIPPED', 'DELIVERED', 'RETURNED', 'CANCELLED'
]),
currentState: z.enum([
'CREATED', 'PAYMENT_PENDING', 'PAYMENT_CONFIRMED',
'FULFILLING', 'SHIPPED', 'DELIVERED', 'RETURNED', 'CANCELLED'
]),
orderValue: z.number().positive(),
currency: z.string().length(3),
fulfillmentCenter: z.string(),
isHighValue: z.boolean(),
channel: z.enum(['web', 'mobile', 'api', 'pos']),
}),
}),
});
// Publish with validation
export async function publishOrderEvent(
client: EventBridgeClient,
event: z.infer<typeof OrderStateChangedV2>
): Promise<void> {
// Validate before publishing — fail fast
OrderStateChangedV2.parse(event);
await client.send(new PutEventsCommand({
Entries: [{
Source: event.source,
DetailType: event['detail-type'],
Detail: JSON.stringify(event.detail),
EventBusName: 'platform-events',
}],
}));
}
We enforce schema validation at publish time — not consume time. A malformed event caught at the producer is 100x cheaper to fix than one that corrupts downstream state.
Pattern 3: Dead-Letter Queues and Replay
Every EventBridge rule targets an SQS dead-letter queue for failed deliveries:
# CDK Infrastructure
const orderEventRule = new events.Rule(this, 'OrderToInventory', {
eventBus: platformBus,
eventPattern: {
source: ['com.platform.orders'],
detailType: ['OrderStateChanged'],
detail: { data: { currentState: ['PAYMENT_CONFIRMED'] } },
},
});
const dlq = new sqs.Queue(this, 'InventoryDLQ', {
retentionPeriod: Duration.days(14),
visibilityTimeout: Duration.minutes(5),
});
orderEventRule.addTarget(new targets.LambdaFunction(inventoryHandler, {
deadLetterQueue: dlq,
retryAttempts: 2,
maxEventAge: Duration.hours(1),
}));
The replay mechanism for DLQ processing:
// dlq-replayer.ts
import { SQSClient, ReceiveMessageCommand, DeleteMessageCommand } from '@aws-sdk/client-sqs';
import { EventBridgeClient, PutEventsCommand } from '@aws-sdk/client-eventbridge';
async function replayDeadLetters(queueUrl: string, batchSize: number = 10): Promise<number> {
const sqs = new SQSClient({});
const eb = new EventBridgeClient({});
let replayed = 0;
while (true) {
const response = await sqs.send(new ReceiveMessageCommand({
QueueUrl: queueUrl,
MaxNumberOfMessages: batchSize,
WaitTimeSeconds: 5,
}));
if (!response.Messages || response.Messages.length === 0) break;
for (const message of response.Messages) {
const event = JSON.parse(message.Body!);
// Re-publish to event bus with replay metadata
await eb.send(new PutEventsCommand({
Entries: [{
Source: event.source,
DetailType: event['detail-type'],
Detail: JSON.stringify({
...event.detail,
metadata: {
...event.detail.metadata,
isReplay: true,
originalTimestamp: event.detail.metadata.timestamp,
replayTimestamp: new Date().toISOString(),
},
}),
EventBusName: 'platform-events',
}],
}));
await sqs.send(new DeleteMessageCommand({
QueueUrl: queueUrl,
ReceiptHandle: message.ReceiptHandle!,
}));
replayed++;
}
}
return replayed;
}
Pattern 4: Event Archive and Replay for Disaster Recovery
EventBridge Archive stores events for replay — critical for disaster recovery and building new consumers:
// Archive configuration
const archive = new events.Archive(this, 'PlatformArchive', {
sourceEventBus: platformBus,
eventPattern: {
source: [{ prefix: 'com.platform' }],
},
archiveName: 'platform-events-archive',
retention: Duration.days(365),
});
// Replay events for a new consumer (backfill)
async function replayForNewConsumer(
startTime: Date,
endTime: Date,
eventPattern: object
): Promise<string> {
const eb = new EventBridgeClient({});
const response = await eb.send(new StartReplayCommand({
ReplayName: `backfill-${Date.now()}`,
EventSourceArn: 'arn:aws:events:us-east-1:123456789012:archive/platform-events-archive',
Destination: {
Arn: 'arn:aws:events:us-east-1:123456789012:event-bus/platform-events',
FilterArns: [],
},
EventStartTime: startTime,
EventEndTime: endTime,
}));
return response.ReplayArn!;
}
We used archive replay when launching a new analytics consumer — replaying 90 days of events to populate dashboards without modifying any producer.
Throughput Analysis and Limits
EventBridge has soft limits that become real constraints at scale:
| Limit | Default | Our Setting | Impact at 2M events/hr |
|---|---|---|---|
| PutEvents TPS | 10,000 | 50,000 (raised) | Hit during flash sales |
| Rules per bus | 300 | 300 | Using 23/300 |
| Targets per rule | 5 | 5 | Necessitates fan-out patterns |
| Event size | 256KB | 256KB | Must trim payloads |
| Invocations per second (Lambda target) | 1,000 | 5,000 (raised) | Critical for burst handling |
The 5-target-per-rule limit is the most constraining. When 8 services need the same event, we use a fan-out pattern:
EventBridge Rule → SNS Topic → 8 SQS Queues → 8 Lambda Functions
This adds ~50ms latency but eliminates the target limit constraint.
Operational Metrics (12 Months)
| Metric | Value |
|---|---|
| Total events processed | 14.2 billion |
| Delivery success rate | 99.97% |
| Events to DLQ | 42,000 (0.0003%) |
| Average delivery latency | 68ms |
| P99 delivery latency | 340ms |
| Schema validation failures | 1,247 (caught at publish) |
| Archive replay operations | 8 (new consumer backfills) |
| Monthly cost breakdown | Events: $2,800 / Archive: $940 / Schema Registry: $460 |
Key Takeaways
- Schema governance prevents cascading failures: Validate at publish time. A producer breaking schema should fail loudly, not silently corrupt consumers.
- Every rule needs a DLQ: EventBridge delivery failures are silent by default. DLQs make them visible and replayable.
- InputTransformer decouples consumers from producers: Send consumers only what they need. This makes schema evolution manageable.
- Archive everything: The ability to replay 90 days of events for a new consumer without touching producers is transformational for team velocity.
- Request limit increases proactively: Default PutEvents TPS (10,000) is surprisingly easy to hit during traffic spikes. Monitor and raise limits before production incidents.
Event-driven architecture with EventBridge is not about the bus — it is about the governance, observability, and failure handling you build around it.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.