Circuit Breaker and Bulkhead Patterns: Preventing Cascade Failures in Microservices
Production implementation of circuit breaker and bulkhead patterns that prevented 23 potential cascade failures in 12 months across a 180-service architecture.

A single slow database query in a payment service should not bring down your recommendation engine. Yet in a naive microservices architecture, that is exactly what happens. Thread pools exhaust, connection pools saturate, back-pressure propagates upstream, and within 90 seconds your entire platform is unresponsive. We know because it happened to us three times in six months before we implemented systematic resilience patterns.
This article details how we deployed circuit breakers and bulkhead isolation across 180 microservices, preventing 23 cascade failures in the subsequent 12 months. The patterns are well-known. The production implementation — tuning thresholds, handling partial failures gracefully, and maintaining observability — is where teams struggle.
The Cascade Failure Anatomy
Before implementing resilience patterns, our typical cascade failure followed this sequence:
- Root cause: Database connection pool saturation in Service A (payment processing)
- T+5s: Service A response times increase from 12ms to 4,200ms
- T+15s: Services B, C, D (which call A) exhaust their HTTP connection pools waiting for responses
- T+30s: Services B, C, D start timing out on their own incoming requests
- T+45s: Upstream services E through M begin failing
- T+90s: Platform-wide degradation. 94% of API requests failing.
- T+12min: Manual intervention restores service after database connection pool is recycled.
Total customer impact: 12 minutes of near-total outage affecting all services, caused by a single service's database issue.
Circuit Breaker Implementation
We implemented a three-state circuit breaker (closed, open, half-open) with adaptive thresholds that adjust based on historical error rates.
State Machine
| State | Behavior | Transition Condition |
|---|---|---|
| Closed | All requests pass through | Error rate exceeds threshold → Open |
| Open | All requests immediately fail with fallback | Timer expires → Half-Open |
| Half-Open | Limited probe requests allowed | Probes succeed → Closed; Probes fail → Open |
Core Implementation
// Production circuit breaker with adaptive thresholds
import { EventEmitter } from 'events';
interface CircuitBreakerConfig {
name: string;
failureThreshold: number; // Error rate % to trip (default: 50)
successThreshold: number; // Successes in half-open to close (default: 5)
timeout: number; // Time in open state before half-open (ms)
rollingWindowSize: number; // Window for calculating error rate (ms)
minimumRequests: number; // Min requests before evaluating threshold
halfOpenMaxConcurrency: number; // Max concurrent requests in half-open
adaptiveEnabled: boolean; // Enable adaptive threshold adjustment
}
enum CircuitState {
CLOSED = 'closed',
OPEN = 'open',
HALF_OPEN = 'half_open'
}
interface CircuitMetrics {
totalRequests: number;
failures: number;
successes: number;
timeouts: number;
rejections: number;
latencyP50: number;
latencyP99: number;
}
class CircuitBreaker extends EventEmitter {
private state: CircuitState = CircuitState.CLOSED;
private failureCount: number = 0;
private successCount: number = 0;
private halfOpenRequests: number = 0;
private lastFailureTime: number = 0;
private rollingWindow: { timestamp: number; success: boolean; latency: number }[] = [];
private adaptiveThreshold: number;
constructor(private config: CircuitBreakerConfig) {
super();
this.adaptiveThreshold = config.failureThreshold;
}
async execute<T>(
operation: () => Promise<T>,
fallback?: () => Promise<T>
): Promise<T> {
if (this.state === CircuitState.OPEN) {
if (this.shouldAttemptReset()) {
this.transitionTo(CircuitState.HALF_OPEN);
} else {
this.recordRejection();
if (fallback) return fallback();
throw new CircuitOpenError(this.config.name, this.getTimeUntilRetry());
}
}
if (this.state === CircuitState.HALF_OPEN) {
if (this.halfOpenRequests >= this.config.halfOpenMaxConcurrency) {
if (fallback) return fallback();
throw new CircuitOpenError(this.config.name, 0);
}
this.halfOpenRequests++;
}
const startTime = Date.now();
try {
const result = await operation();
this.recordSuccess(Date.now() - startTime);
return result;
} catch (error) {
this.recordFailure(Date.now() - startTime);
if (fallback && this.state === CircuitState.OPEN) {
return fallback();
}
throw error;
}
}
private recordSuccess(latency: number): void {
this.rollingWindow.push({ timestamp: Date.now(), success: true, latency });
this.pruneWindow();
if (this.state === CircuitState.HALF_OPEN) {
this.successCount++;
this.halfOpenRequests--;
if (this.successCount >= this.config.successThreshold) {
this.transitionTo(CircuitState.CLOSED);
}
}
}
private recordFailure(latency: number): void {
this.rollingWindow.push({ timestamp: Date.now(), success: false, latency });
this.pruneWindow();
this.lastFailureTime = Date.now();
if (this.state === CircuitState.HALF_OPEN) {
this.halfOpenRequests--;
this.transitionTo(CircuitState.OPEN);
return;
}
if (this.state === CircuitState.CLOSED) {
const errorRate = this.calculateErrorRate();
const threshold = this.config.adaptiveEnabled
? this.adaptiveThreshold
: this.config.failureThreshold;
if (this.rollingWindow.length >= this.config.minimumRequests
&& errorRate >= threshold) {
this.transitionTo(CircuitState.OPEN);
}
}
}
private calculateErrorRate(): number {
if (this.rollingWindow.length === 0) return 0;
const failures = this.rollingWindow.filter(r => !r.success).length;
return (failures / this.rollingWindow.length) * 100;
}
private transitionTo(newState: CircuitState): void {
const oldState = this.state;
this.state = newState;
if (newState === CircuitState.CLOSED) {
this.failureCount = 0;
this.successCount = 0;
} else if (newState === CircuitState.HALF_OPEN) {
this.successCount = 0;
this.halfOpenRequests = 0;
}
this.emit('stateChange', { from: oldState, to: newState, name: this.config.name });
}
}
Bulkhead Pattern Implementation
Circuit breakers prevent calling a failed service. Bulkheads prevent a failed service from consuming shared resources. We implement bulkheads at two levels: thread pool isolation and connection pool isolation.
Semaphore-Based Bulkhead
// Bulkhead implementation with queue and timeout
class Bulkhead {
private activeCalls: number = 0;
private queue: Array<{
resolve: (value: boolean) => void;
timer: NodeJS.Timeout;
}> = [];
constructor(
private readonly name: string,
private readonly maxConcurrent: number,
private readonly maxQueue: number,
private readonly queueTimeout: number
) {}
async execute<T>(operation: () => Promise<T>): Promise<T> {
const permitted = await this.tryAcquire();
if (!permitted) {
throw new BulkheadFullError(
this.name,
this.activeCalls,
this.queue.length
);
}
try {
this.activeCalls++;
return await operation();
} finally {
this.activeCalls--;
this.releaseNext();
}
}
private async tryAcquire(): Promise<boolean> {
if (this.activeCalls < this.maxConcurrent) {
return true;
}
if (this.queue.length >= this.maxQueue) {
return false;
}
return new Promise<boolean>((resolve) => {
const timer = setTimeout(() => {
const index = this.queue.findIndex(q => q.resolve === resolve);
if (index !== -1) {
this.queue.splice(index, 1);
resolve(false);
}
}, this.queueTimeout);
this.queue.push({ resolve, timer });
});
}
private releaseNext(): void {
if (this.queue.length > 0) {
const next = this.queue.shift()!;
clearTimeout(next.timer);
next.resolve(true);
}
}
getMetrics() {
return {
name: this.name,
activeCalls: this.activeCalls,
queueDepth: this.queue.length,
maxConcurrent: this.maxConcurrent,
saturationPct: (this.activeCalls / this.maxConcurrent) * 100
};
}
}
Bulkhead Sizing Strategy
Sizing bulkheads correctly is critical. Too small and you reject legitimate traffic; too large and you lose isolation benefits.
| Service Dependency | Max Concurrent | Queue Size | Queue Timeout | Rationale |
|---|---|---|---|---|
| Payment processor | 50 | 25 | 2000ms | High value, low volume, slow |
| User service | 200 | 100 | 500ms | High volume, fast responses |
| Recommendation engine | 30 | 0 | 1000ms | Non-critical, no queuing |
| Notification service | 100 | 50 | 1000ms | Moderate volume, can queue |
| Analytics pipeline | 20 | 0 | 3000ms | Background, fully degradable |
Our sizing formula:
maxConcurrent = (target_rps × p99_latency_seconds) × 1.5
maxQueue = maxConcurrent × 0.5 (for critical services)
maxQueue = 0 (for non-critical services — fail fast)
Combined Pattern: Resilience Pipeline
In production, circuit breakers and bulkheads work together. Each outbound dependency gets a resilience pipeline: timeout → bulkhead → circuit breaker → retry → fallback.
// Composing resilience patterns into a pipeline
class ResiliencePipeline<T> {
private steps: Array<(fn: () => Promise<T>) => Promise<T>> = [];
withTimeout(ms: number): this {
this.steps.push((fn) => {
return Promise.race([
fn(),
new Promise<T>((_, reject) =>
setTimeout(() => reject(new TimeoutError(ms)), ms)
)
]);
});
return this;
}
withBulkhead(bulkhead: Bulkhead): this {
this.steps.push((fn) => bulkhead.execute(fn));
return this;
}
withCircuitBreaker(cb: CircuitBreaker, fallback?: () => Promise<T>): this {
this.steps.push((fn) => cb.execute(fn, fallback));
return this;
}
withRetry(maxAttempts: number, backoffMs: number): this {
this.steps.push(async (fn) => {
let lastError: Error | undefined;
for (let i = 0; i < maxAttempts; i++) {
try {
return await fn();
} catch (error) {
lastError = error as Error;
if (i < maxAttempts - 1) {
await new Promise(r => setTimeout(r, backoffMs * Math.pow(2, i)));
}
}
}
throw lastError;
});
return this;
}
async execute(operation: () => Promise<T>): Promise<T> {
let current = operation;
// Apply steps in reverse so outermost (timeout) wraps everything
for (const step of [...this.steps].reverse()) {
const prev = current;
current = () => step(prev);
}
return current();
}
}
// Usage
const paymentPipeline = new ResiliencePipeline<PaymentResult>()
.withTimeout(5000)
.withBulkhead(paymentBulkhead)
.withCircuitBreaker(paymentCircuitBreaker, () => cachedPaymentFallback())
.withRetry(3, 200);
const result = await paymentPipeline.execute(() => paymentService.process(order));
Results After 12 Months
| Metric | Before Resilience Patterns | After | Improvement |
|---|---|---|---|
| Cascade failures (platform-wide) | 6 per quarter | 0 | 100% reduction |
| Single-service failures contained | 12% | 96% | 8x improvement |
| Mean blast radius (services affected) | 34 services | 1.2 services | 28x reduction |
| Customer-facing impact duration | 8.4 min average | 0.3 min average | 96% reduction |
| P99 latency during incidents | 12,400ms | 320ms (degraded graceful) | 97% improvement |
Operational Monitoring
Every circuit breaker and bulkhead emits metrics to our observability stack. We monitor:
- Circuit breaker state transitions — Alert on any breaker opening. Page if more than 3 breakers open simultaneously.
- Bulkhead saturation — Alert at 80% saturation, page at 95%.
- Rejection rate — Track requests rejected by open circuits or full bulkheads.
- Fallback invocation rate — High fallback rates indicate degraded user experience even if no errors surface.
Conclusion
Resilience patterns are not optional in a microservices architecture — they are the mechanism that converts a distributed system from a fragile chain into an antifragile network. The technical implementation is straightforward. The operational discipline — tuning thresholds, testing fallbacks, monitoring saturation — is what separates resilient systems from systems that merely have resilience libraries installed.
Start with your most critical dependency path. Instrument it fully. Test failure scenarios in staging weekly. Then expand outward. The patterns compound: each isolated failure domain makes the overall system more predictable, and predictable systems are systems you can sleep through the night operating.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.