Circuit Breaker and Bulkhead Patterns: Preventing Cascade Failures in Microservices

Production implementation of circuit breaker and bulkhead patterns that prevented 23 potential cascade failures in 12 months across a 180-service architecture.

#resilience#circuit-breaker#bulkhead#microservices
Cover image for the article: Circuit Breaker and Bulkhead Patterns: Preventing Cascade Failures in Microservices

A single slow database query in a payment service should not bring down your recommendation engine. Yet in a naive microservices architecture, that is exactly what happens. Thread pools exhaust, connection pools saturate, back-pressure propagates upstream, and within 90 seconds your entire platform is unresponsive. We know because it happened to us three times in six months before we implemented systematic resilience patterns.

This article details how we deployed circuit breakers and bulkhead isolation across 180 microservices, preventing 23 cascade failures in the subsequent 12 months. The patterns are well-known. The production implementation — tuning thresholds, handling partial failures gracefully, and maintaining observability — is where teams struggle.

The Cascade Failure Anatomy

Before implementing resilience patterns, our typical cascade failure followed this sequence:

  1. Root cause: Database connection pool saturation in Service A (payment processing)
  2. T+5s: Service A response times increase from 12ms to 4,200ms
  3. T+15s: Services B, C, D (which call A) exhaust their HTTP connection pools waiting for responses
  4. T+30s: Services B, C, D start timing out on their own incoming requests
  5. T+45s: Upstream services E through M begin failing
  6. T+90s: Platform-wide degradation. 94% of API requests failing.
  7. T+12min: Manual intervention restores service after database connection pool is recycled.

Total customer impact: 12 minutes of near-total outage affecting all services, caused by a single service's database issue.

Cascade Failure Propagation

Circuit Breaker Implementation

We implemented a three-state circuit breaker (closed, open, half-open) with adaptive thresholds that adjust based on historical error rates.

State Machine

StateBehaviorTransition Condition
ClosedAll requests pass throughError rate exceeds threshold → Open
OpenAll requests immediately fail with fallbackTimer expires → Half-Open
Half-OpenLimited probe requests allowedProbes succeed → Closed; Probes fail → Open

Core Implementation

// Production circuit breaker with adaptive thresholds
import { EventEmitter } from 'events';

interface CircuitBreakerConfig {
  name: string;
  failureThreshold: number;       // Error rate % to trip (default: 50)
  successThreshold: number;       // Successes in half-open to close (default: 5)
  timeout: number;                // Time in open state before half-open (ms)
  rollingWindowSize: number;      // Window for calculating error rate (ms)
  minimumRequests: number;        // Min requests before evaluating threshold
  halfOpenMaxConcurrency: number; // Max concurrent requests in half-open
  adaptiveEnabled: boolean;       // Enable adaptive threshold adjustment
}

enum CircuitState {
  CLOSED = 'closed',
  OPEN = 'open',
  HALF_OPEN = 'half_open'
}

interface CircuitMetrics {
  totalRequests: number;
  failures: number;
  successes: number;
  timeouts: number;
  rejections: number;
  latencyP50: number;
  latencyP99: number;
}

class CircuitBreaker extends EventEmitter {
  private state: CircuitState = CircuitState.CLOSED;
  private failureCount: number = 0;
  private successCount: number = 0;
  private halfOpenRequests: number = 0;
  private lastFailureTime: number = 0;
  private rollingWindow: { timestamp: number; success: boolean; latency: number }[] = [];
  private adaptiveThreshold: number;

  constructor(private config: CircuitBreakerConfig) {
    super();
    this.adaptiveThreshold = config.failureThreshold;
  }

  async execute<T>(
    operation: () => Promise<T>,
    fallback?: () => Promise<T>
  ): Promise<T> {
    if (this.state === CircuitState.OPEN) {
      if (this.shouldAttemptReset()) {
        this.transitionTo(CircuitState.HALF_OPEN);
      } else {
        this.recordRejection();
        if (fallback) return fallback();
        throw new CircuitOpenError(this.config.name, this.getTimeUntilRetry());
      }
    }

    if (this.state === CircuitState.HALF_OPEN) {
      if (this.halfOpenRequests >= this.config.halfOpenMaxConcurrency) {
        if (fallback) return fallback();
        throw new CircuitOpenError(this.config.name, 0);
      }
      this.halfOpenRequests++;
    }

    const startTime = Date.now();
    try {
      const result = await operation();
      this.recordSuccess(Date.now() - startTime);
      return result;
    } catch (error) {
      this.recordFailure(Date.now() - startTime);
      if (fallback && this.state === CircuitState.OPEN) {
        return fallback();
      }
      throw error;
    }
  }

  private recordSuccess(latency: number): void {
    this.rollingWindow.push({ timestamp: Date.now(), success: true, latency });
    this.pruneWindow();

    if (this.state === CircuitState.HALF_OPEN) {
      this.successCount++;
      this.halfOpenRequests--;
      if (this.successCount >= this.config.successThreshold) {
        this.transitionTo(CircuitState.CLOSED);
      }
    }
  }

  private recordFailure(latency: number): void {
    this.rollingWindow.push({ timestamp: Date.now(), success: false, latency });
    this.pruneWindow();
    this.lastFailureTime = Date.now();

    if (this.state === CircuitState.HALF_OPEN) {
      this.halfOpenRequests--;
      this.transitionTo(CircuitState.OPEN);
      return;
    }

    if (this.state === CircuitState.CLOSED) {
      const errorRate = this.calculateErrorRate();
      const threshold = this.config.adaptiveEnabled 
        ? this.adaptiveThreshold 
        : this.config.failureThreshold;

      if (this.rollingWindow.length >= this.config.minimumRequests 
          && errorRate >= threshold) {
        this.transitionTo(CircuitState.OPEN);
      }
    }
  }

  private calculateErrorRate(): number {
    if (this.rollingWindow.length === 0) return 0;
    const failures = this.rollingWindow.filter(r => !r.success).length;
    return (failures / this.rollingWindow.length) * 100;
  }

  private transitionTo(newState: CircuitState): void {
    const oldState = this.state;
    this.state = newState;
    
    if (newState === CircuitState.CLOSED) {
      this.failureCount = 0;
      this.successCount = 0;
    } else if (newState === CircuitState.HALF_OPEN) {
      this.successCount = 0;
      this.halfOpenRequests = 0;
    }

    this.emit('stateChange', { from: oldState, to: newState, name: this.config.name });
  }
}

Bulkhead Pattern Implementation

Circuit breakers prevent calling a failed service. Bulkheads prevent a failed service from consuming shared resources. We implement bulkheads at two levels: thread pool isolation and connection pool isolation.

Semaphore-Based Bulkhead

// Bulkhead implementation with queue and timeout
class Bulkhead {
  private activeCalls: number = 0;
  private queue: Array<{
    resolve: (value: boolean) => void;
    timer: NodeJS.Timeout;
  }> = [];

  constructor(
    private readonly name: string,
    private readonly maxConcurrent: number,
    private readonly maxQueue: number,
    private readonly queueTimeout: number
  ) {}

  async execute<T>(operation: () => Promise<T>): Promise<T> {
    const permitted = await this.tryAcquire();
    if (!permitted) {
      throw new BulkheadFullError(
        this.name, 
        this.activeCalls, 
        this.queue.length
      );
    }

    try {
      this.activeCalls++;
      return await operation();
    } finally {
      this.activeCalls--;
      this.releaseNext();
    }
  }

  private async tryAcquire(): Promise<boolean> {
    if (this.activeCalls < this.maxConcurrent) {
      return true;
    }

    if (this.queue.length >= this.maxQueue) {
      return false;
    }

    return new Promise<boolean>((resolve) => {
      const timer = setTimeout(() => {
        const index = this.queue.findIndex(q => q.resolve === resolve);
        if (index !== -1) {
          this.queue.splice(index, 1);
          resolve(false);
        }
      }, this.queueTimeout);

      this.queue.push({ resolve, timer });
    });
  }

  private releaseNext(): void {
    if (this.queue.length > 0) {
      const next = this.queue.shift()!;
      clearTimeout(next.timer);
      next.resolve(true);
    }
  }

  getMetrics() {
    return {
      name: this.name,
      activeCalls: this.activeCalls,
      queueDepth: this.queue.length,
      maxConcurrent: this.maxConcurrent,
      saturationPct: (this.activeCalls / this.maxConcurrent) * 100
    };
  }
}

Bulkhead Sizing Strategy

Sizing bulkheads correctly is critical. Too small and you reject legitimate traffic; too large and you lose isolation benefits.

Service DependencyMax ConcurrentQueue SizeQueue TimeoutRationale
Payment processor50252000msHigh value, low volume, slow
User service200100500msHigh volume, fast responses
Recommendation engine3001000msNon-critical, no queuing
Notification service100501000msModerate volume, can queue
Analytics pipeline2003000msBackground, fully degradable

Our sizing formula:

maxConcurrent = (target_rps × p99_latency_seconds) × 1.5
maxQueue = maxConcurrent × 0.5 (for critical services)
maxQueue = 0 (for non-critical services — fail fast)

Combined Pattern: Resilience Pipeline

In production, circuit breakers and bulkheads work together. Each outbound dependency gets a resilience pipeline: timeout → bulkhead → circuit breaker → retry → fallback.

// Composing resilience patterns into a pipeline
class ResiliencePipeline<T> {
  private steps: Array<(fn: () => Promise<T>) => Promise<T>> = [];

  withTimeout(ms: number): this {
    this.steps.push((fn) => {
      return Promise.race([
        fn(),
        new Promise<T>((_, reject) => 
          setTimeout(() => reject(new TimeoutError(ms)), ms)
        )
      ]);
    });
    return this;
  }

  withBulkhead(bulkhead: Bulkhead): this {
    this.steps.push((fn) => bulkhead.execute(fn));
    return this;
  }

  withCircuitBreaker(cb: CircuitBreaker, fallback?: () => Promise<T>): this {
    this.steps.push((fn) => cb.execute(fn, fallback));
    return this;
  }

  withRetry(maxAttempts: number, backoffMs: number): this {
    this.steps.push(async (fn) => {
      let lastError: Error | undefined;
      for (let i = 0; i < maxAttempts; i++) {
        try {
          return await fn();
        } catch (error) {
          lastError = error as Error;
          if (i < maxAttempts - 1) {
            await new Promise(r => setTimeout(r, backoffMs * Math.pow(2, i)));
          }
        }
      }
      throw lastError;
    });
    return this;
  }

  async execute(operation: () => Promise<T>): Promise<T> {
    let current = operation;
    // Apply steps in reverse so outermost (timeout) wraps everything
    for (const step of [...this.steps].reverse()) {
      const prev = current;
      current = () => step(prev);
    }
    return current();
  }
}

// Usage
const paymentPipeline = new ResiliencePipeline<PaymentResult>()
  .withTimeout(5000)
  .withBulkhead(paymentBulkhead)
  .withCircuitBreaker(paymentCircuitBreaker, () => cachedPaymentFallback())
  .withRetry(3, 200);

const result = await paymentPipeline.execute(() => paymentService.process(order));

Results After 12 Months

MetricBefore Resilience PatternsAfterImprovement
Cascade failures (platform-wide)6 per quarter0100% reduction
Single-service failures contained12%96%8x improvement
Mean blast radius (services affected)34 services1.2 services28x reduction
Customer-facing impact duration8.4 min average0.3 min average96% reduction
P99 latency during incidents12,400ms320ms (degraded graceful)97% improvement

Operational Monitoring

Every circuit breaker and bulkhead emits metrics to our observability stack. We monitor:

  • Circuit breaker state transitions — Alert on any breaker opening. Page if more than 3 breakers open simultaneously.
  • Bulkhead saturation — Alert at 80% saturation, page at 95%.
  • Rejection rate — Track requests rejected by open circuits or full bulkheads.
  • Fallback invocation rate — High fallback rates indicate degraded user experience even if no errors surface.

Conclusion

Resilience patterns are not optional in a microservices architecture — they are the mechanism that converts a distributed system from a fragile chain into an antifragile network. The technical implementation is straightforward. The operational discipline — tuning thresholds, testing fallbacks, monitoring saturation — is what separates resilient systems from systems that merely have resilience libraries installed.

Start with your most critical dependency path. Instrument it fully. Test failure scenarios in staging weekly. Then expand outward. The patterns compound: each isolated failure domain makes the overall system more predictable, and predictable systems are systems you can sleep through the night operating.

Comments

    No comments yet. Be the first to share your thoughts.