Chaos Engineering That Found 12 Critical Failure Modes We Never Anticipated

How systematic chaos experiments in production uncovered hidden dependencies, cascade failures, and timeout misconfigurations across our distributed system

#chaos-engineering#resilience#sre#testing
Cover image for the article: Chaos Engineering That Found 12 Critical Failure Modes We Never Anticipated

We had 99.95% uptime, comprehensive monitoring, and circuit breakers on every external call. We thought our system was resilient. Then we started running chaos experiments in production and discovered 12 critical failure modes that monitoring never caught because they required specific combinations of failures that normal operation never triggered. One of those hidden modes, if triggered organically, would have caused a 4-hour outage affecting all customers. Here is what we found and how we found it.

The Problem: Untested Failure Assumptions

Resilience engineering is full of assumptions. "The circuit breaker will open before cascade failure." "The retry budget will prevent thundering herds." "The fallback cache will serve stale data gracefully." These assumptions are rarely tested under realistic conditions. Unit tests mock the failure. Integration tests test one failure at a time. Production has correlated failures.

Our pre-chaos confidence was based on:

  • Circuit breakers configured on 47 external dependencies
  • Retry logic with exponential backoff on all HTTP clients
  • Health checks on every service with automated restart
  • Multi-AZ deployment across three availability zones
  • 99.95% uptime over the previous 12 months

All of this gave us confidence. Chaos engineering gave us truth.

Chaos Architecture

We built our chaos platform on AWS Fault Injection Service (FIS) combined with custom chaos controllers for application-level faults:

Chaos Engineering Platform Architecture

The platform operates at three levels:

  1. Infrastructure chaos: AZ failures, network partitions, instance termination
  2. Application chaos: Latency injection, error responses, connection pool exhaustion
  3. Dependency chaos: Third-party API failures, database failover, cache eviction

Experiment Framework

Every chaos experiment follows a structured hypothesis-experiment-analysis cycle:

from dataclasses import dataclass, field
from datetime import datetime
from enum import Enum
from typing import Callable

class ExperimentState(Enum):
    DRAFT = "draft"
    APPROVED = "approved"
    RUNNING = "running"
    ANALYZING = "analyzing"
    COMPLETE = "complete"
    ABORTED = "aborted"

@dataclass
class ChaosExperiment:
    """Structured chaos experiment definition."""
    id: str
    name: str
    hypothesis: str
    steady_state: dict  # Metrics that define "normal"
    blast_radius: str   # What could be affected
    abort_conditions: list[str]  # When to stop immediately
    fault_injection: dict  # What we're breaking
    duration_minutes: int
    owner: str
    approved_by: str = ""
    state: ExperimentState = ExperimentState.DRAFT
    findings: list[str] = field(default_factory=list)

    def validate_steady_state(self) -> bool:
        """Verify system is in normal state before injecting faults."""
        for metric, threshold in self.steady_state.items():
            current_value = get_metric_value(metric)
            if not meets_threshold(current_value, threshold):
                print(f"ABORT: Steady state violation - {metric}={current_value}")
                return False
        return True

    def check_abort_conditions(self) -> bool:
        """Continuously check if experiment should be stopped."""
        for condition in self.abort_conditions:
            if evaluate_condition(condition):
                self.state = ExperimentState.ABORTED
                return True
        return False

# Example experiment definition
experiment_001 = ChaosExperiment(
    id="chaos-001",
    name="Payment service AZ failure",
    hypothesis="If us-east-1a becomes unavailable, payment processing "
               "continues via us-east-1b/c with <500ms additional latency",
    steady_state={
        "payment_success_rate": ">= 99.9%",
        "payment_p99_latency": "<= 800ms",
        "error_rate_5xx": "<= 0.1%",
    },
    blast_radius="Payment processing in AZ us-east-1a (33% of traffic)",
    abort_conditions=[
        "payment_success_rate < 99.0%",
        "payment_p99_latency > 5000ms",
        "error_rate_5xx > 1.0%",
    ],
    fault_injection={
        "type": "az_failure",
        "target_az": "us-east-1a",
        "services": ["payment-service"],
        "method": "network_partition",
    },
    duration_minutes=15,
    owner="@sre-team",
    approved_by="@vp-engineering",
)

AWS FIS Experiment Template

For infrastructure-level chaos, we use FIS experiment templates:

{
  "description": "Simulate AZ failure for payment service",
  "targets": {
    "payment-instances": {
      "resourceType": "aws:ec2:instance",
      "selectionMode": "ALL",
      "resourceTags": {
        "service": "payment-service",
        "az": "us-east-1a"
      },
      "filters": [
        {
          "path": "State.Name",
          "values": ["running"]
        }
      ]
    },
    "payment-subnets": {
      "resourceType": "aws:ec2:subnet",
      "selectionMode": "ALL",
      "resourceTags": {
        "service": "payment-service",
        "az": "us-east-1a"
      }
    }
  },
  "actions": {
    "inject-network-partition": {
      "actionId": "aws:network:disrupt-connectivity",
      "parameters": {
        "scope": "all",
        "duration": "PT15M"
      },
      "targets": {
        "Subnets": "payment-subnets"
      }
    },
    "stop-instances": {
      "actionId": "aws:ec2:stop-instances",
      "parameters": {
        "startInstancesAfterDuration": "PT15M"
      },
      "targets": {
        "Instances": "payment-instances"
      },
      "startAfter": ["inject-network-partition"]
    }
  },
  "stopConditions": [
    {
      "source": "aws:cloudwatch:alarm",
      "value": "arn:aws:cloudwatch:us-east-1:123456789:alarm:chaos-abort-payment"
    }
  ],
  "roleArn": "arn:aws:iam::123456789:role/fis-chaos-role"
}

The 12 Critical Findings

Over six months of systematic experimentation, we discovered these failure modes:

#FindingSeverityImpact if Triggered
1DNS cache poisoning during AZ failoverCritical4-hour outage
2Circuit breaker timeout mismatch cascading to all servicesCritical90-minute degradation
3Connection pool leak under partial network partitionHighService unavailability in 45min
4Retry amplification between three services creating feedback loopCriticalTraffic 340x amplification
5Health check passing while service unable to process requestsHighSilent failure for 12 minutes
6Cache stampede after Redis cluster failoverHighDatabase overload for 8 minutes
7Async message ordering violation during consumer restartMediumData inconsistency
8TLS certificate rotation failing under CPU pressureHighmTLS failure after 12 hours
9Graceful shutdown not draining in-flight requestsMedium0.3% request failure during deploys
10Database connection exhaustion from leaked prepared statementsHighProgressive degradation over 6 hours
11Cross-region failover DNS TTL exceeding client cacheCritical30-minute unavailability
12Rate limiter shared state lost during pod reschedulingMediumTemporary rate limit bypass

Deep Dive: Finding #4 - Retry Amplification

The most dangerous finding was a retry amplification loop. Three services (A, B, C) had the following call chain: A calls B calls C. Each had 3 retries configured. When C became slow (not failing, just slow), the math was devastating:

  • C responds slowly (2s instead of 200ms)
  • B retries 3 times against C = 4 requests to C per original request
  • B's total time exceeds A's timeout
  • A retries 3 times against B = 4 requests to B
  • Each B retry generates 4 requests to C
  • Result: 1 user request becomes 16 requests to C

With 1000 concurrent users, C received 16,000 requests instead of 1000. This overwhelmed C further, making it slower, which triggered more retries. The amplification factor was 340x at peak before circuit breakers finally opened.

Fix: We implemented retry budgets at the mesh level, limiting total retry percentage to 20% of baseline traffic regardless of individual service retry configuration.

Deep Dive: Finding #1 - DNS Cache During AZ Failover

The most impactful finding involved DNS resolution during AZ failover. When an AZ failed, ALB targets in that AZ were deregistered. The ALB DNS record updated to point only to healthy AZ endpoints. However:

  • Java services cached DNS for 30 seconds (JVM default)
  • Node.js services cached DNS for 5 seconds (OS default)
  • Go services did not cache DNS at all

During the 30-second window, Java services continued sending traffic to the dead AZ. But worse: some services had custom DNS resolution that cached the ALB's IP addresses rather than re-resolving the CNAME chain, resulting in traffic sent to deregistered targets for up to 5 minutes.

Fix: Standardized DNS TTL handling across all services, implemented client-side health checking that supplemented DNS-based routing, and reduced ALB deregistration delay from 300s to 30s.

Experiment Scheduling and Safety

We run chaos experiments during business hours with the full team present. Our safety framework:

# chaos-schedule.yaml
experiments:
  frequency: weekly
  window:
    day: tuesday
    start: "10:00"
    end: "14:00"
    timezone: "America/New_York"
  
  safety:
    require_approval: true
    approvers: ["sre-lead", "service-owner"]
    abort_on:
      - error_rate_exceeds: "1%"
      - latency_p99_exceeds: "5s"
      - customer_impact_detected: true
    
    excluded_periods:
      - type: "deploy_in_progress"
      - type: "active_incident"
      - type: "traffic_spike"  # >2x normal
      - type: "holiday_freeze"

Results After 6 Months

MetricBefore ChaosAfter RemediationChange
Unplanned outage minutes/quarter22 min4 min-82%
Hidden failure modes discovered012N/A
Blast radius of average incident3.4 services1.2 services-65%
Recovery time (AZ failure)4.5 minutes38 seconds-86%
Confidence in DR proceduresLow (untested)High (tested monthly)Qualitative

Building a Chaos Culture

Technical implementation was the easy part. Cultural adoption required:

  1. Reframe failure as learning. Chaos findings are celebrated, not blamed. Each finding represents an outage prevented.
  2. Start small. First experiments targeted non-critical services in staging. Production experiments came after teams built confidence.
  3. Game days build muscle memory. Monthly game days where the on-call team responds to injected failures built incident response skills without the 3 AM stress.
  4. Share findings broadly. Every chaos finding gets a written report shared company-wide. This builds organizational resilience knowledge.

Conclusion

Chaos engineering revealed that our 99.95% uptime was partially luck. Twelve critical failure modes lurked in our system, each requiring specific failure combinations that normal operation never triggered. Systematic experimentation found them before customers did. The 82% reduction in unplanned outage minutes came not from chaos engineering itself but from fixing what chaos engineering revealed. Start with a hypothesis, define your abort conditions, and break things intentionally before they break accidentally.

Comments

    No comments yet. Be the first to share your thoughts.