Chaos Engineering That Found 12 Critical Failure Modes We Never Anticipated
How systematic chaos experiments in production uncovered hidden dependencies, cascade failures, and timeout misconfigurations across our distributed system

We had 99.95% uptime, comprehensive monitoring, and circuit breakers on every external call. We thought our system was resilient. Then we started running chaos experiments in production and discovered 12 critical failure modes that monitoring never caught because they required specific combinations of failures that normal operation never triggered. One of those hidden modes, if triggered organically, would have caused a 4-hour outage affecting all customers. Here is what we found and how we found it.
The Problem: Untested Failure Assumptions
Resilience engineering is full of assumptions. "The circuit breaker will open before cascade failure." "The retry budget will prevent thundering herds." "The fallback cache will serve stale data gracefully." These assumptions are rarely tested under realistic conditions. Unit tests mock the failure. Integration tests test one failure at a time. Production has correlated failures.
Our pre-chaos confidence was based on:
- Circuit breakers configured on 47 external dependencies
- Retry logic with exponential backoff on all HTTP clients
- Health checks on every service with automated restart
- Multi-AZ deployment across three availability zones
- 99.95% uptime over the previous 12 months
All of this gave us confidence. Chaos engineering gave us truth.
Chaos Architecture
We built our chaos platform on AWS Fault Injection Service (FIS) combined with custom chaos controllers for application-level faults:
The platform operates at three levels:
- Infrastructure chaos: AZ failures, network partitions, instance termination
- Application chaos: Latency injection, error responses, connection pool exhaustion
- Dependency chaos: Third-party API failures, database failover, cache eviction
Experiment Framework
Every chaos experiment follows a structured hypothesis-experiment-analysis cycle:
from dataclasses import dataclass, field
from datetime import datetime
from enum import Enum
from typing import Callable
class ExperimentState(Enum):
DRAFT = "draft"
APPROVED = "approved"
RUNNING = "running"
ANALYZING = "analyzing"
COMPLETE = "complete"
ABORTED = "aborted"
@dataclass
class ChaosExperiment:
"""Structured chaos experiment definition."""
id: str
name: str
hypothesis: str
steady_state: dict # Metrics that define "normal"
blast_radius: str # What could be affected
abort_conditions: list[str] # When to stop immediately
fault_injection: dict # What we're breaking
duration_minutes: int
owner: str
approved_by: str = ""
state: ExperimentState = ExperimentState.DRAFT
findings: list[str] = field(default_factory=list)
def validate_steady_state(self) -> bool:
"""Verify system is in normal state before injecting faults."""
for metric, threshold in self.steady_state.items():
current_value = get_metric_value(metric)
if not meets_threshold(current_value, threshold):
print(f"ABORT: Steady state violation - {metric}={current_value}")
return False
return True
def check_abort_conditions(self) -> bool:
"""Continuously check if experiment should be stopped."""
for condition in self.abort_conditions:
if evaluate_condition(condition):
self.state = ExperimentState.ABORTED
return True
return False
# Example experiment definition
experiment_001 = ChaosExperiment(
id="chaos-001",
name="Payment service AZ failure",
hypothesis="If us-east-1a becomes unavailable, payment processing "
"continues via us-east-1b/c with <500ms additional latency",
steady_state={
"payment_success_rate": ">= 99.9%",
"payment_p99_latency": "<= 800ms",
"error_rate_5xx": "<= 0.1%",
},
blast_radius="Payment processing in AZ us-east-1a (33% of traffic)",
abort_conditions=[
"payment_success_rate < 99.0%",
"payment_p99_latency > 5000ms",
"error_rate_5xx > 1.0%",
],
fault_injection={
"type": "az_failure",
"target_az": "us-east-1a",
"services": ["payment-service"],
"method": "network_partition",
},
duration_minutes=15,
owner="@sre-team",
approved_by="@vp-engineering",
)
AWS FIS Experiment Template
For infrastructure-level chaos, we use FIS experiment templates:
{
"description": "Simulate AZ failure for payment service",
"targets": {
"payment-instances": {
"resourceType": "aws:ec2:instance",
"selectionMode": "ALL",
"resourceTags": {
"service": "payment-service",
"az": "us-east-1a"
},
"filters": [
{
"path": "State.Name",
"values": ["running"]
}
]
},
"payment-subnets": {
"resourceType": "aws:ec2:subnet",
"selectionMode": "ALL",
"resourceTags": {
"service": "payment-service",
"az": "us-east-1a"
}
}
},
"actions": {
"inject-network-partition": {
"actionId": "aws:network:disrupt-connectivity",
"parameters": {
"scope": "all",
"duration": "PT15M"
},
"targets": {
"Subnets": "payment-subnets"
}
},
"stop-instances": {
"actionId": "aws:ec2:stop-instances",
"parameters": {
"startInstancesAfterDuration": "PT15M"
},
"targets": {
"Instances": "payment-instances"
},
"startAfter": ["inject-network-partition"]
}
},
"stopConditions": [
{
"source": "aws:cloudwatch:alarm",
"value": "arn:aws:cloudwatch:us-east-1:123456789:alarm:chaos-abort-payment"
}
],
"roleArn": "arn:aws:iam::123456789:role/fis-chaos-role"
}
The 12 Critical Findings
Over six months of systematic experimentation, we discovered these failure modes:
| # | Finding | Severity | Impact if Triggered |
|---|---|---|---|
| 1 | DNS cache poisoning during AZ failover | Critical | 4-hour outage |
| 2 | Circuit breaker timeout mismatch cascading to all services | Critical | 90-minute degradation |
| 3 | Connection pool leak under partial network partition | High | Service unavailability in 45min |
| 4 | Retry amplification between three services creating feedback loop | Critical | Traffic 340x amplification |
| 5 | Health check passing while service unable to process requests | High | Silent failure for 12 minutes |
| 6 | Cache stampede after Redis cluster failover | High | Database overload for 8 minutes |
| 7 | Async message ordering violation during consumer restart | Medium | Data inconsistency |
| 8 | TLS certificate rotation failing under CPU pressure | High | mTLS failure after 12 hours |
| 9 | Graceful shutdown not draining in-flight requests | Medium | 0.3% request failure during deploys |
| 10 | Database connection exhaustion from leaked prepared statements | High | Progressive degradation over 6 hours |
| 11 | Cross-region failover DNS TTL exceeding client cache | Critical | 30-minute unavailability |
| 12 | Rate limiter shared state lost during pod rescheduling | Medium | Temporary rate limit bypass |
Deep Dive: Finding #4 - Retry Amplification
The most dangerous finding was a retry amplification loop. Three services (A, B, C) had the following call chain: A calls B calls C. Each had 3 retries configured. When C became slow (not failing, just slow), the math was devastating:
- C responds slowly (2s instead of 200ms)
- B retries 3 times against C = 4 requests to C per original request
- B's total time exceeds A's timeout
- A retries 3 times against B = 4 requests to B
- Each B retry generates 4 requests to C
- Result: 1 user request becomes 16 requests to C
With 1000 concurrent users, C received 16,000 requests instead of 1000. This overwhelmed C further, making it slower, which triggered more retries. The amplification factor was 340x at peak before circuit breakers finally opened.
Fix: We implemented retry budgets at the mesh level, limiting total retry percentage to 20% of baseline traffic regardless of individual service retry configuration.
Deep Dive: Finding #1 - DNS Cache During AZ Failover
The most impactful finding involved DNS resolution during AZ failover. When an AZ failed, ALB targets in that AZ were deregistered. The ALB DNS record updated to point only to healthy AZ endpoints. However:
- Java services cached DNS for 30 seconds (JVM default)
- Node.js services cached DNS for 5 seconds (OS default)
- Go services did not cache DNS at all
During the 30-second window, Java services continued sending traffic to the dead AZ. But worse: some services had custom DNS resolution that cached the ALB's IP addresses rather than re-resolving the CNAME chain, resulting in traffic sent to deregistered targets for up to 5 minutes.
Fix: Standardized DNS TTL handling across all services, implemented client-side health checking that supplemented DNS-based routing, and reduced ALB deregistration delay from 300s to 30s.
Experiment Scheduling and Safety
We run chaos experiments during business hours with the full team present. Our safety framework:
# chaos-schedule.yaml
experiments:
frequency: weekly
window:
day: tuesday
start: "10:00"
end: "14:00"
timezone: "America/New_York"
safety:
require_approval: true
approvers: ["sre-lead", "service-owner"]
abort_on:
- error_rate_exceeds: "1%"
- latency_p99_exceeds: "5s"
- customer_impact_detected: true
excluded_periods:
- type: "deploy_in_progress"
- type: "active_incident"
- type: "traffic_spike" # >2x normal
- type: "holiday_freeze"
Results After 6 Months
| Metric | Before Chaos | After Remediation | Change |
|---|---|---|---|
| Unplanned outage minutes/quarter | 22 min | 4 min | -82% |
| Hidden failure modes discovered | 0 | 12 | N/A |
| Blast radius of average incident | 3.4 services | 1.2 services | -65% |
| Recovery time (AZ failure) | 4.5 minutes | 38 seconds | -86% |
| Confidence in DR procedures | Low (untested) | High (tested monthly) | Qualitative |
Building a Chaos Culture
Technical implementation was the easy part. Cultural adoption required:
- Reframe failure as learning. Chaos findings are celebrated, not blamed. Each finding represents an outage prevented.
- Start small. First experiments targeted non-critical services in staging. Production experiments came after teams built confidence.
- Game days build muscle memory. Monthly game days where the on-call team responds to injected failures built incident response skills without the 3 AM stress.
- Share findings broadly. Every chaos finding gets a written report shared company-wide. This builds organizational resilience knowledge.
Conclusion
Chaos engineering revealed that our 99.95% uptime was partially luck. Twelve critical failure modes lurked in our system, each requiring specific failure combinations that normal operation never triggered. Systematic experimentation found them before customers did. The 82% reduction in unplanned outage minutes came not from chaos engineering itself but from fixing what chaos engineering revealed. Start with a hypothesis, define your abort conditions, and break things intentionally before they break accidentally.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.