Automated Incident Response Playbooks: Reducing MTTR by 78%
How we built automated incident response runbooks with AWS Systems Manager and Step Functions, cutting mean time to resolution from 47 to 10 minutes

At 2:47 AM, a database connection pool exhaustion alert fires. The on-call engineer wakes up, opens their laptop, tries to remember the runbook, SSH'es into a bastion, runs diagnostic commands, identifies the issue, and applies the fix. Total time: 47 minutes average. Most of that time is not problem-solving but mechanical execution of known procedures. We automated 23 incident response playbooks using AWS Systems Manager and Step Functions, reducing our mean time to resolution from 47 minutes to 10 minutes.
The Problem: Human Latency in Known Scenarios
Analysis of 18 months of incident data revealed a pattern: 72% of our incidents fell into known categories with documented resolution procedures. The problem was not diagnosis but execution speed. Humans are slow at 3 AM.
Incident response timeline breakdown (average):
- Wake up and acknowledge: 4.2 minutes
- Open laptop and connect: 3.8 minutes
- Read runbook and orient: 8.4 minutes
- Execute diagnostic commands: 12.1 minutes
- Apply remediation: 9.3 minutes
- Verify resolution: 6.7 minutes
- Document and close: 2.5 minutes
- Total: 47 minutes
Of these 47 minutes, only the diagnosis step (12.1 minutes) requires human judgment. Everything else is mechanical.
Architecture: Event-Driven Response
The automated response system watches for specific alert patterns, matches them to playbooks, executes automated diagnostics, applies safe remediations, and escalates to humans only when automation cannot resolve the issue.
The system integrates with our existing alerting stack (CloudWatch Alarms via SNS to PagerDuty) by intercepting the alert pipeline and attempting automated resolution before human notification.
Playbook: Database Connection Pool Exhaustion
Our most common incident type. The automated playbook handles the full lifecycle:
# ssm-playbooks/db-connection-pool-exhaustion.yaml
description: "Automated response for RDS connection pool exhaustion"
schemaVersion: '0.3'
assumeRole: "{{ AutomationAssumeRole }}"
parameters:
DBInstanceIdentifier:
type: String
description: "RDS instance identifier"
AlarmName:
type: String
description: "CloudWatch alarm that triggered"
MaxConnectionThreshold:
type: Integer
default: 90
description: "Connection percentage threshold"
mainSteps:
- name: DiagnoseConnectionState
action: aws:executeScript
inputs:
Runtime: python3.11
Handler: handler
Script: |
import boto3
def handler(event, context):
rds = boto3.client('rds')
cw = boto3.client('cloudwatch')
db_id = event['DBInstanceIdentifier']
# Get current connection count
metrics = cw.get_metric_statistics(
Namespace='AWS/RDS',
MetricName='DatabaseConnections',
Dimensions=[{'Name': 'DBInstanceIdentifier', 'Value': db_id}],
StartTime=datetime.utcnow() - timedelta(minutes=10),
EndTime=datetime.utcnow(),
Period=60,
Statistics=['Maximum', 'Average']
)
# Get instance class max connections
instance = rds.describe_db_instances(
DBInstanceIdentifier=db_id
)['DBInstances'][0]
max_connections = get_max_connections(instance['DBInstanceClass'])
current = metrics['Datapoints'][-1]['Maximum']
utilization = (current / max_connections) * 100
return {
'current_connections': int(current),
'max_connections': max_connections,
'utilization_percent': round(utilization, 1),
'instance_class': instance['DBInstanceClass'],
'requires_intervention': utilization > 95
}
outputs:
- Name: Diagnosis
Selector: $.Payload
Type: StringMap
- name: IdentifyLeakingServices
action: aws:executeScript
inputs:
Runtime: python3.11
Handler: handler
Script: |
import boto3
import json
def handler(event, context):
# Query performance insights for top connection holders
pi = boto3.client('pi')
response = pi.get_resource_metrics(
ServiceType='RDS',
Identifier=f"db-{event['DBInstanceIdentifier']}",
MetricQueries=[{
'Metric': 'db.Connections',
'GroupBy': {'Group': 'db.application.name'}
}],
StartTime=datetime.utcnow() - timedelta(minutes=30),
EndTime=datetime.utcnow(),
PeriodInSeconds=300
)
# Identify services holding disproportionate connections
top_consumers = sorted(
response['MetricList'][0]['DataPoints'],
key=lambda x: x['Value'],
reverse=True
)[:5]
return {
'top_consumers': top_consumers,
'recommended_action': determine_action(top_consumers)
}
- name: ApplyRemediation
action: aws:branch
inputs:
Choices:
- NextStep: RestartLeakingService
Variable: "{{ IdentifyLeakingServices.recommended_action }}"
StringEquals: "restart_service"
- NextStep: ScaleConnectionPool
Variable: "{{ IdentifyLeakingServices.recommended_action }}"
StringEquals: "scale_pool"
- NextStep: EscalateToHuman
Variable: "{{ IdentifyLeakingServices.recommended_action }}"
StringEquals: "escalate"
- name: RestartLeakingService
action: aws:executeScript
inputs:
Runtime: python3.11
Handler: handler
Script: |
import boto3
def handler(event, context):
ecs = boto3.client('ecs')
service_name = event['leaking_service']
# Rolling restart - not a hard kill
ecs.update_service(
cluster='production',
service=service_name,
forceNewDeployment=True
)
return {'action': 'rolling_restart', 'service': service_name}
- name: VerifyResolution
action: aws:waitForAwsResourceProperty
timeoutSeconds: 300
inputs:
Service: cloudwatch
Api: DescribeAlarms
AlarmNames:
- "{{ AlarmName }}"
PropertySelector: "$.MetricAlarms[0].StateValue"
DesiredValues:
- "OK"
This playbook typically resolves connection pool exhaustion in 3-4 minutes versus the previous 47-minute manual process.
Step Functions Orchestrator
For complex incidents requiring multi-step coordination, we use Step Functions to orchestrate multiple SSM playbooks:
{
"Comment": "Incident Response Orchestrator",
"StartAt": "ClassifyIncident",
"States": {
"ClassifyIncident": {
"Type": "Task",
"Resource": "arn:aws:lambda:us-east-1:123456789:function:classify-incident",
"Next": "SelectPlaybook",
"Catch": [{
"ErrorEquals": ["States.ALL"],
"Next": "EscalateToHuman"
}]
},
"SelectPlaybook": {
"Type": "Choice",
"Choices": [
{
"Variable": "$.incident_type",
"StringEquals": "connection_pool_exhaustion",
"Next": "RunDBConnectionPlaybook"
},
{
"Variable": "$.incident_type",
"StringEquals": "memory_pressure",
"Next": "RunMemoryPlaybook"
},
{
"Variable": "$.incident_type",
"StringEquals": "disk_space",
"Next": "RunDiskPlaybook"
}
],
"Default": "EscalateToHuman"
},
"RunDBConnectionPlaybook": {
"Type": "Task",
"Resource": "arn:aws:ssm:us-east-1:123456789:automation-definition/db-connection-pool-exhaustion",
"Next": "ValidateResolution",
"TimeoutSeconds": 600
},
"ValidateResolution": {
"Type": "Task",
"Resource": "arn:aws:lambda:us-east-1:123456789:function:validate-resolution",
"Next": "ResolutionCheck"
},
"ResolutionCheck": {
"Type": "Choice",
"Choices": [{
"Variable": "$.resolved",
"BooleanEquals": true,
"Next": "DocumentAndClose"
}],
"Default": "EscalateToHuman"
},
"EscalateToHuman": {
"Type": "Task",
"Resource": "arn:aws:lambda:us-east-1:123456789:function:page-oncall",
"End": true
},
"DocumentAndClose": {
"Type": "Task",
"Resource": "arn:aws:lambda:us-east-1:123456789:function:close-incident",
"End": true
}
}
}
The orchestrator follows a strict principle: if automation cannot resolve the issue within 10 minutes, escalate to a human immediately. We never let automation thrash on a novel problem.
Safety Guardrails
Automated remediation without guardrails is dangerous. We implemented several safety mechanisms:
- Blast radius limits: Automation can restart at most one service at a time, never multiple simultaneously.
- Cool-down periods: After an automated remediation, the same playbook cannot trigger again for 15 minutes.
- Human override: Any on-call engineer can disable automation for a specific service via a Slack command.
- Dry-run mode: New playbooks run in observation mode for two weeks before enabling remediation actions.
- Audit trail: Every automated action is logged to CloudTrail and our incident management system.
Results: 6-Month Metrics
| Metric | Before | After | Improvement |
|---|---|---|---|
| Mean time to resolution | 47 min | 10.2 min | -78% |
| Incidents resolved without human wake | 0% | 64% | +64% |
| On-call pages per week | 23 | 8.3 | -64% |
| Incidents auto-resolved | 0/month | 89/month | N/A |
| Playbook coverage of incident types | 0% | 72% | +72% |
| False positive auto-remediation | N/A | 2.1% | Acceptable |
The 64% of incidents resolved without waking anyone up meant engineers got significantly more uninterrupted sleep. The 2.1% false positive rate (automation applied remediation when it was not needed) was acceptable because all automated remediations are inherently safe actions (rolling restarts, scaling up, clearing caches).
Playbook Development Process
We follow a structured process for adding new automated playbooks:
- Incident pattern identification: Three or more incidents of the same type triggers playbook development.
- Manual execution documentation: An engineer runs the manual process while recording every command and decision point.
- Decision tree mapping: We identify which steps require judgment and which are mechanical.
- Automation of mechanical steps: Only well-understood, safe actions are automated.
- Two-week dry-run: The playbook runs in observation mode, logging what it would do without taking action.
- Gradual enablement: First enabled in staging, then production with human confirmation, then fully automated.
Conclusion
Incident response automation is not about replacing engineers. It is about eliminating the mechanical latency that makes known-resolution incidents take 47 minutes instead of 4. The 78% MTTR reduction came from automating what humans already knew how to do but were slow to execute at 3 AM. Start with your most common incident type, automate the diagnostics first, add safe remediation second, and always maintain a fast escalation path to humans. The goal is fewer pages, not fewer engineers.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.