Automated Incident Response Playbooks: Reducing MTTR by 78%

How we built automated incident response runbooks with AWS Systems Manager and Step Functions, cutting mean time to resolution from 47 to 10 minutes

#incident-response#automation#sre#aws
Cover image for the article: Automated Incident Response Playbooks: Reducing MTTR by 78%

At 2:47 AM, a database connection pool exhaustion alert fires. The on-call engineer wakes up, opens their laptop, tries to remember the runbook, SSH'es into a bastion, runs diagnostic commands, identifies the issue, and applies the fix. Total time: 47 minutes average. Most of that time is not problem-solving but mechanical execution of known procedures. We automated 23 incident response playbooks using AWS Systems Manager and Step Functions, reducing our mean time to resolution from 47 minutes to 10 minutes.

The Problem: Human Latency in Known Scenarios

Analysis of 18 months of incident data revealed a pattern: 72% of our incidents fell into known categories with documented resolution procedures. The problem was not diagnosis but execution speed. Humans are slow at 3 AM.

Incident response timeline breakdown (average):

  • Wake up and acknowledge: 4.2 minutes
  • Open laptop and connect: 3.8 minutes
  • Read runbook and orient: 8.4 minutes
  • Execute diagnostic commands: 12.1 minutes
  • Apply remediation: 9.3 minutes
  • Verify resolution: 6.7 minutes
  • Document and close: 2.5 minutes
  • Total: 47 minutes

Of these 47 minutes, only the diagnosis step (12.1 minutes) requires human judgment. Everything else is mechanical.

Architecture: Event-Driven Response

The automated response system watches for specific alert patterns, matches them to playbooks, executes automated diagnostics, applies safe remediations, and escalates to humans only when automation cannot resolve the issue.

Incident Response Automation Architecture

The system integrates with our existing alerting stack (CloudWatch Alarms via SNS to PagerDuty) by intercepting the alert pipeline and attempting automated resolution before human notification.

Playbook: Database Connection Pool Exhaustion

Our most common incident type. The automated playbook handles the full lifecycle:

# ssm-playbooks/db-connection-pool-exhaustion.yaml
description: "Automated response for RDS connection pool exhaustion"
schemaVersion: '0.3'
assumeRole: "{{ AutomationAssumeRole }}"
parameters:
  DBInstanceIdentifier:
    type: String
    description: "RDS instance identifier"
  AlarmName:
    type: String
    description: "CloudWatch alarm that triggered"
  MaxConnectionThreshold:
    type: Integer
    default: 90
    description: "Connection percentage threshold"

mainSteps:
  - name: DiagnoseConnectionState
    action: aws:executeScript
    inputs:
      Runtime: python3.11
      Handler: handler
      Script: |
        import boto3

        def handler(event, context):
            rds = boto3.client('rds')
            cw = boto3.client('cloudwatch')
            
            db_id = event['DBInstanceIdentifier']
            
            # Get current connection count
            metrics = cw.get_metric_statistics(
                Namespace='AWS/RDS',
                MetricName='DatabaseConnections',
                Dimensions=[{'Name': 'DBInstanceIdentifier', 'Value': db_id}],
                StartTime=datetime.utcnow() - timedelta(minutes=10),
                EndTime=datetime.utcnow(),
                Period=60,
                Statistics=['Maximum', 'Average']
            )
            
            # Get instance class max connections
            instance = rds.describe_db_instances(
                DBInstanceIdentifier=db_id
            )['DBInstances'][0]
            
            max_connections = get_max_connections(instance['DBInstanceClass'])
            current = metrics['Datapoints'][-1]['Maximum']
            utilization = (current / max_connections) * 100
            
            return {
                'current_connections': int(current),
                'max_connections': max_connections,
                'utilization_percent': round(utilization, 1),
                'instance_class': instance['DBInstanceClass'],
                'requires_intervention': utilization > 95
            }
    outputs:
      - Name: Diagnosis
        Selector: $.Payload
        Type: StringMap

  - name: IdentifyLeakingServices
    action: aws:executeScript
    inputs:
      Runtime: python3.11
      Handler: handler
      Script: |
        import boto3
        import json

        def handler(event, context):
            # Query performance insights for top connection holders
            pi = boto3.client('pi')
            
            response = pi.get_resource_metrics(
                ServiceType='RDS',
                Identifier=f"db-{event['DBInstanceIdentifier']}",
                MetricQueries=[{
                    'Metric': 'db.Connections',
                    'GroupBy': {'Group': 'db.application.name'}
                }],
                StartTime=datetime.utcnow() - timedelta(minutes=30),
                EndTime=datetime.utcnow(),
                PeriodInSeconds=300
            )
            
            # Identify services holding disproportionate connections
            top_consumers = sorted(
                response['MetricList'][0]['DataPoints'],
                key=lambda x: x['Value'],
                reverse=True
            )[:5]
            
            return {
                'top_consumers': top_consumers,
                'recommended_action': determine_action(top_consumers)
            }

  - name: ApplyRemediation
    action: aws:branch
    inputs:
      Choices:
        - NextStep: RestartLeakingService
          Variable: "{{ IdentifyLeakingServices.recommended_action }}"
          StringEquals: "restart_service"
        - NextStep: ScaleConnectionPool
          Variable: "{{ IdentifyLeakingServices.recommended_action }}"
          StringEquals: "scale_pool"
        - NextStep: EscalateToHuman
          Variable: "{{ IdentifyLeakingServices.recommended_action }}"
          StringEquals: "escalate"

  - name: RestartLeakingService
    action: aws:executeScript
    inputs:
      Runtime: python3.11
      Handler: handler
      Script: |
        import boto3

        def handler(event, context):
            ecs = boto3.client('ecs')
            service_name = event['leaking_service']
            
            # Rolling restart - not a hard kill
            ecs.update_service(
                cluster='production',
                service=service_name,
                forceNewDeployment=True
            )
            
            return {'action': 'rolling_restart', 'service': service_name}

  - name: VerifyResolution
    action: aws:waitForAwsResourceProperty
    timeoutSeconds: 300
    inputs:
      Service: cloudwatch
      Api: DescribeAlarms
      AlarmNames:
        - "{{ AlarmName }}"
      PropertySelector: "$.MetricAlarms[0].StateValue"
      DesiredValues:
        - "OK"

This playbook typically resolves connection pool exhaustion in 3-4 minutes versus the previous 47-minute manual process.

Step Functions Orchestrator

For complex incidents requiring multi-step coordination, we use Step Functions to orchestrate multiple SSM playbooks:

{
  "Comment": "Incident Response Orchestrator",
  "StartAt": "ClassifyIncident",
  "States": {
    "ClassifyIncident": {
      "Type": "Task",
      "Resource": "arn:aws:lambda:us-east-1:123456789:function:classify-incident",
      "Next": "SelectPlaybook",
      "Catch": [{
        "ErrorEquals": ["States.ALL"],
        "Next": "EscalateToHuman"
      }]
    },
    "SelectPlaybook": {
      "Type": "Choice",
      "Choices": [
        {
          "Variable": "$.incident_type",
          "StringEquals": "connection_pool_exhaustion",
          "Next": "RunDBConnectionPlaybook"
        },
        {
          "Variable": "$.incident_type",
          "StringEquals": "memory_pressure",
          "Next": "RunMemoryPlaybook"
        },
        {
          "Variable": "$.incident_type",
          "StringEquals": "disk_space",
          "Next": "RunDiskPlaybook"
        }
      ],
      "Default": "EscalateToHuman"
    },
    "RunDBConnectionPlaybook": {
      "Type": "Task",
      "Resource": "arn:aws:ssm:us-east-1:123456789:automation-definition/db-connection-pool-exhaustion",
      "Next": "ValidateResolution",
      "TimeoutSeconds": 600
    },
    "ValidateResolution": {
      "Type": "Task",
      "Resource": "arn:aws:lambda:us-east-1:123456789:function:validate-resolution",
      "Next": "ResolutionCheck"
    },
    "ResolutionCheck": {
      "Type": "Choice",
      "Choices": [{
        "Variable": "$.resolved",
        "BooleanEquals": true,
        "Next": "DocumentAndClose"
      }],
      "Default": "EscalateToHuman"
    },
    "EscalateToHuman": {
      "Type": "Task",
      "Resource": "arn:aws:lambda:us-east-1:123456789:function:page-oncall",
      "End": true
    },
    "DocumentAndClose": {
      "Type": "Task",
      "Resource": "arn:aws:lambda:us-east-1:123456789:function:close-incident",
      "End": true
    }
  }
}

The orchestrator follows a strict principle: if automation cannot resolve the issue within 10 minutes, escalate to a human immediately. We never let automation thrash on a novel problem.

Safety Guardrails

Automated remediation without guardrails is dangerous. We implemented several safety mechanisms:

  1. Blast radius limits: Automation can restart at most one service at a time, never multiple simultaneously.
  2. Cool-down periods: After an automated remediation, the same playbook cannot trigger again for 15 minutes.
  3. Human override: Any on-call engineer can disable automation for a specific service via a Slack command.
  4. Dry-run mode: New playbooks run in observation mode for two weeks before enabling remediation actions.
  5. Audit trail: Every automated action is logged to CloudTrail and our incident management system.

Results: 6-Month Metrics

MetricBeforeAfterImprovement
Mean time to resolution47 min10.2 min-78%
Incidents resolved without human wake0%64%+64%
On-call pages per week238.3-64%
Incidents auto-resolved0/month89/monthN/A
Playbook coverage of incident types0%72%+72%
False positive auto-remediationN/A2.1%Acceptable

The 64% of incidents resolved without waking anyone up meant engineers got significantly more uninterrupted sleep. The 2.1% false positive rate (automation applied remediation when it was not needed) was acceptable because all automated remediations are inherently safe actions (rolling restarts, scaling up, clearing caches).

Playbook Development Process

We follow a structured process for adding new automated playbooks:

  1. Incident pattern identification: Three or more incidents of the same type triggers playbook development.
  2. Manual execution documentation: An engineer runs the manual process while recording every command and decision point.
  3. Decision tree mapping: We identify which steps require judgment and which are mechanical.
  4. Automation of mechanical steps: Only well-understood, safe actions are automated.
  5. Two-week dry-run: The playbook runs in observation mode, logging what it would do without taking action.
  6. Gradual enablement: First enabled in staging, then production with human confirmation, then fully automated.

Conclusion

Incident response automation is not about replacing engineers. It is about eliminating the mechanical latency that makes known-resolution incidents take 47 minutes instead of 4. The 78% MTTR reduction came from automating what humans already knew how to do but were slow to execute at 3 AM. Start with your most common incident type, automate the diagnostics first, add safe remediation second, and always maintain a fast escalation path to humans. The goal is fewer pages, not fewer engineers.

Comments

    No comments yet. Be the first to share your thoughts.