Self-Healing Systems That Resolve 67% of Alerts Automatically
How to build runbook automation that transforms manual incident response into self-healing infrastructure, reducing human intervention to only the incidents that truly need it.

The Problem: Humans as Middleware
Every mature operations team has runbooks—step-by-step documents for resolving known issues. The irony: if the steps are well-defined enough to document, they're well-defined enough to automate. Yet most teams still wake engineers at 3 AM to follow a script that a computer could execute in seconds.
Our SRE team maintained 83 runbooks covering known failure modes across 47 microservices. We tracked resolution patterns over 6 months and found that 71% of incidents followed an existing runbook exactly. Engineers were acting as middleware—receiving an alert, looking up the runbook, executing the steps, then going back to sleep. The median resolution time was 14 minutes, but median acknowledgment time was 8 minutes. The system knew what to do; it just needed permission.
After implementing runbook automation, 67% of previously-paging alerts now self-heal without human intervention. Mean time to resolution for automated incidents dropped from 14 minutes to 47 seconds.
Architecture: The Self-Healing Loop
Our self-healing system follows a detect-diagnose-remediate-verify loop:
- Detect: Alertmanager fires an alert with structured labels
- Diagnose: Automation engine matches alert to a remediation playbook
- Remediate: Execute the automated fix with safety bounds
- Verify: Confirm the fix worked using the same SLI that triggered the alert
- Notify: Create a ticket for review (not a page)
Remediation Engine Design
The core of our system is a remediation engine that matches alerts to playbooks and executes them with safety constraints:
# remediation/engine.py
from dataclasses import dataclass, field
from enum import Enum
from datetime import datetime, timedelta
import asyncio
class RemediationStatus(Enum):
PENDING = "pending"
EXECUTING = "executing"
VERIFYING = "verifying"
SUCCESS = "success"
FAILED = "failed"
ESCALATED = "escalated"
@dataclass
class SafetyBounds:
max_executions_per_hour: int = 3
cooldown_minutes: int = 15
max_blast_radius_pods: int = 5
require_canary: bool = True
rollback_on_failure: bool = True
@dataclass
class Playbook:
alert_name: str
remediation_steps: list[str]
safety_bounds: SafetyBounds
verification_query: str
verification_timeout_seconds: int = 120
escalate_after_failures: int = 2
_execution_history: list[datetime] = field(default_factory=list)
def can_execute(self) -> tuple[bool, str]:
"""Check if safety bounds allow execution."""
recent = [
t for t in self._execution_history
if t > datetime.now() - timedelta(hours=1)
]
if len(recent) >= self.safety_bounds.max_executions_per_hour:
return False, f"Rate limit: {len(recent)} executions in last hour"
if self._execution_history:
last = self._execution_history[-1]
cooldown_end = last + timedelta(minutes=self.safety_bounds.cooldown_minutes)
if datetime.now() < cooldown_end:
return False, f"Cooldown active until {cooldown_end}"
return True, "OK"
Playbook Definitions
We define playbooks in YAML for each known failure mode:
# playbooks/pod-crashloop.yaml
name: pod-crashloop-recovery
alert_match:
alertname: PodCrashLoopBackOff
namespace: production
safety_bounds:
max_executions_per_hour: 3
cooldown_minutes: 10
max_blast_radius_pods: 3
require_canary: false
rollback_on_failure: true
diagnosis:
- check: oom_killed
query: |
kube_pod_container_status_last_terminated_reason{
reason="OOMKilled",
pod=~"{{ .Labels.pod }}"
}
if_true: escalate # OOM needs human investigation
- check: exit_code
query: |
kube_pod_container_status_last_terminated_exitcode{
pod=~"{{ .Labels.pod }}"
}
if_value_gt: 1
action: restart_with_increased_memory
remediation_steps:
restart_clean:
- action: kubectl_delete_pod
target: "{{ .Labels.pod }}"
namespace: "{{ .Labels.namespace }}"
wait_for_ready: true
timeout: 120s
restart_with_increased_memory:
- action: kubectl_patch_deployment
target: "{{ .Labels.deployment }}"
patch: |
spec:
template:
spec:
containers:
- name: "{{ .Labels.container }}"
resources:
requests:
memory: "{{ mul .CurrentMemory 1.5 }}"
limits:
memory: "{{ mul .CurrentMemory 2 }}"
- action: kubectl_rollout_restart
target: "{{ .Labels.deployment }}"
verification:
query: |
kube_pod_status_ready{
pod=~"{{ .Labels.deployment }}.*",
namespace="{{ .Labels.namespace }}"
} == 1
expected: true
timeout: 180s
escalation:
after_failures: 2
channel: pagerduty
context: "Auto-remediation failed twice for {{ .Labels.pod }}"
Webhook Integration with Alertmanager
Alertmanager sends alerts to our remediation engine via webhook:
# alertmanager-config.yaml
receivers:
- name: self-healing
webhook_configs:
- url: 'http://remediation-engine.sre:8080/alerts'
send_resolved: true
max_alerts: 10
route:
routes:
# Self-healing eligible alerts go to automation first
- match:
self_healing: "enabled"
receiver: self-healing
continue: true # Also notify humans if automation fails
- match:
self_healing: "enabled"
auto_remediation_failed: "true"
receiver: pagerduty-oncall
Safety: The Non-Negotiable Constraints
Self-healing without safety constraints is a recipe for cascade failures. Our system enforces:
# remediation/safety.py
from dataclasses import dataclass
@dataclass
class GlobalSafetyPolicy:
"""System-wide constraints that override playbook settings."""
# Never auto-remediate during active incidents
halt_during_incident: bool = True
# Maximum concurrent remediations across all services
max_concurrent_global: int = 3
# Never auto-remediate data stores
protected_namespaces: list[str] = None
# Require at least N healthy pods before touching any
min_healthy_ratio: float = 0.5
# Circuit breaker: stop all automation if failure rate exceeds threshold
circuit_breaker_threshold: float = 0.3 # 30% failure rate
circuit_breaker_window_minutes: int = 60
def __post_init__(self):
if self.protected_namespaces is None:
self.protected_namespaces = [
"database", "kafka", "elasticsearch",
"vault", "cert-manager"
]
def is_safe_to_remediate(self, context: dict) -> tuple[bool, str]:
if context.get("active_incident"):
return False, "Active incident in progress"
if context["namespace"] in self.protected_namespaces:
return False, f"Namespace {context['namespace']} is protected"
healthy_ratio = context["healthy_pods"] / context["desired_pods"]
if healthy_ratio < self.min_healthy_ratio:
return False, f"Healthy ratio {healthy_ratio:.0%} below minimum"
return True, "Safe to proceed"
Measuring Success
We track automation effectiveness with these metrics:
| Metric | Month 1 | Month 6 |
|---|---|---|
| Alerts eligible for automation | 23% | 67% |
| Auto-resolved successfully | 78% | 94% |
| MTTR (automated incidents) | 3.2 min | 47 sec |
| False remediations | 4 | 0 |
| Escalations from automation | 22% | 6% |
| Human pages eliminated/week | 4.8 | 11.2 |
Graduating Runbooks to Automation
Not every runbook should be automated on day one. We use a maturity model:
- Level 0 - Document: Write the runbook. Track how often it's used.
- Level 1 - Assist: Automation gathers diagnostics and suggests next steps. Human executes.
- Level 2 - Supervised: Automation proposes a fix and executes on human approval (Slack button).
- Level 3 - Autonomous: Automation executes within safety bounds. Human notified after.
- Level 4 - Invisible: Issue detected and fixed before it becomes an alert.
Playbooks graduate through levels based on execution success rate:
- 10+ successful supervised executions → eligible for autonomous
- 0 false remediations in 30 days → eligible for invisible
Key Takeaways
-
Start with the highest-frequency alerts. Automating your top 5 most common alerts eliminates more toil than automating 20 rare ones.
-
Safety bounds are more important than the automation itself. Rate limits, cooldowns, circuit breakers, and protected namespaces prevent automation from becoming a force multiplier for outages.
-
Diagnosis before remediation. The same symptom can have different root causes. Automated diagnosis (is it OOM? Is it a bad deploy? Is it upstream?) determines the correct response.
-
Graduate playbooks through maturity levels. Don't automate everything at once. Build trust through supervised execution before granting autonomy.
-
Always verify the fix. A remediation that doesn't verify it worked is just creating a different problem.
The goal isn't zero human involvement—it's human involvement only for novel problems that genuinely need creative thinking.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.