Real-Time Incident Classification and Routing with Claude
How we built a Claude-powered incident triage system that classifies severity, identifies root causes, and routes to the right team in under 30 seconds.

At 3 AM, when PagerDuty fires and you're staring at a wall of alerts, the last thing you want is to spend 15 minutes figuring out which team should handle this and how bad it really is. We built a Claude-powered incident triage system that classifies severity, correlates related alerts, identifies probable root causes, and routes to the correct on-call team — all within 30 seconds of the first alert firing.
The system reduced our mean-time-to-engage (MTTE) from 14 minutes to 47 seconds and cut misrouted incidents by 84%.
The Problem
Our platform processes 2.3 million requests per minute across 47 microservices. When something breaks:
- Alert fatigue: 300-400 alerts fire per incident, burying the signal in noise
- Misrouting: 31% of incidents were initially routed to the wrong team
- Severity underestimation: Critical incidents were classified as low-severity 18% of the time
- Slow correlation: Engineers spent 8-12 minutes just understanding which alerts were related
Every minute of downtime cost $4,200. The 14-minute MTTE was burning money.
Architecture
The system sits between our alerting infrastructure (Datadog, PagerDuty) and our incident management platform.
Alert Ingestion and Correlation
The first challenge is grouping related alerts into a single incident context before Claude analyzes anything.
import Anthropic from '@anthropic-ai/sdk';
interface Alert {
id: string;
source: string;
severity: string;
service: string;
message: string;
metrics: Record<string, number>;
timestamp: Date;
labels: Record<string, string>;
}
interface IncidentContext {
alerts: Alert[];
correlatedGroup: string;
timeWindow: { start: Date; end: Date };
affectedServices: string[];
serviceTopology: Record<string, string[]>;
recentDeployments: Deployment[];
historicalIncidents: HistoricalIncident[];
}
class AlertCorrelator {
private alertBuffer: Map<string, Alert[]> = new Map();
private correlationWindow = 120_000; // 2 minutes
async correlate(incomingAlert: Alert): Promise<IncidentContext | null> {
// Buffer alerts and look for patterns
const key = this.getCorrelationKey(incomingAlert);
const buffer = this.alertBuffer.get(key) || [];
buffer.push(incomingAlert);
this.alertBuffer.set(key, buffer);
// Wait for correlation window or trigger on critical alerts immediately
if (incomingAlert.severity === 'critical' || buffer.length >= 5) {
return this.buildIncidentContext(buffer);
}
// Set timeout to process buffer if no more alerts arrive
setTimeout(() => this.processBuffer(key), this.correlationWindow);
return null;
}
private async buildIncidentContext(alerts: Alert[]): Promise<IncidentContext> {
const affectedServices = [...new Set(alerts.map(a => a.service))];
// Fetch service dependency topology
const topology = await this.fetchServiceTopology(affectedServices);
// Check for recent deployments
const deployments = await this.fetchRecentDeployments(affectedServices, '1h');
// Find similar historical incidents
const historical = await this.searchHistoricalIncidents(alerts);
return {
alerts,
correlatedGroup: this.generateGroupId(alerts),
timeWindow: {
start: new Date(Math.min(...alerts.map(a => a.timestamp.getTime()))),
end: new Date(Math.max(...alerts.map(a => a.timestamp.getTime())))
},
affectedServices,
serviceTopology: topology,
recentDeployments: deployments,
historicalIncidents: historical
};
}
private getCorrelationKey(alert: Alert): string {
// Correlate by service cluster and time proximity
return `${alert.service.split('-')[0]}_${Math.floor(alert.timestamp.getTime() / 60000)}`;
}
}
Claude-Powered Triage
Once we have correlated context, Claude performs multi-dimensional analysis.
import anthropic
import json
from datetime import datetime
class IncidentTriageEngine:
def __init__(self):
self.client = anthropic.Anthropic()
self.runbooks = self._load_runbooks()
self.team_routing = self._load_routing_rules()
def triage(self, context: dict) -> dict:
"""Perform full incident triage: classify, identify root cause, route."""
response = self.client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=4096,
system=f"""You are an expert SRE performing incident triage.
You have access to:
- Alert data and correlated metrics
- Service dependency topology
- Recent deployment history
- Historical incident patterns
Team routing rules:
{json.dumps(self.team_routing, indent=2)}
Your job:
1. Classify severity accurately (SEV1-SEV4)
2. Identify the most likely root cause
3. Determine the correct owning team
4. Suggest immediate mitigation steps
5. Estimate blast radius (users affected)
Be decisive. Wrong fast is better than right slow in incident response.""",
messages=[{
"role": "user",
"content": f"""Triage this incident.
## Correlated Alerts ({len(context['alerts'])} total)
{self._format_alerts(context['alerts'][:20])}
## Affected Services
{json.dumps(context['affectedServices'])}
## Service Topology (dependencies)
{json.dumps(context['serviceTopology'], indent=2)}
## Recent Deployments
{self._format_deployments(context['recentDeployments'])}
## Similar Historical Incidents
{self._format_historical(context['historicalIncidents'][:5])}
## Current Metrics Snapshot
{self._format_metrics(context)}
Respond as JSON:
{{
"severity": "SEV1|SEV2|SEV3|SEV4",
"severity_reasoning": "why this severity level",
"probable_root_cause": "most likely cause",
"root_cause_confidence": 0.0-1.0,
"alternative_causes": ["other possible causes"],
"owning_team": "team name",
"routing_reasoning": "why this team",
"blast_radius": {{
"users_affected_estimate": number,
"services_affected": ["list"],
"revenue_impact_per_minute": number
}},
"immediate_actions": ["ordered list of mitigation steps"],
"relevant_runbook": "runbook name or null",
"escalation_needed": true/false,
"communication_template": "customer-facing status update"
}}"""
}]
)
triage_result = json.loads(response.content[0].text)
# Post-processing: validate routing against team schedules
triage_result = self._validate_routing(triage_result)
return triage_result
def _format_alerts(self, alerts: list) -> str:
return "\n".join([
f"[{a['timestamp']}] [{a['severity']}] {a['service']}: {a['message']}"
for a in alerts
])
def _format_deployments(self, deployments: list) -> str:
if not deployments:
return "No recent deployments in affected services."
return "\n".join([
f"[{d['timestamp']}] {d['service']} - {d['author']}: {d['description']}"
for d in deployments
])
Continuous Learning from Incident Retrospectives
The system improves over time by learning from post-incident reviews.
class TriageLearningLoop:
def __init__(self):
self.client = anthropic.Anthropic()
def process_retrospective(self, incident_id: str, retro_data: dict) -> dict:
"""Extract learnings from incident retrospective to improve future triage."""
response = self.client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=2048,
messages=[{
"role": "user",
"content": f"""Analyze this incident retrospective to improve our triage system.
## Original Triage Decision
{json.dumps(retro_data['original_triage'], indent=2)}
## Actual Outcome
- Real root cause: {retro_data['actual_root_cause']}
- Real severity: {retro_data['actual_severity']}
- Correct team: {retro_data['correct_team']}
- Time to resolve: {retro_data['resolution_time_minutes']} minutes
## What Was Missed
{retro_data.get('missed_signals', 'None documented')}
Extract:
1. What signals should have indicated the correct classification?
2. Should routing rules be updated? How?
3. What new pattern should we recognize in future?
4. Any new runbook needed?
Return actionable improvements as JSON."""
}]
)
improvements = json.loads(response.content[0].text)
self._apply_improvements(improvements)
return improvements
Benchmarks
After 4 months in production handling 127 incidents:
| Metric | Before | After | Improvement |
|---|---|---|---|
| Mean time to engage (MTTE) | 14 min | 47 sec | 95% faster |
| Misrouted incidents | 31% | 5% | 84% reduction |
| Severity misclassification | 18% | 4% | 78% reduction |
| Alert-to-context time | 8-12 min | <5 sec | ~99% faster |
| Incidents escalated unnecessarily | 22% | 8% | 64% reduction |
| Root cause correctly identified | 41% | 73% | +32 points |
Real-World Example
A recent SEV1 incident flow:
- 00:00 — Error rate spike alert fires for
payment-service - 00:03 — 12 additional alerts fire across 4 services
- 00:05 — Correlator groups all 13 alerts, builds context
- 00:05 — Claude triage completes in 4.2 seconds:
- Severity: SEV1 (payment processing affected)
- Root cause: Database connection pool exhaustion (confidence: 0.87)
- Signal: recent deployment added a query without connection release
- Route to: Platform team (database infrastructure owners)
- Immediate action: Roll back deployment
deploy-4521
- 00:06 — PagerDuty pages platform team with full context and suggested action
- 00:08 — On-call engineer confirms and initiates rollback
- 00:11 — Service recovered
Total time from first alert to resolution: 11 minutes. Previous average for similar incidents: 43 minutes.
Cost
- Claude API for triage: ~$0.08 per incident (avg. context size)
- At 30 incidents/month: $2.40/month in API costs
- Infrastructure (correlation engine): $120/month
- Value of 13 minutes saved per incident × $4,200/min downtime cost = $54,600/month saved on SEV1s alone
Limitations
- Novel failure modes: The system is less accurate for failure types not seen in historical data
- Cascading failures: When everything fails simultaneously, root cause identification drops to ~50% accuracy
- Human factors: Can't detect "someone accidentally deleted the database" without deployment/change logs
Conclusion
Incident triage is a classification problem with extremely high stakes and tight time constraints — exactly where AI provides the most value. The system doesn't replace human judgment for resolution, but it compresses the "figure out what's happening" phase from minutes to seconds. The ROI is immediate and measurable: faster engagement means less downtime, and less downtime means less revenue loss. Build the correlation engine first, then add Claude for classification and routing.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.