Reducing Alert Volume by 83% While Catching More Incidents
How we eliminated alert fatigue by restructuring our monitoring from symptom-based to SLO-based alerting, reducing pages from 147/week to 25 while improving detection

Our on-call engineers received 147 alerts per week. They acknowledged most within seconds and immediately closed them as non-actionable. The signal-to-noise ratio was so poor that when a real incident occurred at 2:14 AM, the on-call engineer dismissed the first three alerts as false positives before realizing the fourth was genuine. That 8-minute delay turned a minor blip into a customer-facing outage. We restructured our entire alerting philosophy from resource metrics to SLO-based alerting, reducing volume by 83% while actually catching incidents faster.
The Problem: Alert Fatigue Is a Safety Issue
Alert fatigue is not an inconvenience. It is a safety issue. When engineers learn to ignore alerts, they ignore real ones too. Our post-incident review revealed a pattern: the on-call engineer had received 23 alerts in the previous 4 hours, all false positives. By the time the real alert fired, their conditioned response was to dismiss.
Before alerting overhaul:
- 147 alerts per week across 4 on-call rotations
- 89% of alerts required no action (false positive or auto-resolved)
- Mean acknowledgment time increasing monthly (fatigue signal)
- 3 incidents in 6 months where real alerts were initially dismissed
- On-call satisfaction score: 2.1/5
Root Cause: Resource Alerts vs. Symptom Alerts
Analysis of our 147 weekly alerts revealed the distribution:
| Alert Category | Count/Week | Actionable | Root Cause |
|---|---|---|---|
| CPU > 80% | 34 | 2% | Normal autoscaling behavior |
| Memory > 75% | 28 | 5% | Java heap patterns |
| Disk > 70% | 19 | 8% | Log rotation timing |
| Pod restarts > 0 | 22 | 12% | Spot instance reclamation |
| Error rate > 0.1% | 18 | 31% | Too sensitive threshold |
| Latency spike | 14 | 45% | Mixed - some meaningful |
| Health check failure | 12 | 15% | Transient network blips |
The pattern was clear: we were alerting on causes (CPU high) rather than symptoms (users affected). CPU at 85% is not an incident if users experience normal latency and error rates.
The SLO-Based Alerting Model
We replaced resource-based alerts with SLO (Service Level Objective) burn-rate alerts. Instead of asking "is a resource metric abnormal?", we ask "are we burning through our error budget faster than sustainable?"
The core concept: if our SLO is 99.9% availability (43.8 minutes of allowed downtime per month), we alert when we are consuming that budget faster than the month allows.
Implementing Multi-Window Burn Rate Alerts
We implemented the Google SRE burn-rate alerting model with multiple windows to catch both fast burns (outages) and slow burns (degradation):
# prometheus-rules/slo-alerts.yaml
groups:
- name: slo-burn-rate-alerts
rules:
# Fast burn: consuming 14.4x budget (will exhaust in 5 hours)
# Window: 5m rate over 1h lookback
- alert: SLOBurnRateCritical
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
) > (14.4 * 0.001)
and
(
sum(rate(http_requests_total{status=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
) > (14.4 * 0.001)
for: 2m
labels:
severity: critical
slo: availability
burn_rate: "14.4x"
annotations:
summary: "SLO burn rate critical - error budget exhaustion in ~5 hours"
description: |
Service {{ $labels.service }} is burning error budget at 14.4x rate.
Current error rate: {{ $value | humanizePercentage }}
At this rate, monthly budget exhausts in ~5 hours.
Action required immediately.
runbook_url: "https://runbooks.company.com/slo-burn-rate"
# Medium burn: consuming 6x budget (will exhaust in 12 hours)
# Window: 30m rate over 6h lookback
- alert: SLOBurnRateHigh
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[30m]))
/
sum(rate(http_requests_total[30m]))
) > (6 * 0.001)
and
(
sum(rate(http_requests_total{status=~"5.."}[6h]))
/
sum(rate(http_requests_total[6h]))
) > (6 * 0.001)
for: 5m
labels:
severity: warning
slo: availability
burn_rate: "6x"
annotations:
summary: "SLO burn rate elevated - error budget exhaustion in ~12 hours"
# Slow burn: consuming 3x budget (will exhaust in 2.4 days)
# Window: 2h rate over 24h lookback
- alert: SLOBurnRateSlow
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[2h]))
/
sum(rate(http_requests_total[2h]))
) > (3 * 0.001)
and
(
sum(rate(http_requests_total{status=~"5.."}[1d]))
/
sum(rate(http_requests_total[1d]))
) > (3 * 0.001)
for: 15m
labels:
severity: warning
slo: availability
burn_rate: "3x"
notification: slack_only
The multi-window approach eliminates both false positives (brief spikes that self-resolve) and false negatives (slow degradation that burns budget without dramatic spikes).
Alert Routing and Escalation
Not all alerts deserve a page. We implemented tiered routing:
from dataclasses import dataclass
from enum import Enum
class AlertAction(Enum):
PAGE = "page" # PagerDuty, wake someone up
SLACK_URGENT = "slack" # Slack channel, respond within 15 min
TICKET = "ticket" # Create JIRA ticket, respond within 24h
DASHBOARD = "dashboard" # Log to dashboard only
@dataclass
class AlertRoute:
burn_rate: float
time_to_exhaustion: str
action: AlertAction
response_time: str
ROUTING_TABLE = [
AlertRoute(burn_rate=14.4, time_to_exhaustion="5 hours", action=AlertAction.PAGE, response_time="5 minutes"),
AlertRoute(burn_rate=6.0, time_to_exhaustion="12 hours", action=AlertAction.SLACK_URGENT, response_time="15 minutes"),
AlertRoute(burn_rate=3.0, time_to_exhaustion="2.4 days", action=AlertAction.TICKET, response_time="4 hours"),
AlertRoute(burn_rate=1.0, time_to_exhaustion="30 days", action=AlertAction.DASHBOARD, response_time="N/A"),
]
def route_alert(burn_rate: float, service: str, slo_type: str) -> AlertAction:
"""Determine alert routing based on burn rate severity."""
for route in ROUTING_TABLE:
if burn_rate >= route.burn_rate:
return route.action
return AlertAction.DASHBOARD
Only 14.4x burn rate (budget exhaustion within 5 hours) pages the on-call engineer. Everything else goes to Slack or ticketing systems for business-hours response.
Alert Correlation and Deduplication
Multiple SLO alerts often fire simultaneously during an incident. We implemented correlation logic to group related alerts into a single notification:
interface AlertGroup {
incident_id: string;
primary_alert: Alert;
correlated_alerts: Alert[];
blast_radius: string[];
suggested_runbook: string;
}
function correlateAlerts(alerts: Alert[], timeWindow: number = 300): AlertGroup[] {
const groups: AlertGroup[] = [];
const processed = new Set<string>();
for (const alert of alerts) {
if (processed.has(alert.id)) continue;
// Find alerts that fired within the time window and share dependencies
const correlated = alerts.filter(a =>
!processed.has(a.id) &&
a.id !== alert.id &&
Math.abs(a.fired_at - alert.fired_at) < timeWindow &&
hasSharedDependency(alert.service, a.service)
);
const group: AlertGroup = {
incident_id: generateIncidentId(),
primary_alert: alert, // Highest severity or earliest
correlated_alerts: correlated,
blast_radius: [alert.service, ...correlated.map(a => a.service)],
suggested_runbook: selectRunbook(alert, correlated),
};
groups.push(group);
processed.add(alert.id);
correlated.forEach(a => processed.add(a.id));
}
return groups;
}
During a database outage that previously generated 14 separate alerts (one per dependent service), the engineer now receives a single grouped notification: "Database connectivity SLO violation affecting 14 services" with a link to the database runbook.
Error Budget Dashboard
Visibility into error budget consumption replaced gut-feel decisions about when to page:
The dashboard shows:
- Remaining error budget for the month (in minutes and percentage)
- Current burn rate trajectory
- Historical budget consumption patterns
- Projected budget exhaustion date at current rate
Teams that can see their error budget make better decisions about risk tolerance without needing prescriptive alerting rules.
Results After 3 Months
| Metric | Before | After | Change |
|---|---|---|---|
| Alerts per week | 147 | 25 | -83% |
| Actionable alert percentage | 11% | 78% | +609% |
| Mean time to acknowledge (real incidents) | 8.2 min | 1.4 min | -83% |
| Incidents initially dismissed | 3/6 months | 0/6 months | -100% |
| On-call satisfaction score | 2.1/5 | 4.3/5 | +105% |
| Pages per on-call shift | 5.2 | 0.9 | -83% |
| Incidents detected by alerting (vs customer report) | 72% | 94% | +31% |
The most significant result: we catch more incidents (94% vs 72% detected before customer report) with 83% fewer alerts. This is not a tradeoff. Better signal-to-noise ratio means engineers trust and respond to alerts immediately.
Implementation Mistakes We Made
Mistake 1: Removing resource alerts entirely. Some resource alerts (disk at 95%) are genuinely useful as early warnings. We kept a small set as Slack-only notifications, not pages.
Mistake 2: Setting burn rates too sensitively initially. Our first SLO thresholds were too tight (99.99% for non-critical services). This generated excessive alerts. We calibrated to realistic SLOs (99.9% for most services, 99.95% for critical paths).
Mistake 3: Not accounting for deploy-related error spikes. Deploys cause brief error spikes. We added deploy annotations to suppress alerts during rolling deployments.
On-Call Quality of Life
The qualitative impact was as significant as the quantitative. Engineers reported:
- Sleeping through on-call shifts without interruption (previously rare)
- Trusting that a page means something genuinely needs attention
- Less burnout and longer tenure in on-call rotations
- Volunteering for on-call rather than avoiding it
Conclusion
Alert fatigue is solved by changing what you alert on, not by adding suppression rules to noisy alerts. SLO-based burn-rate alerting asks the right question: "are users being impacted?" instead of "is a metric abnormal?" The 83% volume reduction came from eliminating alerts that never required action. The 31% improvement in incident detection came from engineers who trust and immediately respond to the alerts that remain. Start by classifying your current alerts into actionable vs noise, define SLOs for each service, implement multi-window burn rates, and route by severity. Your on-call engineers will thank you.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.