Reducing Alert Volume by 83% While Catching More Incidents

How we eliminated alert fatigue by restructuring our monitoring from symptom-based to SLO-based alerting, reducing pages from 147/week to 25 while improving detection

#alerting#observability#sre#on-call
Cover image for the article: Reducing Alert Volume by 83% While Catching More Incidents

Our on-call engineers received 147 alerts per week. They acknowledged most within seconds and immediately closed them as non-actionable. The signal-to-noise ratio was so poor that when a real incident occurred at 2:14 AM, the on-call engineer dismissed the first three alerts as false positives before realizing the fourth was genuine. That 8-minute delay turned a minor blip into a customer-facing outage. We restructured our entire alerting philosophy from resource metrics to SLO-based alerting, reducing volume by 83% while actually catching incidents faster.

The Problem: Alert Fatigue Is a Safety Issue

Alert fatigue is not an inconvenience. It is a safety issue. When engineers learn to ignore alerts, they ignore real ones too. Our post-incident review revealed a pattern: the on-call engineer had received 23 alerts in the previous 4 hours, all false positives. By the time the real alert fired, their conditioned response was to dismiss.

Before alerting overhaul:

  • 147 alerts per week across 4 on-call rotations
  • 89% of alerts required no action (false positive or auto-resolved)
  • Mean acknowledgment time increasing monthly (fatigue signal)
  • 3 incidents in 6 months where real alerts were initially dismissed
  • On-call satisfaction score: 2.1/5

Root Cause: Resource Alerts vs. Symptom Alerts

Analysis of our 147 weekly alerts revealed the distribution:

Alert CategoryCount/WeekActionableRoot Cause
CPU > 80%342%Normal autoscaling behavior
Memory > 75%285%Java heap patterns
Disk > 70%198%Log rotation timing
Pod restarts > 02212%Spot instance reclamation
Error rate > 0.1%1831%Too sensitive threshold
Latency spike1445%Mixed - some meaningful
Health check failure1215%Transient network blips

The pattern was clear: we were alerting on causes (CPU high) rather than symptoms (users affected). CPU at 85% is not an incident if users experience normal latency and error rates.

The SLO-Based Alerting Model

We replaced resource-based alerts with SLO (Service Level Objective) burn-rate alerts. Instead of asking "is a resource metric abnormal?", we ask "are we burning through our error budget faster than sustainable?"

SLO-Based Alerting Architecture

The core concept: if our SLO is 99.9% availability (43.8 minutes of allowed downtime per month), we alert when we are consuming that budget faster than the month allows.

Implementing Multi-Window Burn Rate Alerts

We implemented the Google SRE burn-rate alerting model with multiple windows to catch both fast burns (outages) and slow burns (degradation):

# prometheus-rules/slo-alerts.yaml
groups:
  - name: slo-burn-rate-alerts
    rules:
      # Fast burn: consuming 14.4x budget (will exhaust in 5 hours)
      # Window: 5m rate over 1h lookback
      - alert: SLOBurnRateCritical
        expr: |
          (
            sum(rate(http_requests_total{status=~"5.."}[5m]))
            /
            sum(rate(http_requests_total[5m]))
          ) > (14.4 * 0.001)
          and
          (
            sum(rate(http_requests_total{status=~"5.."}[1h]))
            /
            sum(rate(http_requests_total[1h]))
          ) > (14.4 * 0.001)
        for: 2m
        labels:
          severity: critical
          slo: availability
          burn_rate: "14.4x"
        annotations:
          summary: "SLO burn rate critical - error budget exhaustion in ~5 hours"
          description: |
            Service {{ $labels.service }} is burning error budget at 14.4x rate.
            Current error rate: {{ $value | humanizePercentage }}
            At this rate, monthly budget exhausts in ~5 hours.
            Action required immediately.
          runbook_url: "https://runbooks.company.com/slo-burn-rate"

      # Medium burn: consuming 6x budget (will exhaust in 12 hours)
      # Window: 30m rate over 6h lookback
      - alert: SLOBurnRateHigh
        expr: |
          (
            sum(rate(http_requests_total{status=~"5.."}[30m]))
            /
            sum(rate(http_requests_total[30m]))
          ) > (6 * 0.001)
          and
          (
            sum(rate(http_requests_total{status=~"5.."}[6h]))
            /
            sum(rate(http_requests_total[6h]))
          ) > (6 * 0.001)
        for: 5m
        labels:
          severity: warning
          slo: availability
          burn_rate: "6x"
        annotations:
          summary: "SLO burn rate elevated - error budget exhaustion in ~12 hours"

      # Slow burn: consuming 3x budget (will exhaust in 2.4 days)
      # Window: 2h rate over 24h lookback
      - alert: SLOBurnRateSlow
        expr: |
          (
            sum(rate(http_requests_total{status=~"5.."}[2h]))
            /
            sum(rate(http_requests_total[2h]))
          ) > (3 * 0.001)
          and
          (
            sum(rate(http_requests_total{status=~"5.."}[1d]))
            /
            sum(rate(http_requests_total[1d]))
          ) > (3 * 0.001)
        for: 15m
        labels:
          severity: warning
          slo: availability
          burn_rate: "3x"
          notification: slack_only

The multi-window approach eliminates both false positives (brief spikes that self-resolve) and false negatives (slow degradation that burns budget without dramatic spikes).

Alert Routing and Escalation

Not all alerts deserve a page. We implemented tiered routing:

from dataclasses import dataclass
from enum import Enum

class AlertAction(Enum):
    PAGE = "page"               # PagerDuty, wake someone up
    SLACK_URGENT = "slack"      # Slack channel, respond within 15 min
    TICKET = "ticket"           # Create JIRA ticket, respond within 24h
    DASHBOARD = "dashboard"     # Log to dashboard only

@dataclass
class AlertRoute:
    burn_rate: float
    time_to_exhaustion: str
    action: AlertAction
    response_time: str

ROUTING_TABLE = [
    AlertRoute(burn_rate=14.4, time_to_exhaustion="5 hours",  action=AlertAction.PAGE, response_time="5 minutes"),
    AlertRoute(burn_rate=6.0,  time_to_exhaustion="12 hours", action=AlertAction.SLACK_URGENT, response_time="15 minutes"),
    AlertRoute(burn_rate=3.0,  time_to_exhaustion="2.4 days", action=AlertAction.TICKET, response_time="4 hours"),
    AlertRoute(burn_rate=1.0,  time_to_exhaustion="30 days",  action=AlertAction.DASHBOARD, response_time="N/A"),
]

def route_alert(burn_rate: float, service: str, slo_type: str) -> AlertAction:
    """Determine alert routing based on burn rate severity."""
    for route in ROUTING_TABLE:
        if burn_rate >= route.burn_rate:
            return route.action
    
    return AlertAction.DASHBOARD

Only 14.4x burn rate (budget exhaustion within 5 hours) pages the on-call engineer. Everything else goes to Slack or ticketing systems for business-hours response.

Alert Correlation and Deduplication

Multiple SLO alerts often fire simultaneously during an incident. We implemented correlation logic to group related alerts into a single notification:

interface AlertGroup {
  incident_id: string;
  primary_alert: Alert;
  correlated_alerts: Alert[];
  blast_radius: string[];
  suggested_runbook: string;
}

function correlateAlerts(alerts: Alert[], timeWindow: number = 300): AlertGroup[] {
  const groups: AlertGroup[] = [];
  const processed = new Set<string>();

  for (const alert of alerts) {
    if (processed.has(alert.id)) continue;

    // Find alerts that fired within the time window and share dependencies
    const correlated = alerts.filter(a => 
      !processed.has(a.id) &&
      a.id !== alert.id &&
      Math.abs(a.fired_at - alert.fired_at) < timeWindow &&
      hasSharedDependency(alert.service, a.service)
    );

    const group: AlertGroup = {
      incident_id: generateIncidentId(),
      primary_alert: alert,  // Highest severity or earliest
      correlated_alerts: correlated,
      blast_radius: [alert.service, ...correlated.map(a => a.service)],
      suggested_runbook: selectRunbook(alert, correlated),
    };

    groups.push(group);
    processed.add(alert.id);
    correlated.forEach(a => processed.add(a.id));
  }

  return groups;
}

During a database outage that previously generated 14 separate alerts (one per dependent service), the engineer now receives a single grouped notification: "Database connectivity SLO violation affecting 14 services" with a link to the database runbook.

Error Budget Dashboard

Visibility into error budget consumption replaced gut-feel decisions about when to page:

Error Budget Dashboard

The dashboard shows:

  • Remaining error budget for the month (in minutes and percentage)
  • Current burn rate trajectory
  • Historical budget consumption patterns
  • Projected budget exhaustion date at current rate

Teams that can see their error budget make better decisions about risk tolerance without needing prescriptive alerting rules.

Results After 3 Months

MetricBeforeAfterChange
Alerts per week14725-83%
Actionable alert percentage11%78%+609%
Mean time to acknowledge (real incidents)8.2 min1.4 min-83%
Incidents initially dismissed3/6 months0/6 months-100%
On-call satisfaction score2.1/54.3/5+105%
Pages per on-call shift5.20.9-83%
Incidents detected by alerting (vs customer report)72%94%+31%

The most significant result: we catch more incidents (94% vs 72% detected before customer report) with 83% fewer alerts. This is not a tradeoff. Better signal-to-noise ratio means engineers trust and respond to alerts immediately.

Implementation Mistakes We Made

Mistake 1: Removing resource alerts entirely. Some resource alerts (disk at 95%) are genuinely useful as early warnings. We kept a small set as Slack-only notifications, not pages.

Mistake 2: Setting burn rates too sensitively initially. Our first SLO thresholds were too tight (99.99% for non-critical services). This generated excessive alerts. We calibrated to realistic SLOs (99.9% for most services, 99.95% for critical paths).

Mistake 3: Not accounting for deploy-related error spikes. Deploys cause brief error spikes. We added deploy annotations to suppress alerts during rolling deployments.

On-Call Quality of Life

The qualitative impact was as significant as the quantitative. Engineers reported:

  • Sleeping through on-call shifts without interruption (previously rare)
  • Trusting that a page means something genuinely needs attention
  • Less burnout and longer tenure in on-call rotations
  • Volunteering for on-call rather than avoiding it

Conclusion

Alert fatigue is solved by changing what you alert on, not by adding suppression rules to noisy alerts. SLO-based burn-rate alerting asks the right question: "are users being impacted?" instead of "is a metric abnormal?" The 83% volume reduction came from eliminating alerts that never required action. The 31% improvement in incident detection came from engineers who trust and immediately respond to the alerts that remain. Start by classifying your current alerts into actionable vs noise, define SLOs for each service, implement multi-window burn rates, and route by severity. Your on-call engineers will thank you.

Comments

    No comments yet. Be the first to share your thoughts.