Sustainable On-Call That Does Not Burn Out Engineers

How to design on-call rotations that maintain reliability without sacrificing engineer well-being—covering compensation, escalation policies, alert hygiene, and rotation structures.

#on-call#sre#burnout#engineering-management
Cover image for the article: Sustainable On-Call That Does Not Burn Out Engineers

The Problem: On-Call as a Retention Risk

On-call is the tax engineers pay for running production systems. When that tax is too high, your best people leave. In our annual engagement survey, on-call burden was the #1 complaint for three consecutive years. Two senior SREs resigned citing "unsustainable pager load" as their primary reason.

The data was damning: engineers on our primary rotation were paged an average of 11 times per week, with 3.2 pages occurring between midnight and 6 AM. Only 34% of alerts were actionable. The rest were false positives, informational noise, or issues that resolved themselves before anyone could respond.

After a complete overhaul of our on-call system, we reduced pages to 2.1 per week (81% reduction), eliminated non-actionable alerts entirely, and saw on-call satisfaction scores rise from 2.1/5 to 4.3/5.

Diagnosing On-Call Health

Before fixing anything, we needed data. We built a dashboard tracking on-call health metrics:

# scripts/oncall_health_report.py
from dataclasses import dataclass
from datetime import datetime, timedelta

@dataclass
class OnCallHealthMetrics:
    period: str
    total_pages: int
    actionable_pages: int
    false_positives: int
    auto_resolved: int
    mttr_minutes: float
    sleep_interruptions: int  # Pages between 22:00-07:00
    escalations: int
    unique_alerts: int  # Distinct alert names
    toil_hours: float

    @property
    def actionable_rate(self) -> float:
        return self.actionable_pages / self.total_pages if self.total_pages > 0 else 0

    @property
    def noise_score(self) -> str:
        rate = self.actionable_rate
        if rate >= 0.9:
            return "healthy"
        elif rate >= 0.7:
            return "noisy"
        else:
            return "broken"

    @property
    def burnout_risk(self) -> str:
        weekly_sleep_interrupts = self.sleep_interruptions
        if weekly_sleep_interrupts <= 1:
            return "low"
        elif weekly_sleep_interrupts <= 3:
            return "moderate"
        else:
            return "high"

On-Call Health Dashboard

Principle 1: Every Alert Must Be Actionable

The single most impactful change: every alert in the paging rotation must require a human to do something that cannot be automated. We created strict criteria:

An alert is page-worthy if and only if:

  • A customer is currently impacted OR will be within 30 minutes
  • The issue cannot be resolved automatically
  • The on-call engineer can meaningfully reduce impact by responding

Everything else becomes a ticket, a dashboard signal, or gets automated away.

Alert Audit Process

We audited every alert over a 30-day period:

# alert-audit-template.yaml
alert_name: "HighMemoryUsage"
trigger_count_30d: 47
actionable_count: 3
resolution_pattern: "Pod restarted itself within 2 minutes"
recommendation: "REMOVE - convert to auto-restart with ticket if recurring"
---
alert_name: "DatabaseConnectionPoolExhausted"
trigger_count_30d: 8
actionable_count: 8
resolution_pattern: "Required query optimization or connection limit increase"
recommendation: "KEEP - always requires human investigation"
---
alert_name: "CertificateExpiringSoon"
trigger_count_30d: 12
actionable_count: 12
resolution_pattern: "Engineer manually triggered cert renewal"
recommendation: "AUTOMATE - implement cert-manager auto-renewal, alert only on failure"

After the audit, we went from 147 paging alerts to 23. The remaining 23 are genuinely actionable.

Principle 2: Rotation Structure Matters

Our original rotation: one person on-call for 7 days straight. This meant one terrible week per month. We switched to a "follow-the-sun" model with shorter shifts:

# rotation-config.yaml
rotation:
  name: platform-oncall
  type: follow-the-sun
  handoff_overlap: 30m
  
  shifts:
    - name: americas
      hours: "08:00-18:00 America/New_York"
      team_size: 6
      rotation_length: 1d  # Daily rotation during business hours
      
    - name: emea
      hours: "08:00-18:00 Europe/London"  
      team_size: 4
      rotation_length: 1d
      
    - name: overnight
      hours: "18:00-08:00 America/New_York"
      team_size: 8
      rotation_length: 1_night  # One night at a time
      compensation: 2x_time_off  # Next day off guaranteed
      
  escalation:
    - level: 1
      target: current_oncall
      timeout: 10m
    - level: 2
      target: secondary_oncall
      timeout: 15m
    - level: 3
      target: engineering_manager
      timeout: 20m

Key design decisions:

  • No one does more than one overnight shift per week. We expanded the overnight pool to 8 engineers.
  • Overnight pages always grant the next day off. Non-negotiable—this is policy, not a favor.
  • 30-minute handoff overlap ensures context transfer between shifts.

Principle 3: Compensate Fairly

On-call compensation must reflect the actual burden. Our model:

ComponentCompensation
Carrying the pager (business hours)$200/day flat
Carrying the pager (overnight/weekend)$400/day flat
Each page responded to$50 per incident
Sleep interruption (22:00-07:00)Guaranteed next day off + $100
Extended incident (> 2 hours)Time-and-a-half for duration

This costs more than "on-call is just part of the job." But it's cheaper than replacing senior engineers who leave due to burnout.

Principle 4: Automate the Repetitive Responses

67% of our actionable alerts had the same resolution steps every time. We automated them:

# automation/auto_remediation.py
from typing import Callable
import subprocess

REMEDIATIONS: dict[str, Callable] = {
    "PodCrashLoopBackOff": restart_and_scale,
    "DiskSpaceHigh": cleanup_and_expand,
    "ConnectionPoolExhausted": kill_idle_connections,
    "CertificateExpiring": trigger_renewal,
}

def restart_and_scale(alert_context: dict) -> dict:
    """Restart crashed pod and temporarily scale up."""
    namespace = alert_context["labels"]["namespace"]
    deployment = alert_context["labels"]["deployment"]
    
    # Restart the problematic pod
    subprocess.run([
        "kubectl", "rollout", "restart",
        f"deployment/{deployment}",
        f"--namespace={namespace}"
    ], check=True)
    
    # Scale up temporarily to maintain capacity
    current_replicas = get_replica_count(namespace, deployment)
    subprocess.run([
        "kubectl", "scale",
        f"deployment/{deployment}",
        f"--replicas={current_replicas + 1}",
        f"--namespace={namespace}"
    ], check=True)
    
    return {
        "action": "restarted_and_scaled",
        "notify": "ticket",  # Create ticket, don't page
        "auto_revert_minutes": 30,
    }

Principle 5: Blameless Escalation

Engineers must never feel penalized for escalating. We explicitly track and celebrate escalations:

  • Escalating does not count against you
  • Managers review unescalated long incidents as potential problems
  • Weekly reviews highlight good escalation decisions
  • "I don't know how to fix this" is a valid and respected response

Results After Implementation

MetricBeforeAfter
Pages per week11.02.1
Actionable rate34%98%
Sleep interruptions/week3.20.4
Mean time to acknowledge8 min3 min
On-call satisfaction2.1/54.3/5
Attrition citing on-call3 eng/year0 eng/year
Auto-resolved incidents0%67%

On-Call Metrics Improvement Timeline

Key Takeaways

  1. Audit every alert ruthlessly. If an alert fired 47 times in 30 days and was only actionable 3 times, it's noise. Remove it or automate the response.

  2. Shorter shifts prevent burnout. One night at a time with a guaranteed day off is sustainable. Seven consecutive days is not.

  3. Compensate the burden, not just the work. Carrying a pager has a psychological cost even when it doesn't ring. Pay for availability, not just response.

  4. Automate repetitive responses. If the same runbook runs more than 3 times a month, it should be a script that pages only on failure.

  5. Make escalation safe. Every unescalated long incident is a failure of culture, not of the individual engineer.

The goal of on-call isn't to have someone awake at 3 AM fixing things. It's to build systems that rarely need humans—and to treat those humans well when they do.

Comments

    No comments yet. Be the first to share your thoughts.