Sustainable On-Call That Does Not Burn Out Engineers
How to design on-call rotations that maintain reliability without sacrificing engineer well-being—covering compensation, escalation policies, alert hygiene, and rotation structures.

The Problem: On-Call as a Retention Risk
On-call is the tax engineers pay for running production systems. When that tax is too high, your best people leave. In our annual engagement survey, on-call burden was the #1 complaint for three consecutive years. Two senior SREs resigned citing "unsustainable pager load" as their primary reason.
The data was damning: engineers on our primary rotation were paged an average of 11 times per week, with 3.2 pages occurring between midnight and 6 AM. Only 34% of alerts were actionable. The rest were false positives, informational noise, or issues that resolved themselves before anyone could respond.
After a complete overhaul of our on-call system, we reduced pages to 2.1 per week (81% reduction), eliminated non-actionable alerts entirely, and saw on-call satisfaction scores rise from 2.1/5 to 4.3/5.
Diagnosing On-Call Health
Before fixing anything, we needed data. We built a dashboard tracking on-call health metrics:
# scripts/oncall_health_report.py
from dataclasses import dataclass
from datetime import datetime, timedelta
@dataclass
class OnCallHealthMetrics:
period: str
total_pages: int
actionable_pages: int
false_positives: int
auto_resolved: int
mttr_minutes: float
sleep_interruptions: int # Pages between 22:00-07:00
escalations: int
unique_alerts: int # Distinct alert names
toil_hours: float
@property
def actionable_rate(self) -> float:
return self.actionable_pages / self.total_pages if self.total_pages > 0 else 0
@property
def noise_score(self) -> str:
rate = self.actionable_rate
if rate >= 0.9:
return "healthy"
elif rate >= 0.7:
return "noisy"
else:
return "broken"
@property
def burnout_risk(self) -> str:
weekly_sleep_interrupts = self.sleep_interruptions
if weekly_sleep_interrupts <= 1:
return "low"
elif weekly_sleep_interrupts <= 3:
return "moderate"
else:
return "high"
Principle 1: Every Alert Must Be Actionable
The single most impactful change: every alert in the paging rotation must require a human to do something that cannot be automated. We created strict criteria:
An alert is page-worthy if and only if:
- A customer is currently impacted OR will be within 30 minutes
- The issue cannot be resolved automatically
- The on-call engineer can meaningfully reduce impact by responding
Everything else becomes a ticket, a dashboard signal, or gets automated away.
Alert Audit Process
We audited every alert over a 30-day period:
# alert-audit-template.yaml
alert_name: "HighMemoryUsage"
trigger_count_30d: 47
actionable_count: 3
resolution_pattern: "Pod restarted itself within 2 minutes"
recommendation: "REMOVE - convert to auto-restart with ticket if recurring"
---
alert_name: "DatabaseConnectionPoolExhausted"
trigger_count_30d: 8
actionable_count: 8
resolution_pattern: "Required query optimization or connection limit increase"
recommendation: "KEEP - always requires human investigation"
---
alert_name: "CertificateExpiringSoon"
trigger_count_30d: 12
actionable_count: 12
resolution_pattern: "Engineer manually triggered cert renewal"
recommendation: "AUTOMATE - implement cert-manager auto-renewal, alert only on failure"
After the audit, we went from 147 paging alerts to 23. The remaining 23 are genuinely actionable.
Principle 2: Rotation Structure Matters
Our original rotation: one person on-call for 7 days straight. This meant one terrible week per month. We switched to a "follow-the-sun" model with shorter shifts:
# rotation-config.yaml
rotation:
name: platform-oncall
type: follow-the-sun
handoff_overlap: 30m
shifts:
- name: americas
hours: "08:00-18:00 America/New_York"
team_size: 6
rotation_length: 1d # Daily rotation during business hours
- name: emea
hours: "08:00-18:00 Europe/London"
team_size: 4
rotation_length: 1d
- name: overnight
hours: "18:00-08:00 America/New_York"
team_size: 8
rotation_length: 1_night # One night at a time
compensation: 2x_time_off # Next day off guaranteed
escalation:
- level: 1
target: current_oncall
timeout: 10m
- level: 2
target: secondary_oncall
timeout: 15m
- level: 3
target: engineering_manager
timeout: 20m
Key design decisions:
- No one does more than one overnight shift per week. We expanded the overnight pool to 8 engineers.
- Overnight pages always grant the next day off. Non-negotiable—this is policy, not a favor.
- 30-minute handoff overlap ensures context transfer between shifts.
Principle 3: Compensate Fairly
On-call compensation must reflect the actual burden. Our model:
| Component | Compensation |
|---|---|
| Carrying the pager (business hours) | $200/day flat |
| Carrying the pager (overnight/weekend) | $400/day flat |
| Each page responded to | $50 per incident |
| Sleep interruption (22:00-07:00) | Guaranteed next day off + $100 |
| Extended incident (> 2 hours) | Time-and-a-half for duration |
This costs more than "on-call is just part of the job." But it's cheaper than replacing senior engineers who leave due to burnout.
Principle 4: Automate the Repetitive Responses
67% of our actionable alerts had the same resolution steps every time. We automated them:
# automation/auto_remediation.py
from typing import Callable
import subprocess
REMEDIATIONS: dict[str, Callable] = {
"PodCrashLoopBackOff": restart_and_scale,
"DiskSpaceHigh": cleanup_and_expand,
"ConnectionPoolExhausted": kill_idle_connections,
"CertificateExpiring": trigger_renewal,
}
def restart_and_scale(alert_context: dict) -> dict:
"""Restart crashed pod and temporarily scale up."""
namespace = alert_context["labels"]["namespace"]
deployment = alert_context["labels"]["deployment"]
# Restart the problematic pod
subprocess.run([
"kubectl", "rollout", "restart",
f"deployment/{deployment}",
f"--namespace={namespace}"
], check=True)
# Scale up temporarily to maintain capacity
current_replicas = get_replica_count(namespace, deployment)
subprocess.run([
"kubectl", "scale",
f"deployment/{deployment}",
f"--replicas={current_replicas + 1}",
f"--namespace={namespace}"
], check=True)
return {
"action": "restarted_and_scaled",
"notify": "ticket", # Create ticket, don't page
"auto_revert_minutes": 30,
}
Principle 5: Blameless Escalation
Engineers must never feel penalized for escalating. We explicitly track and celebrate escalations:
- Escalating does not count against you
- Managers review unescalated long incidents as potential problems
- Weekly reviews highlight good escalation decisions
- "I don't know how to fix this" is a valid and respected response
Results After Implementation
| Metric | Before | After |
|---|---|---|
| Pages per week | 11.0 | 2.1 |
| Actionable rate | 34% | 98% |
| Sleep interruptions/week | 3.2 | 0.4 |
| Mean time to acknowledge | 8 min | 3 min |
| On-call satisfaction | 2.1/5 | 4.3/5 |
| Attrition citing on-call | 3 eng/year | 0 eng/year |
| Auto-resolved incidents | 0% | 67% |
Key Takeaways
-
Audit every alert ruthlessly. If an alert fired 47 times in 30 days and was only actionable 3 times, it's noise. Remove it or automate the response.
-
Shorter shifts prevent burnout. One night at a time with a guaranteed day off is sustainable. Seven consecutive days is not.
-
Compensate the burden, not just the work. Carrying a pager has a psychological cost even when it doesn't ring. Pay for availability, not just response.
-
Automate repetitive responses. If the same runbook runs more than 3 times a month, it should be a script that pages only on failure.
-
Make escalation safe. Every unescalated long incident is a failure of culture, not of the individual engineer.
The goal of on-call isn't to have someone awake at 3 AM fixing things. It's to build systems that rarely need humans—and to treat those humans well when they do.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.