Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

The Problem: Engineers Trapped in Repetitive Work
Our SRE team of 8 spent 40% of their time on toil—manual, repetitive, automatable work that scaled linearly with service count. As we grew from 30 to 47 services, toil grew proportionally, but headcount didn't. Engineers were drowning in access provisioning, certificate rotations, capacity adjustments, and deploy babysitting.
The breaking point came during Q3 planning. Our VP of Engineering asked: "What would you build if you had 40% more engineering capacity?" The answer was obvious—we already had the capacity. It was being consumed by work that machines should do.
After a focused 6-month toil reduction initiative, we reduced toil from 40% to 12% of engineering time, freeing 2.2 FTE-equivalents for reliability projects. The team shipped 3x more reliability improvements in Q4 than Q3, and no one worked overtime.
Defining Toil Precisely
Toil isn't "work I don't enjoy." Google's SRE book defines it with specific characteristics. Work is toil if it is:
- Manual: A human runs it (not automated)
- Repetitive: Done more than once or twice
- Automatable: A machine could do it
- Tactical: Reactive, interrupt-driven
- Scales linearly: More services/traffic = more work
- No enduring value: Doesn't permanently improve the system
Critically, toil is NOT:
- Overhead (meetings, planning, reviews)
- Engineering work that happens to be boring
- One-time manual tasks (migrations, investigations)
- Creative problem-solving during incidents
Step 1: Measure Current Toil
You can't reduce what you don't measure. We implemented a lightweight time-tracking system focused on categorization, not precision:
# toil-tracker/categories.py
from enum import Enum
from dataclasses import dataclass
from datetime import datetime, timedelta
class WorkCategory(Enum):
TOIL = "toil"
ENGINEERING = "engineering"
OVERHEAD = "overhead"
INCIDENT = "incident"
class ToilSubcategory(Enum):
ACCESS_PROVISIONING = "access_provisioning"
CERTIFICATE_ROTATION = "certificate_rotation"
CAPACITY_ADJUSTMENT = "capacity_adjustment"
DEPLOY_BABYSITTING = "deploy_babysitting"
DATA_EXPORT = "data_export"
CONFIG_CHANGE = "config_change"
SECRET_ROTATION = "secret_rotation"
ENVIRONMENT_SETUP = "environment_setup"
MANUAL_SCALING = "manual_scaling"
LOG_CLEANUP = "log_cleanup"
@dataclass
class ToilEntry:
engineer: str
category: WorkCategory
subcategory: ToilSubcategory | None
duration_minutes: int
service_affected: str
description: str
timestamp: datetime
could_automate: bool = True
automation_difficulty: str = "medium" # low, medium, high
@dataclass
class ToilReport:
period: str
total_engineer_hours: float
toil_hours: float
entries: list[ToilEntry]
@property
def toil_percentage(self) -> float:
return (self.toil_hours / self.total_engineer_hours) * 100
def top_categories(self, n: int = 5) -> list[tuple[str, float]]:
"""Top N toil categories by hours consumed."""
category_hours: dict[str, float] = {}
for entry in self.entries:
if entry.category == WorkCategory.TOIL and entry.subcategory:
key = entry.subcategory.value
category_hours[key] = category_hours.get(key, 0) + entry.duration_minutes / 60
sorted_categories = sorted(
category_hours.items(),
key=lambda x: x[1],
reverse=True
)
return sorted_categories[:n]
Initial Measurement Results
After 4 weeks of tracking:
| Toil Category | Hours/Week | % of Total Toil |
|---|---|---|
| Access provisioning | 12.4 | 24% |
| Deploy babysitting | 9.8 | 19% |
| Certificate rotation | 7.2 | 14% |
| Manual capacity adjustments | 6.5 | 13% |
| Config changes across environments | 5.8 | 11% |
| Data exports/reports | 4.9 | 10% |
| Secret rotation | 3.1 | 6% |
| Other | 1.6 | 3% |
| Total | 51.3 | 100% |
51.3 hours per week / (8 engineers * 32 productive hours) = 40.1% toil
Step 2: Prioritize by Impact and Effort
Not all toil is equally worth automating. We scored each category using an impact/effort matrix:
# toil-prioritization.yaml
automation_candidates:
- category: access_provisioning
weekly_hours: 12.4
automation_effort: medium # 2-3 weeks
risk: low
score: 9.2 # High hours, medium effort = high priority
approach: "Self-service portal with auto-approval for standard roles"
- category: deploy_babysitting
weekly_hours: 9.8
automation_effort: low # 1-2 weeks (canary automation exists)
risk: medium
score: 9.5 # High hours, low effort = highest priority
approach: "Automated canary analysis with auto-promote/rollback"
- category: certificate_rotation
weekly_hours: 7.2
automation_effort: low # cert-manager handles this
risk: low
score: 9.0
approach: "cert-manager with auto-renewal, alert only on failure"
- category: capacity_adjustments
weekly_hours: 6.5
automation_effort: medium
risk: medium
score: 7.8
approach: "Kubernetes VPA + custom HPA policies"
- category: config_changes
weekly_hours: 5.8
automation_effort: high # Requires GitOps overhaul
risk: low
score: 5.4
approach: "GitOps with ArgoCD, PR-based config changes"
Step 3: Automate Top-Priority Toil
Access Provisioning (12.4 → 0.8 hours/week)
Before: Engineers file a ticket, SRE manually runs kubectl commands to grant RBAC roles. Average turnaround: 4 hours.
After: Self-service portal with pre-approved role templates:
# access-automation/role-templates.yaml
templates:
- name: developer-readonly
auto_approve: true
max_duration: 90d
namespaces: ["{{ .Team }}-*"]
rbac:
- apiGroups: [""]
resources: ["pods", "services", "configmaps"]
verbs: ["get", "list", "watch"]
- apiGroups: ["apps"]
resources: ["deployments", "replicasets"]
verbs: ["get", "list", "watch"]
- name: developer-deploy
auto_approve: true
max_duration: 30d
namespaces: ["{{ .Team }}-staging", "{{ .Team }}-production"]
rbac:
- apiGroups: ["apps"]
resources: ["deployments"]
verbs: ["get", "list", "patch", "update"]
- name: admin-access
auto_approve: false # Requires manager approval
max_duration: 8h # Short-lived for break-glass
approval_chain: ["team-lead", "sre-oncall"]
Deploy Babysitting (9.8 → 0.4 hours/week)
Before: Engineer watches deployment metrics for 15 minutes after each deploy, manually rolls back if metrics degrade.
After: Fully automated canary with self-rollback. Human involvement only on automation failure:
#!/bin/bash
# scripts/automated-deploy.sh
set -euo pipefail
SERVICE=$1
VERSION=$2
CANARY_DURATION=600 # 10 minutes
ERROR_THRESHOLD=0.01 # 1% error rate
echo "Deploying $SERVICE:$VERSION with canary analysis"
# Deploy canary (10% traffic)
kubectl set image deployment/$SERVICE \
app=$SERVICE:$VERSION \
--namespace=production
kubectl rollout status deployment/$SERVICE --timeout=120s
# Automated canary analysis
echo "Running canary analysis for ${CANARY_DURATION}s..."
RESULT=$(curl -s "http://canary-analyzer.sre/analyze?service=$SERVICE&duration=$CANARY_DURATION&threshold=$ERROR_THRESHOLD")
STATUS=$(echo $RESULT | jq -r '.verdict')
if [ "$STATUS" = "pass" ]; then
echo "Canary passed. Promoting to full rollout."
kubectl scale deployment/$SERVICE --replicas=$(kubectl get deployment $SERVICE -o jsonpath='{.spec.replicas}')
# Post to Slack
notify_slack "Deployment $SERVICE:$VERSION completed successfully (automated)"
elif [ "$STATUS" = "fail" ]; then
echo "Canary failed. Auto-rolling back."
kubectl rollout undo deployment/$SERVICE
notify_slack "Deployment $SERVICE:$VERSION auto-rolled back. Error rate: $(echo $RESULT | jq -r '.error_rate')"
# Only page if rollback also fails
if ! kubectl rollout status deployment/$SERVICE --timeout=60s; then
page_oncall "Automated rollback failed for $SERVICE"
fi
fi
Certificate Rotation (7.2 → 0 hours/week)
Before: SRE manually tracks cert expiry dates in a spreadsheet, runs renewal scripts monthly.
After: cert-manager handles everything. We only get paged if auto-renewal fails:
# cert-manager/cluster-issuer.yaml
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
server: https://acme-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: letsencrypt-prod-key
solvers:
- dns01:
route53:
region: us-east-1
---
# Alert only on renewal failure
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: cert-alerts
spec:
groups:
- name: certificates
rules:
- alert: CertificateRenewalFailed
expr: certmanager_certificate_ready_status == 0
for: 1h
labels:
severity: warning
annotations:
summary: "Certificate {{ $labels.name }} renewal failed"
Step 4: Track Progress Over Time
We run monthly toil reports and display trends on our team dashboard:
# Prometheus metrics from toil tracker
toil_hours_weekly{team="sre"} /
(team_size{team="sre"} * 32) * 100
# Trend over 6 months
avg_over_time(toil_percentage{team="sre"}[30d])
| Month | Toil % | Automation Shipped | Hours Freed/Week |
|---|---|---|---|
| Month 1 | 40% | Deploy automation | 9.4 |
| Month 2 | 32% | cert-manager + access portal | 18.6 |
| Month 3 | 24% | Capacity auto-scaling | 5.2 |
| Month 4 | 19% | Config GitOps | 5.8 |
| Month 5 | 15% | Data export self-service | 4.9 |
| Month 6 | 12% | Secret rotation automation | 3.1 |
The 50% Rule
Google's SRE book recommends that SRE teams spend no more than 50% of their time on toil. We target 15% as our steady-state maximum. When toil exceeds 15%, we pause project work and focus exclusively on automation until it drops below threshold.
This sounds extreme, but the math is clear: reducing toil from 40% to 12% freed 28 percentage points—equivalent to 2.2 full-time engineers doing creative work instead of repetitive tasks.
Key Takeaways
-
Measure toil weekly, not annually. Monthly tracking reveals trends and creates accountability. Annual surveys are too imprecise to drive change.
-
Prioritize by hours * (1/effort). The highest-impact automation targets are high-frequency, low-effort categories. Our deploy babysitting automation took 1 week and saved 9.8 hours/week.
-
Self-service beats automation-for-others. The access provisioning portal doesn't just save SRE time—it saves the requesting engineer's wait time too.
-
Set a toil budget like an error budget. When toil exceeds your threshold, stop feature work and automate. This makes toil reduction a team-level commitment, not a backlog item that never gets prioritized.
-
Celebrate the freed capacity. Track what your team ships with the freed time. "We reduced toil by 28 points" is abstract. "We shipped VPA, chaos testing, and cross-region failover with the freed time" is concrete.
Toil is a tax on your team's creativity. Every hour spent on repetitive work is an hour not spent making the system genuinely better.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Synthetic Monitors From 12 Regions Catching Issues Before Users
How to implement global synthetic monitoring that detects availability and performance degradation from every major region—catching issues minutes before real users are impacted.

Comments
No comments yet. Be the first to share your thoughts.