Measuring and Eliminating Toil: From 40% to 12% of Engineering Time

A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

#toil#automation#sre#productivity
Cover image for the article: Measuring and Eliminating Toil: From 40% to 12% of Engineering Time

The Problem: Engineers Trapped in Repetitive Work

Our SRE team of 8 spent 40% of their time on toil—manual, repetitive, automatable work that scaled linearly with service count. As we grew from 30 to 47 services, toil grew proportionally, but headcount didn't. Engineers were drowning in access provisioning, certificate rotations, capacity adjustments, and deploy babysitting.

The breaking point came during Q3 planning. Our VP of Engineering asked: "What would you build if you had 40% more engineering capacity?" The answer was obvious—we already had the capacity. It was being consumed by work that machines should do.

After a focused 6-month toil reduction initiative, we reduced toil from 40% to 12% of engineering time, freeing 2.2 FTE-equivalents for reliability projects. The team shipped 3x more reliability improvements in Q4 than Q3, and no one worked overtime.

Defining Toil Precisely

Toil isn't "work I don't enjoy." Google's SRE book defines it with specific characteristics. Work is toil if it is:

  • Manual: A human runs it (not automated)
  • Repetitive: Done more than once or twice
  • Automatable: A machine could do it
  • Tactical: Reactive, interrupt-driven
  • Scales linearly: More services/traffic = more work
  • No enduring value: Doesn't permanently improve the system

Toil Identification Decision Tree

Critically, toil is NOT:

  • Overhead (meetings, planning, reviews)
  • Engineering work that happens to be boring
  • One-time manual tasks (migrations, investigations)
  • Creative problem-solving during incidents

Step 1: Measure Current Toil

You can't reduce what you don't measure. We implemented a lightweight time-tracking system focused on categorization, not precision:

# toil-tracker/categories.py
from enum import Enum
from dataclasses import dataclass
from datetime import datetime, timedelta

class WorkCategory(Enum):
    TOIL = "toil"
    ENGINEERING = "engineering"
    OVERHEAD = "overhead"
    INCIDENT = "incident"

class ToilSubcategory(Enum):
    ACCESS_PROVISIONING = "access_provisioning"
    CERTIFICATE_ROTATION = "certificate_rotation"
    CAPACITY_ADJUSTMENT = "capacity_adjustment"
    DEPLOY_BABYSITTING = "deploy_babysitting"
    DATA_EXPORT = "data_export"
    CONFIG_CHANGE = "config_change"
    SECRET_ROTATION = "secret_rotation"
    ENVIRONMENT_SETUP = "environment_setup"
    MANUAL_SCALING = "manual_scaling"
    LOG_CLEANUP = "log_cleanup"
    
@dataclass
class ToilEntry:
    engineer: str
    category: WorkCategory
    subcategory: ToilSubcategory | None
    duration_minutes: int
    service_affected: str
    description: str
    timestamp: datetime
    could_automate: bool = True
    automation_difficulty: str = "medium"  # low, medium, high

@dataclass
class ToilReport:
    period: str
    total_engineer_hours: float
    toil_hours: float
    entries: list[ToilEntry]
    
    @property
    def toil_percentage(self) -> float:
        return (self.toil_hours / self.total_engineer_hours) * 100
    
    def top_categories(self, n: int = 5) -> list[tuple[str, float]]:
        """Top N toil categories by hours consumed."""
        category_hours: dict[str, float] = {}
        for entry in self.entries:
            if entry.category == WorkCategory.TOIL and entry.subcategory:
                key = entry.subcategory.value
                category_hours[key] = category_hours.get(key, 0) + entry.duration_minutes / 60
        
        sorted_categories = sorted(
            category_hours.items(), 
            key=lambda x: x[1], 
            reverse=True
        )
        return sorted_categories[:n]

Initial Measurement Results

After 4 weeks of tracking:

Toil CategoryHours/Week% of Total Toil
Access provisioning12.424%
Deploy babysitting9.819%
Certificate rotation7.214%
Manual capacity adjustments6.513%
Config changes across environments5.811%
Data exports/reports4.910%
Secret rotation3.16%
Other1.63%
Total51.3100%

51.3 hours per week / (8 engineers * 32 productive hours) = 40.1% toil

Step 2: Prioritize by Impact and Effort

Not all toil is equally worth automating. We scored each category using an impact/effort matrix:

# toil-prioritization.yaml
automation_candidates:
  - category: access_provisioning
    weekly_hours: 12.4
    automation_effort: medium  # 2-3 weeks
    risk: low
    score: 9.2  # High hours, medium effort = high priority
    approach: "Self-service portal with auto-approval for standard roles"
    
  - category: deploy_babysitting
    weekly_hours: 9.8
    automation_effort: low  # 1-2 weeks (canary automation exists)
    risk: medium
    score: 9.5  # High hours, low effort = highest priority
    approach: "Automated canary analysis with auto-promote/rollback"
    
  - category: certificate_rotation
    weekly_hours: 7.2
    automation_effort: low  # cert-manager handles this
    risk: low
    score: 9.0
    approach: "cert-manager with auto-renewal, alert only on failure"
    
  - category: capacity_adjustments
    weekly_hours: 6.5
    automation_effort: medium
    risk: medium
    score: 7.8
    approach: "Kubernetes VPA + custom HPA policies"
    
  - category: config_changes
    weekly_hours: 5.8
    automation_effort: high  # Requires GitOps overhaul
    risk: low
    score: 5.4
    approach: "GitOps with ArgoCD, PR-based config changes"

Step 3: Automate Top-Priority Toil

Access Provisioning (12.4 → 0.8 hours/week)

Before: Engineers file a ticket, SRE manually runs kubectl commands to grant RBAC roles. Average turnaround: 4 hours.

After: Self-service portal with pre-approved role templates:

# access-automation/role-templates.yaml
templates:
  - name: developer-readonly
    auto_approve: true
    max_duration: 90d
    namespaces: ["{{ .Team }}-*"]
    rbac:
      - apiGroups: [""]
        resources: ["pods", "services", "configmaps"]
        verbs: ["get", "list", "watch"]
      - apiGroups: ["apps"]
        resources: ["deployments", "replicasets"]
        verbs: ["get", "list", "watch"]
        
  - name: developer-deploy
    auto_approve: true
    max_duration: 30d
    namespaces: ["{{ .Team }}-staging", "{{ .Team }}-production"]
    rbac:
      - apiGroups: ["apps"]
        resources: ["deployments"]
        verbs: ["get", "list", "patch", "update"]
        
  - name: admin-access
    auto_approve: false  # Requires manager approval
    max_duration: 8h     # Short-lived for break-glass
    approval_chain: ["team-lead", "sre-oncall"]

Deploy Babysitting (9.8 → 0.4 hours/week)

Before: Engineer watches deployment metrics for 15 minutes after each deploy, manually rolls back if metrics degrade.

After: Fully automated canary with self-rollback. Human involvement only on automation failure:

#!/bin/bash
# scripts/automated-deploy.sh
set -euo pipefail

SERVICE=$1
VERSION=$2
CANARY_DURATION=600  # 10 minutes
ERROR_THRESHOLD=0.01  # 1% error rate

echo "Deploying $SERVICE:$VERSION with canary analysis"

# Deploy canary (10% traffic)
kubectl set image deployment/$SERVICE \
  app=$SERVICE:$VERSION \
  --namespace=production

kubectl rollout status deployment/$SERVICE --timeout=120s

# Automated canary analysis
echo "Running canary analysis for ${CANARY_DURATION}s..."
RESULT=$(curl -s "http://canary-analyzer.sre/analyze?service=$SERVICE&duration=$CANARY_DURATION&threshold=$ERROR_THRESHOLD")
STATUS=$(echo $RESULT | jq -r '.verdict')

if [ "$STATUS" = "pass" ]; then
  echo "Canary passed. Promoting to full rollout."
  kubectl scale deployment/$SERVICE --replicas=$(kubectl get deployment $SERVICE -o jsonpath='{.spec.replicas}')
  # Post to Slack
  notify_slack "Deployment $SERVICE:$VERSION completed successfully (automated)"
elif [ "$STATUS" = "fail" ]; then
  echo "Canary failed. Auto-rolling back."
  kubectl rollout undo deployment/$SERVICE
  notify_slack "Deployment $SERVICE:$VERSION auto-rolled back. Error rate: $(echo $RESULT | jq -r '.error_rate')"
  # Only page if rollback also fails
  if ! kubectl rollout status deployment/$SERVICE --timeout=60s; then
    page_oncall "Automated rollback failed for $SERVICE"
  fi
fi

Certificate Rotation (7.2 → 0 hours/week)

Before: SRE manually tracks cert expiry dates in a spreadsheet, runs renewal scripts monthly.

After: cert-manager handles everything. We only get paged if auto-renewal fails:

# cert-manager/cluster-issuer.yaml
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-prod
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    privateKeySecretRef:
      name: letsencrypt-prod-key
    solvers:
      - dns01:
          route53:
            region: us-east-1
---
# Alert only on renewal failure
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: cert-alerts
spec:
  groups:
    - name: certificates
      rules:
        - alert: CertificateRenewalFailed
          expr: certmanager_certificate_ready_status == 0
          for: 1h
          labels:
            severity: warning
          annotations:
            summary: "Certificate {{ $labels.name }} renewal failed"

Step 4: Track Progress Over Time

We run monthly toil reports and display trends on our team dashboard:

# Prometheus metrics from toil tracker
toil_hours_weekly{team="sre"} / 
  (team_size{team="sre"} * 32) * 100

# Trend over 6 months
avg_over_time(toil_percentage{team="sre"}[30d])

Toil Reduction Progress Over 6 Months

MonthToil %Automation ShippedHours Freed/Week
Month 140%Deploy automation9.4
Month 232%cert-manager + access portal18.6
Month 324%Capacity auto-scaling5.2
Month 419%Config GitOps5.8
Month 515%Data export self-service4.9
Month 612%Secret rotation automation3.1

The 50% Rule

Google's SRE book recommends that SRE teams spend no more than 50% of their time on toil. We target 15% as our steady-state maximum. When toil exceeds 15%, we pause project work and focus exclusively on automation until it drops below threshold.

This sounds extreme, but the math is clear: reducing toil from 40% to 12% freed 28 percentage points—equivalent to 2.2 full-time engineers doing creative work instead of repetitive tasks.

Key Takeaways

  1. Measure toil weekly, not annually. Monthly tracking reveals trends and creates accountability. Annual surveys are too imprecise to drive change.

  2. Prioritize by hours * (1/effort). The highest-impact automation targets are high-frequency, low-effort categories. Our deploy babysitting automation took 1 week and saved 9.8 hours/week.

  3. Self-service beats automation-for-others. The access provisioning portal doesn't just save SRE time—it saves the requesting engineer's wait time too.

  4. Set a toil budget like an error budget. When toil exceeds your threshold, stop feature work and automate. This makes toil reduction a team-level commitment, not a backlog item that never gets prioritized.

  5. Celebrate the freed capacity. Track what your team ships with the freed time. "We reduced toil by 28 points" is abstract. "We shipped VPA, chaos testing, and cross-region failover with the freed time" is concrete.

Toil is a tax on your team's creativity. Every hour spent on repetitive work is an hour not spent making the system genuinely better.

Comments

    No comments yet. Be the first to share your thoughts.