Implementing Error Budgets That Drive Engineering Decisions
A practical guide to implementing SRE error budget policies that align reliability targets with product velocity and empower teams to make data-driven tradeoffs.

The Problem: Reliability vs. Velocity Without a Framework
Every engineering organization faces the same tension: ship faster or keep things stable. Without a formal mechanism to balance these priorities, teams oscillate between "move fast and break things" sprints and reactive firefighting periods where all feature work stops.
At our organization, we experienced this firsthand. Our platform team maintained 47 microservices serving 12 million daily active users. We had SLOs defined on paper, but no enforcement mechanism. The result: quarterly reliability averaged 99.2% against a 99.9% target, and engineers had no objective framework for deciding when to slow down.
After implementing a formal error budget policy, our reliability improved to 99.94%, deployment frequency increased by 34%, and—critically—cross-team conflicts about "should we ship this?" dropped by 80%.
Defining the Error Budget
An error budget is the inverse of your SLO. If your SLO is 99.9% availability over a 30-day rolling window, your error budget is 0.1%—roughly 43.2 minutes of allowed downtime per month.
The key insight: this budget belongs to the product team, not the SRE team. Product decides how to spend it—on risky deployments, infrastructure migrations, or experiments.
# error-budget-policy.yaml
apiVersion: sre.company.io/v1
kind: ErrorBudgetPolicy
metadata:
name: payments-service-policy
spec:
service: payments-service
slo:
target: 99.95
window: 30d
metric: successful_requests / total_requests
budget:
total_minutes: 21.6 # 0.05% of 30 days
thresholds:
- level: green
remaining: ">50%"
actions:
- allow_risky_deployments
- allow_experiments
- level: yellow
remaining: "25%-50%"
actions:
- require_canary_deployments
- increase_monitoring
- level: red
remaining: "<25%"
actions:
- freeze_feature_deployments
- mandatory_reliability_sprint
- level: exhausted
remaining: "0%"
actions:
- full_deployment_freeze
- incident_review_required
- escalate_to_vp_engineering
Implementation Architecture
Our error budget system consists of four components: SLI collection, budget calculation, policy enforcement, and reporting.
SLI Collection
We use Prometheus to collect Service Level Indicators. For request-driven services, we track request success rate and latency distribution:
# SLI: Request success rate (excluding client errors)
sum(rate(http_requests_total{
service="payments",
status_code!~"5.."
}[5m])) /
sum(rate(http_requests_total{
service="payments"
}[5m]))
# SLI: Latency - percentage of requests under 300ms
sum(rate(http_request_duration_seconds_bucket{
service="payments",
le="0.3"
}[5m])) /
sum(rate(http_request_duration_seconds_count{
service="payments"
}[5m]))
Budget Calculation Engine
We built a service that computes remaining budget every minute and publishes the state to both Prometheus (for alerting) and a PostgreSQL store (for historical analysis):
from datetime import datetime, timedelta
from dataclasses import dataclass
@dataclass
class ErrorBudget:
service: str
slo_target: float
window_days: int
bad_minutes_consumed: float
@property
def total_budget_minutes(self) -> float:
total_window_minutes = self.window_days * 24 * 60
return total_window_minutes * (1 - self.slo_target / 100)
@property
def remaining_percentage(self) -> float:
remaining = self.total_budget_minutes - self.bad_minutes_consumed
return max(0, (remaining / self.total_budget_minutes) * 100)
@property
def policy_level(self) -> str:
pct = self.remaining_percentage
if pct > 50:
return "green"
elif pct > 25:
return "yellow"
elif pct > 0:
return "red"
return "exhausted"
def can_deploy(self, risk_level: str) -> bool:
allowed = {
"green": ["low", "medium", "high"],
"yellow": ["low", "medium"],
"red": ["low"],
"exhausted": [],
}
return risk_level in allowed[self.policy_level]
Policy Enforcement in CI/CD
The critical piece: integrating budget checks into your deployment pipeline. We added a gate in our CI/CD system that queries the budget service before allowing deployments:
# .github/workflows/deploy.yaml (simplified)
jobs:
error-budget-check:
runs-on: ubuntu-latest
steps:
- name: Check Error Budget
run: |
BUDGET_STATUS=$(curl -s https://sre-api.internal/budget/payments-service)
LEVEL=$(echo $BUDGET_STATUS | jq -r '.policy_level')
RISK=$(echo $BUDGET_STATUS | jq -r '.deployment_risk')
if [ "$LEVEL" = "exhausted" ]; then
echo "ERROR: Budget exhausted. Deployment blocked."
exit 1
fi
if [ "$LEVEL" = "red" ] && [ "$RISK" != "low" ]; then
echo "ERROR: Budget critical. Only low-risk deployments allowed."
exit 1
fi
Operationalizing the Policy
Weekly Budget Reviews
Every Monday, each service team receives an automated budget report. This 5-minute review replaces hour-long reliability meetings:
- Current budget remaining (percentage and minutes)
- Biggest budget consumers from the previous week
- Projected exhaustion date at current burn rate
- Recommended actions based on policy level
Burn Rate Alerts
Rather than alerting only when the budget is exhausted, we alert on burn rate—how fast the budget is being consumed relative to the window:
- 1x burn rate: budget will exactly exhaust at window end (informational)
- 2x burn rate: budget will exhaust in half the window (warning)
- 10x burn rate: budget will exhaust in 3 days (page)
- 100x burn rate: active incident consuming budget rapidly (page immediately)
Stakeholder Buy-In
The hardest part isn't technical—it's cultural. We secured buy-in by:
- Starting with a 3-month observation period (no enforcement, just reporting)
- Letting product managers see how budget correlated with user complaints
- Framing budgets as "freedom to ship" rather than "reliability police"
- Making the VP of Engineering the escalation point for exhausted budgets
Results After 12 Months
| Metric | Before | After | Change |
|---|---|---|---|
| Monthly availability | 99.2% | 99.94% | +0.74% |
| Deployment frequency | 12/week | 16/week | +34% |
| Mean time to recovery | 47 min | 18 min | -62% |
| Cross-team escalations | 23/quarter | 4/quarter | -83% |
| Engineer satisfaction (reliability) | 3.1/5 | 4.4/5 | +42% |
Key Takeaways
-
Error budgets turn reliability into a measurable resource, not an abstract goal. Teams can make rational decisions about risk because they know exactly how much budget remains.
-
Automate enforcement early. Manual enforcement creates political conflicts. Automated CI/CD gates remove the human from the decision loop.
-
Start with observation, not enforcement. Give teams 2-3 months to understand their burn patterns before any deployment freezes kick in.
-
The budget belongs to product, not SRE. This framing is essential. SRE defines the SLO; product decides how to spend the budget.
-
Pair budgets with burn-rate alerts. Knowing you have 40% remaining is less actionable than knowing you'll exhaust it in 3 days at the current rate.
Error budgets aren't just an SRE tool—they're an organizational alignment mechanism. When everyone shares the same reliability currency, the velocity-vs-stability debate becomes a simple accounting problem.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.