From Weekly to Hourly Deployments: Optimizing DORA Metrics
How we increased deployment frequency from weekly batches to hourly continuous delivery by systematically removing bottlenecks in our CI/CD pipeline.

The Problem: Batched Deployments Create Risk
Our team shipped once a week. Every Thursday afternoon, a release manager would bundle 15-30 commits into a release branch, run manual QA for two hours, then deploy during a maintenance window. This felt safe. It wasn't.
Large batches mean large blast radii. When Thursday's deployment caused a P1 incident, we couldn't pinpoint which of the 27 commits introduced the regression. Rollbacks reverted a week of work. Engineers avoided merging on Wednesdays because "the release is tomorrow."
The DORA research is clear: elite performers deploy multiple times per day with lower change failure rates than teams deploying weekly. Smaller changes are safer changes. We set a goal: reduce lead time from commit to production from 5 days to under 1 hour.
Baseline Metrics
Before optimization, our DORA metrics painted a clear picture:
| Metric | Our Baseline | Elite Benchmark |
|---|---|---|
| Deployment Frequency | 1/week | Multiple/day |
| Lead Time for Changes | 5 days | < 1 hour |
| Change Failure Rate | 18% | < 5% |
| Mean Time to Recovery | 4 hours | < 1 hour |
The Bottleneck Map
We mapped every step from commit to production and measured each phase:
Commit → PR Review (18h) → CI Build (22min) → Manual QA (3h) →
Staging Deploy (15min) → Staging Validation (2h) → Release Bundle (24h wait) →
Production Deploy (20min) → Smoke Tests (10min)
Total: ~120 hours (5 days). Our target: < 60 minutes.
The biggest bottlenecks:
- PR Review wait time (18 hours average)—not review time, wait time
- Manual QA (3 hours)—testing the same paths CI should cover
- Release bundling (24 hours)—artificial batching with no technical reason
- Staging validation (2 hours)—duplicate of production smoke tests
Phase 1: Eliminate Batching (Week 1-4)
The highest-leverage change: stop batching. Every merged PR deploys independently.
# .github/workflows/continuous-deploy.yaml
name: Continuous Deployment
on:
push:
branches: [main]
concurrency:
group: production-deploy
cancel-in-progress: false
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Build
run: npm run build
- name: Test
run: npm run test:ci
- name: Deploy to Production
run: |
./scripts/deploy.sh \
--canary-percentage=5 \
--canary-duration=300 \
--auto-promote \
--auto-rollback-on-error-rate=0.01
The concurrency block ensures deployments don't overlap. Each deployment goes through a 5-minute canary phase where 5% of traffic hits the new version. If error rates exceed 1%, automatic rollback triggers.
Canary Analysis Automation
We built an automated canary analysis tool that compares key metrics between the canary and baseline:
# scripts/canary_analysis.py
import statistics
from dataclasses import dataclass
@dataclass
class CanaryResult:
metric: str
baseline_value: float
canary_value: float
threshold: float
passed: bool
def analyze_canary(
prometheus_url: str,
canary_label: str,
baseline_label: str,
duration_minutes: int = 5,
) -> list[CanaryResult]:
metrics_config = [
{
"name": "error_rate",
"query_template": 'sum(rate(http_requests_total{{status=~"5..",version="{version}"}}[5m])) / sum(rate(http_requests_total{{version="{version}"}}[5m]))',
"threshold": 0.01, # Canary must not exceed baseline + 1%
},
{
"name": "p99_latency",
"query_template": 'histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket{{version="{version}"}}[5m])) by (le))',
"threshold": 1.2, # Canary must not exceed 120% of baseline
},
]
results = []
for config in metrics_config:
baseline_val = query_prometheus(
prometheus_url,
config["query_template"].format(version=baseline_label),
)
canary_val = query_prometheus(
prometheus_url,
config["query_template"].format(version=canary_label),
)
if config["name"] == "error_rate":
passed = canary_val <= baseline_val + config["threshold"]
else:
passed = canary_val <= baseline_val * config["threshold"]
results.append(CanaryResult(
metric=config["name"],
baseline_value=baseline_val,
canary_value=canary_val,
threshold=config["threshold"],
passed=passed,
))
return results
Phase 2: Accelerate PR Reviews (Week 2-6)
PR review wait time was our longest bottleneck. Not because reviews took long—median review time was 23 minutes—but because PRs sat in queue for hours.
Changes we made:
- Review SLA: PRs must receive first review within 2 hours during business hours
- Auto-assign reviewers: Round-robin assignment on PR creation
- Small PRs only: Hard limit of 400 lines changed (enforced by CI)
- Stacking: Allowed merge of dependent PRs in sequence
# .github/workflows/pr-size-check.yaml
name: PR Size Check
on: [pull_request]
jobs:
check-size:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Check PR Size
run: |
ADDITIONS=$(git diff --numstat origin/main...HEAD | awk '{s+=$1} END {print s}')
DELETIONS=$(git diff --numstat origin/main...HEAD | awk '{s+=$2} END {print s}')
TOTAL=$((ADDITIONS + DELETIONS))
if [ "$TOTAL" -gt 400 ]; then
echo "::error::PR too large ($TOTAL lines). Max 400. Please split."
exit 1
fi
echo "PR size: $TOTAL lines (limit: 400)"
Phase 3: Replace Manual QA with Automated Verification (Week 4-8)
Manual QA tested the same 12 user flows every deployment. We automated all of them:
// e2e/critical-paths.spec.ts
import { test, expect } from '@playwright/test';
const CRITICAL_FLOWS = [
'user-registration',
'login-password',
'login-sso',
'payment-checkout',
'subscription-upgrade',
'api-key-generation',
] as const;
test.describe('Critical Path Verification', () => {
test('checkout flow completes successfully', async ({ page }) => {
await page.goto('/pricing');
await page.click('[data-testid="plan-pro-select"]');
await page.fill('#email', 'test@example.com');
await page.fill('#card-number', '4242424242424242');
await page.click('[data-testid="submit-payment"]');
await expect(page.locator('[data-testid="success-message"]'))
.toBeVisible({ timeout: 10000 });
// Verify webhook processed
const response = await page.request.get('/api/subscription/status');
const data = await response.json();
expect(data.status).toBe('active');
});
});
Phase 4: Feature Flags for Decoupled Deploys (Week 6-10)
The final unlock: separating deployment from release. Feature flags let us deploy code to production without exposing it to users:
// lib/feature-flags.ts
interface FeatureFlag {
name: string;
enabled: boolean;
rolloutPercentage: number;
allowedUsers: string[];
}
export function isFeatureEnabled(
flag: string,
userId: string,
flags: FeatureFlag[]
): boolean {
const config = flags.find(f => f.name === flag);
if (!config || !config.enabled) return false;
// Check allowlist first
if (config.allowedUsers.includes(userId)) return true;
// Percentage rollout using consistent hashing
const hash = hashString(`${flag}:${userId}`);
return (hash % 100) < config.rolloutPercentage;
}
Results After 6 Months
| Metric | Before | After | Improvement |
|---|---|---|---|
| Deployment Frequency | 1/week | 8-12/day | 60x |
| Lead Time for Changes | 5 days | 47 minutes | 153x |
| Change Failure Rate | 18% | 3.2% | -82% |
| Mean Time to Recovery | 4 hours | 8 minutes | 30x |
| Deploy-related incidents | 4.2/month | 0.8/month | -81% |
Key Takeaways
-
Eliminate batching first. The single highest-leverage change is deploying every merge independently. Smaller changes fail less and are easier to debug.
-
Measure wait time, not work time. Our PR reviews took 23 minutes but had 18 hours of wait. Reducing wait time through auto-assignment and SLAs delivered 10x more improvement than faster reviews.
-
Automate the manual QA gate. If humans test the same flows every deployment, those flows should be automated tests that run in CI.
-
Canary deployments make speed safe. Automated canary analysis with auto-rollback lets you deploy hourly with lower risk than weekly manual deployments.
-
Feature flags decouple deploy from release. When deployment is just a code sync with no user impact, engineers stop fearing the deploy button.
Speed and safety aren't opposites—they're the same thing. The fastest teams are also the most reliable.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.