From Weekly to Hourly Deployments: Optimizing DORA Metrics

How we increased deployment frequency from weekly batches to hourly continuous delivery by systematically removing bottlenecks in our CI/CD pipeline.

#dora-metrics#deployment#cicd#devops
Cover image for the article: From Weekly to Hourly Deployments: Optimizing DORA Metrics

The Problem: Batched Deployments Create Risk

Our team shipped once a week. Every Thursday afternoon, a release manager would bundle 15-30 commits into a release branch, run manual QA for two hours, then deploy during a maintenance window. This felt safe. It wasn't.

Large batches mean large blast radii. When Thursday's deployment caused a P1 incident, we couldn't pinpoint which of the 27 commits introduced the regression. Rollbacks reverted a week of work. Engineers avoided merging on Wednesdays because "the release is tomorrow."

The DORA research is clear: elite performers deploy multiple times per day with lower change failure rates than teams deploying weekly. Smaller changes are safer changes. We set a goal: reduce lead time from commit to production from 5 days to under 1 hour.

Baseline Metrics

Before optimization, our DORA metrics painted a clear picture:

MetricOur BaselineElite Benchmark
Deployment Frequency1/weekMultiple/day
Lead Time for Changes5 days< 1 hour
Change Failure Rate18%< 5%
Mean Time to Recovery4 hours< 1 hour

DORA Metrics Baseline vs Target

The Bottleneck Map

We mapped every step from commit to production and measured each phase:

Commit → PR Review (18h) → CI Build (22min) → Manual QA (3h) → 
Staging Deploy (15min) → Staging Validation (2h) → Release Bundle (24h wait) →
Production Deploy (20min) → Smoke Tests (10min)

Total: ~120 hours (5 days). Our target: < 60 minutes.

The biggest bottlenecks:

  1. PR Review wait time (18 hours average)—not review time, wait time
  2. Manual QA (3 hours)—testing the same paths CI should cover
  3. Release bundling (24 hours)—artificial batching with no technical reason
  4. Staging validation (2 hours)—duplicate of production smoke tests

Phase 1: Eliminate Batching (Week 1-4)

The highest-leverage change: stop batching. Every merged PR deploys independently.

# .github/workflows/continuous-deploy.yaml
name: Continuous Deployment
on:
  push:
    branches: [main]

concurrency:
  group: production-deploy
  cancel-in-progress: false

jobs:
  deploy:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      
      - name: Build
        run: npm run build
        
      - name: Test
        run: npm run test:ci
        
      - name: Deploy to Production
        run: |
          ./scripts/deploy.sh \
            --canary-percentage=5 \
            --canary-duration=300 \
            --auto-promote \
            --auto-rollback-on-error-rate=0.01

The concurrency block ensures deployments don't overlap. Each deployment goes through a 5-minute canary phase where 5% of traffic hits the new version. If error rates exceed 1%, automatic rollback triggers.

Canary Analysis Automation

We built an automated canary analysis tool that compares key metrics between the canary and baseline:

# scripts/canary_analysis.py
import statistics
from dataclasses import dataclass

@dataclass
class CanaryResult:
    metric: str
    baseline_value: float
    canary_value: float
    threshold: float
    passed: bool

def analyze_canary(
    prometheus_url: str,
    canary_label: str,
    baseline_label: str,
    duration_minutes: int = 5,
) -> list[CanaryResult]:
    metrics_config = [
        {
            "name": "error_rate",
            "query_template": 'sum(rate(http_requests_total{{status=~"5..",version="{version}"}}[5m])) / sum(rate(http_requests_total{{version="{version}"}}[5m]))',
            "threshold": 0.01,  # Canary must not exceed baseline + 1%
        },
        {
            "name": "p99_latency",
            "query_template": 'histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket{{version="{version}"}}[5m])) by (le))',
            "threshold": 1.2,  # Canary must not exceed 120% of baseline
        },
    ]

    results = []
    for config in metrics_config:
        baseline_val = query_prometheus(
            prometheus_url,
            config["query_template"].format(version=baseline_label),
        )
        canary_val = query_prometheus(
            prometheus_url,
            config["query_template"].format(version=canary_label),
        )

        if config["name"] == "error_rate":
            passed = canary_val &#x3C;= baseline_val + config["threshold"]
        else:
            passed = canary_val &#x3C;= baseline_val * config["threshold"]

        results.append(CanaryResult(
            metric=config["name"],
            baseline_value=baseline_val,
            canary_value=canary_val,
            threshold=config["threshold"],
            passed=passed,
        ))

    return results

Phase 2: Accelerate PR Reviews (Week 2-6)

PR review wait time was our longest bottleneck. Not because reviews took long—median review time was 23 minutes—but because PRs sat in queue for hours.

Changes we made:

  1. Review SLA: PRs must receive first review within 2 hours during business hours
  2. Auto-assign reviewers: Round-robin assignment on PR creation
  3. Small PRs only: Hard limit of 400 lines changed (enforced by CI)
  4. Stacking: Allowed merge of dependent PRs in sequence
# .github/workflows/pr-size-check.yaml
name: PR Size Check
on: [pull_request]

jobs:
  check-size:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - name: Check PR Size
        run: |
          ADDITIONS=$(git diff --numstat origin/main...HEAD | awk '{s+=$1} END {print s}')
          DELETIONS=$(git diff --numstat origin/main...HEAD | awk '{s+=$2} END {print s}')
          TOTAL=$((ADDITIONS + DELETIONS))
          
          if [ "$TOTAL" -gt 400 ]; then
            echo "::error::PR too large ($TOTAL lines). Max 400. Please split."
            exit 1
          fi
          
          echo "PR size: $TOTAL lines (limit: 400)"

Phase 3: Replace Manual QA with Automated Verification (Week 4-8)

Manual QA tested the same 12 user flows every deployment. We automated all of them:

// e2e/critical-paths.spec.ts
import { test, expect } from '@playwright/test';

const CRITICAL_FLOWS = [
  'user-registration',
  'login-password',
  'login-sso', 
  'payment-checkout',
  'subscription-upgrade',
  'api-key-generation',
] as const;

test.describe('Critical Path Verification', () => {
  test('checkout flow completes successfully', async ({ page }) => {
    await page.goto('/pricing');
    await page.click('[data-testid="plan-pro-select"]');
    await page.fill('#email', 'test@example.com');
    await page.fill('#card-number', '4242424242424242');
    await page.click('[data-testid="submit-payment"]');
    
    await expect(page.locator('[data-testid="success-message"]'))
      .toBeVisible({ timeout: 10000 });
    
    // Verify webhook processed
    const response = await page.request.get('/api/subscription/status');
    const data = await response.json();
    expect(data.status).toBe('active');
  });
});

Phase 4: Feature Flags for Decoupled Deploys (Week 6-10)

The final unlock: separating deployment from release. Feature flags let us deploy code to production without exposing it to users:

// lib/feature-flags.ts
interface FeatureFlag {
  name: string;
  enabled: boolean;
  rolloutPercentage: number;
  allowedUsers: string[];
}

export function isFeatureEnabled(
  flag: string, 
  userId: string,
  flags: FeatureFlag[]
): boolean {
  const config = flags.find(f => f.name === flag);
  if (!config || !config.enabled) return false;
  
  // Check allowlist first
  if (config.allowedUsers.includes(userId)) return true;
  
  // Percentage rollout using consistent hashing
  const hash = hashString(`${flag}:${userId}`);
  return (hash % 100) &#x3C; config.rolloutPercentage;
}

Deployment Pipeline Before and After

Results After 6 Months

MetricBeforeAfterImprovement
Deployment Frequency1/week8-12/day60x
Lead Time for Changes5 days47 minutes153x
Change Failure Rate18%3.2%-82%
Mean Time to Recovery4 hours8 minutes30x
Deploy-related incidents4.2/month0.8/month-81%

Key Takeaways

  1. Eliminate batching first. The single highest-leverage change is deploying every merge independently. Smaller changes fail less and are easier to debug.

  2. Measure wait time, not work time. Our PR reviews took 23 minutes but had 18 hours of wait. Reducing wait time through auto-assignment and SLAs delivered 10x more improvement than faster reviews.

  3. Automate the manual QA gate. If humans test the same flows every deployment, those flows should be automated tests that run in CI.

  4. Canary deployments make speed safe. Automated canary analysis with auto-rollback lets you deploy hourly with lower risk than weekly manual deployments.

  5. Feature flags decouple deploy from release. When deployment is just a code sync with no user impact, engineers stop fearing the deploy button.

Speed and safety aren't opposites—they're the same thing. The fastest teams are also the most reliable.

Comments

    No comments yet. Be the first to share your thoughts.