Feature Flags for Infrastructure Changes: Progressive Rollouts Without the Risk

Using feature flags to progressively roll out infrastructure changes, enabling instant rollback and percentage-based traffic shifting for risky modifications.

#feature-flags#deployment#progressive-rollout#aws
Cover image for the article: Feature Flags for Infrastructure Changes: Progressive Rollouts Without the Risk

Feature flags are universally adopted for application code. Deploy the change behind a flag, expose it to 5% of users, validate, and expand. But infrastructure changes — new load balancer configurations, database connection pool sizes, caching layers, DNS routing — are typically deployed as all-or-nothing operations. A bad infrastructure change hits 100% of traffic immediately.

We extended feature flag patterns to infrastructure rollouts, enabling percentage-based exposure and instant rollback for changes that previously required emergency hotfixes. Here is the architecture.

The Problem: Infrastructure Changes Are Binary

Consider these real incidents from our post-mortem database:

  • Connection pool increase from 20 to 50 caused database connection exhaustion across all instances simultaneously. Recovery time: 23 minutes.
  • CDN cache rule change invalidated the entire cache fleet globally. Origin servers overloaded. Recovery time: 41 minutes.
  • New WAF rule blocked legitimate API traffic from a mobile client. Impact: 100% of mobile users for 17 minutes.
  • Load balancer idle timeout change broke long-polling connections for all connected clients. Recovery time: 8 minutes.

Every one of these could have been a 30-second rollback if exposed to 2% of traffic first.

Architecture: Infrastructure Feature Flag System

Our system sits between the deployment pipeline and infrastructure configuration, intercepting changes and exposing them progressively.

Infrastructure Feature Flags Architecture

Flag Definition for Infrastructure

Infrastructure flags differ from application flags. They operate at the resource level, not the user level:

// infra-flags/src/types.ts
interface InfrastructureFlag {
  id: string;
  name: string;
  description: string;
  resourceType: 'alb' | 'security-group' | 'rds' | 'elasticache' | 'cloudfront' | 'route53';
  rolloutStrategy: RolloutStrategy;
  rollbackTrigger: RollbackTrigger;
  currentState: FlagState;
}

interface RolloutStrategy {
  type: 'percentage' | 'canary-instance' | 'regional' | 'time-based';
  stages: RolloutStage[];
  bakeTime: number;  // minutes between stages
  autoAdvance: boolean;
}

interface RolloutStage {
  percentage: number;
  duration: number;    // minutes to hold at this stage
  successCriteria: MetricCondition[];
}

interface RollbackTrigger {
  metrics: MetricCondition[];
  evaluationWindow: number;  // seconds
  action: 'immediate' | 'gradual';
}

Weighted Resource Configuration

For changes that can be split by traffic weight, we use ALB-level feature flags:

# infra-flags/rollout_controller.py
import boto3
from dataclasses import dataclass

@dataclass
class InfraFlagConfig:
    flag_id: str
    old_target_group: str
    new_target_group: str
    current_percentage: int
    max_percentage: int
    step_size: int
    bake_time_minutes: int

class InfraFlagRolloutController:
    def __init__(self):
        self.elbv2 = boto3.client('elbv2')
        self.cloudwatch = boto3.client('cloudwatch')

    async def advance_rollout(self, config: InfraFlagConfig) -> dict:
        # Check health metrics before advancing
        if not await self._check_health(config):
            return {'action': 'hold', 'reason': 'Health check failed'}

        new_percentage = min(
            config.current_percentage + config.step_size,
            config.max_percentage
        )

        # Update ALB target group weights
        await self._update_weights(
            config.old_target_group,
            config.new_target_group,
            new_percentage
        )

        return {
            'action': 'advanced',
            'previous_percentage': config.current_percentage,
            'new_percentage': new_percentage,
            'next_advance_at': datetime.utcnow() + timedelta(minutes=config.bake_time_minutes)
        }

    async def instant_rollback(self, config: InfraFlagConfig) -> dict:
        await self._update_weights(
            config.old_target_group,
            config.new_target_group,
            0  # Shift 100% back to old
        )
        return {'action': 'rollback', 'completed_at': datetime.utcnow().isoformat()}

    async def _update_weights(self, old_tg: str, new_tg: str, new_pct: int) -> None:
        self.elbv2.modify_listener(
            ListenerArn=self._get_listener_arn(),
            DefaultActions=[{
                'Type': 'forward',
                'ForwardConfig': {
                    'TargetGroups': [
                        {'TargetGroupArn': old_tg, 'Weight': 100 - new_pct},
                        {'TargetGroupArn': new_tg, 'Weight': new_pct}
                    ]
                }
            }]
        )

    async def _check_health(self, config: InfraFlagConfig) -> bool:
        metrics = await self.cloudwatch.get_metric_data(
            MetricDataQueries=[
                {
                    'Id': 'error_rate',
                    'MetricStat': {
                        'Metric': {
                            'Namespace': 'AWS/ApplicationELB',
                            'MetricName': 'HTTPCode_Target_5XX_Count',
                            'Dimensions': [
                                {'Name': 'TargetGroup', 'Value': config.new_target_group}
                            ]
                        },
                        'Period': 60,
                        'Stat': 'Sum'
                    }
                }
            ],
            StartTime=datetime.utcnow() - timedelta(minutes=5),
            EndTime=datetime.utcnow()
        )
        error_count = sum(metrics['MetricDataResults'][0]['Values'])
        return error_count < 10  # Threshold for new target group

Instance-Level Flag for Non-Splittable Changes

Some infrastructure changes cannot be split by traffic. For these, we use instance-level canary flags — applying the change to a subset of instances:

# infra-flags/configs/connection-pool-increase.yaml
flag:
  id: infra-flag-conn-pool-50
  name: "Increase DB connection pool to 50"
  description: "Progressive rollout of connection pool increase from 20 to 50"
  resource_type: rds
  rollout_strategy:
    type: canary-instance
    stages:
      - percentage: 10   # 1 of 10 instances
        duration: 30
        success_criteria:
          - metric: DatabaseConnections
            threshold: 45
            comparison: LessThan
          - metric: CPUUtilization
            namespace: AWS/RDS
            threshold: 80
            comparison: LessThan
      - percentage: 30
        duration: 30
        success_criteria:
          - metric: DatabaseConnections
            threshold: 45
            comparison: LessThan
      - percentage: 50
        duration: 60
        success_criteria:
          - metric: ReadLatency
            threshold: 10
            comparison: LessThan
      - percentage: 100
        duration: 0
  rollback_trigger:
    metrics:
      - metric: DatabaseConnections
        threshold: 48
        comparison: GreaterThan
        action: immediate

Real-World Example: WAF Rule Rollout

The WAF rule incident mentioned earlier was resolved by flag-gating new rules:

# waf/flagged-rule.tf
resource "aws_wafv2_rule_group" "new_bot_protection" {
  name     = "bot-protection-v2"
  scope    = "REGIONAL"
  capacity = 100

  rule {
    name     = "block-automated-scanners"
    priority = 1

    action {
      # During rollout: COUNT mode (log but don't block)
      # After validation: BLOCK mode
      dynamic "count" {
        for_each = var.waf_bot_rule_mode == "monitor" ? [1] : []
        content {}
      }
      dynamic "block" {
        for_each = var.waf_bot_rule_mode == "enforce" ? [1] : []
        content {}
      }
    }

    statement {
      rate_based_statement {
        limit              = 2000
        aggregate_key_type = "IP"
      }
    }

    visibility_config {
      sampled_requests_enabled   = true
      cloudwatch_metrics_enabled = true
      metric_name               = "BotProtectionV2"
    }
  }
}

The rollout follows: deploy in COUNT mode → monitor for 48 hours → analyze blocked-would-be requests → switch to BLOCK mode for 5% via weighted routing → validate → expand.

Metrics and Outcomes

Infrastructure Flag Rollout Results

MetricBefore FlagsAfter FlagsImprovement
Infra change incidents/quarter7.30.889% reduction
Mean blast radius100% traffic4.2% traffic96% reduction
Mean rollback time18 min12 sec99% faster
Infra change lead time1 week (fear)1 day7x faster
Change failure rate12%1.8%85% reduction

Integration with Deployment Pipeline

Infrastructure flags integrate into the standard CI/CD pipeline:

# .github/workflows/infra-change.yml
name: Infrastructure Change with Feature Flag
on:
  push:
    paths: ['terraform/modules/**']

jobs:
  deploy-flagged:
    steps:
      - name: Terraform Apply (flagged at 0%)
        run: |
          terraform apply -var="flag_percentage=0" -auto-approve

      - name: Start Progressive Rollout
        run: |
          infra-flag advance \
            --flag-id=${{ env.FLAG_ID }} \
            --target-percentage=5 \
            --bake-time=30m \
            --auto-advance=true \
            --rollback-on="error_rate > 1%"

      - name: Monitor Rollout
        run: |
          infra-flag watch \
            --flag-id=${{ env.FLAG_ID }} \
            --timeout=4h \
            --success-criteria="100% reached"

Key Takeaways

  1. Treat infrastructure like code: If you would not deploy application code to 100% of users without a flag, do not deploy infrastructure changes to 100% of traffic without progressive rollout.

  2. Not all changes are splittable: Some changes (like connection pool sizes) require instance-level canary patterns rather than traffic-level splitting.

  3. Monitor-then-enforce for security rules: WAF rules, rate limits, and access controls should always deploy in monitor mode first, regardless of confidence level.

  4. Automate the advancement: Manual flag progression introduces human delay and error. Automate advancement with metric gates and let rollback be the manual override.

  5. Speed follows safety: Teams that trust their rollback mechanism deploy infrastructure changes 7x faster. Safety enables velocity, not the other way around.

Infrastructure feature flags bridged the gap between "infrastructure changes are scary" and "we ship infrastructure changes multiple times daily" — transforming our most conservative operational practice into our most agile.

Comments

    No comments yet. Be the first to share your thoughts.