Feature Flags for Infrastructure Changes: Progressive Rollouts Without the Risk
Using feature flags to progressively roll out infrastructure changes, enabling instant rollback and percentage-based traffic shifting for risky modifications.

Feature flags are universally adopted for application code. Deploy the change behind a flag, expose it to 5% of users, validate, and expand. But infrastructure changes — new load balancer configurations, database connection pool sizes, caching layers, DNS routing — are typically deployed as all-or-nothing operations. A bad infrastructure change hits 100% of traffic immediately.
We extended feature flag patterns to infrastructure rollouts, enabling percentage-based exposure and instant rollback for changes that previously required emergency hotfixes. Here is the architecture.
The Problem: Infrastructure Changes Are Binary
Consider these real incidents from our post-mortem database:
- Connection pool increase from 20 to 50 caused database connection exhaustion across all instances simultaneously. Recovery time: 23 minutes.
- CDN cache rule change invalidated the entire cache fleet globally. Origin servers overloaded. Recovery time: 41 minutes.
- New WAF rule blocked legitimate API traffic from a mobile client. Impact: 100% of mobile users for 17 minutes.
- Load balancer idle timeout change broke long-polling connections for all connected clients. Recovery time: 8 minutes.
Every one of these could have been a 30-second rollback if exposed to 2% of traffic first.
Architecture: Infrastructure Feature Flag System
Our system sits between the deployment pipeline and infrastructure configuration, intercepting changes and exposing them progressively.
Flag Definition for Infrastructure
Infrastructure flags differ from application flags. They operate at the resource level, not the user level:
// infra-flags/src/types.ts
interface InfrastructureFlag {
id: string;
name: string;
description: string;
resourceType: 'alb' | 'security-group' | 'rds' | 'elasticache' | 'cloudfront' | 'route53';
rolloutStrategy: RolloutStrategy;
rollbackTrigger: RollbackTrigger;
currentState: FlagState;
}
interface RolloutStrategy {
type: 'percentage' | 'canary-instance' | 'regional' | 'time-based';
stages: RolloutStage[];
bakeTime: number; // minutes between stages
autoAdvance: boolean;
}
interface RolloutStage {
percentage: number;
duration: number; // minutes to hold at this stage
successCriteria: MetricCondition[];
}
interface RollbackTrigger {
metrics: MetricCondition[];
evaluationWindow: number; // seconds
action: 'immediate' | 'gradual';
}
Weighted Resource Configuration
For changes that can be split by traffic weight, we use ALB-level feature flags:
# infra-flags/rollout_controller.py
import boto3
from dataclasses import dataclass
@dataclass
class InfraFlagConfig:
flag_id: str
old_target_group: str
new_target_group: str
current_percentage: int
max_percentage: int
step_size: int
bake_time_minutes: int
class InfraFlagRolloutController:
def __init__(self):
self.elbv2 = boto3.client('elbv2')
self.cloudwatch = boto3.client('cloudwatch')
async def advance_rollout(self, config: InfraFlagConfig) -> dict:
# Check health metrics before advancing
if not await self._check_health(config):
return {'action': 'hold', 'reason': 'Health check failed'}
new_percentage = min(
config.current_percentage + config.step_size,
config.max_percentage
)
# Update ALB target group weights
await self._update_weights(
config.old_target_group,
config.new_target_group,
new_percentage
)
return {
'action': 'advanced',
'previous_percentage': config.current_percentage,
'new_percentage': new_percentage,
'next_advance_at': datetime.utcnow() + timedelta(minutes=config.bake_time_minutes)
}
async def instant_rollback(self, config: InfraFlagConfig) -> dict:
await self._update_weights(
config.old_target_group,
config.new_target_group,
0 # Shift 100% back to old
)
return {'action': 'rollback', 'completed_at': datetime.utcnow().isoformat()}
async def _update_weights(self, old_tg: str, new_tg: str, new_pct: int) -> None:
self.elbv2.modify_listener(
ListenerArn=self._get_listener_arn(),
DefaultActions=[{
'Type': 'forward',
'ForwardConfig': {
'TargetGroups': [
{'TargetGroupArn': old_tg, 'Weight': 100 - new_pct},
{'TargetGroupArn': new_tg, 'Weight': new_pct}
]
}
}]
)
async def _check_health(self, config: InfraFlagConfig) -> bool:
metrics = await self.cloudwatch.get_metric_data(
MetricDataQueries=[
{
'Id': 'error_rate',
'MetricStat': {
'Metric': {
'Namespace': 'AWS/ApplicationELB',
'MetricName': 'HTTPCode_Target_5XX_Count',
'Dimensions': [
{'Name': 'TargetGroup', 'Value': config.new_target_group}
]
},
'Period': 60,
'Stat': 'Sum'
}
}
],
StartTime=datetime.utcnow() - timedelta(minutes=5),
EndTime=datetime.utcnow()
)
error_count = sum(metrics['MetricDataResults'][0]['Values'])
return error_count < 10 # Threshold for new target group
Instance-Level Flag for Non-Splittable Changes
Some infrastructure changes cannot be split by traffic. For these, we use instance-level canary flags — applying the change to a subset of instances:
# infra-flags/configs/connection-pool-increase.yaml
flag:
id: infra-flag-conn-pool-50
name: "Increase DB connection pool to 50"
description: "Progressive rollout of connection pool increase from 20 to 50"
resource_type: rds
rollout_strategy:
type: canary-instance
stages:
- percentage: 10 # 1 of 10 instances
duration: 30
success_criteria:
- metric: DatabaseConnections
threshold: 45
comparison: LessThan
- metric: CPUUtilization
namespace: AWS/RDS
threshold: 80
comparison: LessThan
- percentage: 30
duration: 30
success_criteria:
- metric: DatabaseConnections
threshold: 45
comparison: LessThan
- percentage: 50
duration: 60
success_criteria:
- metric: ReadLatency
threshold: 10
comparison: LessThan
- percentage: 100
duration: 0
rollback_trigger:
metrics:
- metric: DatabaseConnections
threshold: 48
comparison: GreaterThan
action: immediate
Real-World Example: WAF Rule Rollout
The WAF rule incident mentioned earlier was resolved by flag-gating new rules:
# waf/flagged-rule.tf
resource "aws_wafv2_rule_group" "new_bot_protection" {
name = "bot-protection-v2"
scope = "REGIONAL"
capacity = 100
rule {
name = "block-automated-scanners"
priority = 1
action {
# During rollout: COUNT mode (log but don't block)
# After validation: BLOCK mode
dynamic "count" {
for_each = var.waf_bot_rule_mode == "monitor" ? [1] : []
content {}
}
dynamic "block" {
for_each = var.waf_bot_rule_mode == "enforce" ? [1] : []
content {}
}
}
statement {
rate_based_statement {
limit = 2000
aggregate_key_type = "IP"
}
}
visibility_config {
sampled_requests_enabled = true
cloudwatch_metrics_enabled = true
metric_name = "BotProtectionV2"
}
}
}
The rollout follows: deploy in COUNT mode → monitor for 48 hours → analyze blocked-would-be requests → switch to BLOCK mode for 5% via weighted routing → validate → expand.
Metrics and Outcomes
| Metric | Before Flags | After Flags | Improvement |
|---|---|---|---|
| Infra change incidents/quarter | 7.3 | 0.8 | 89% reduction |
| Mean blast radius | 100% traffic | 4.2% traffic | 96% reduction |
| Mean rollback time | 18 min | 12 sec | 99% faster |
| Infra change lead time | 1 week (fear) | 1 day | 7x faster |
| Change failure rate | 12% | 1.8% | 85% reduction |
Integration with Deployment Pipeline
Infrastructure flags integrate into the standard CI/CD pipeline:
# .github/workflows/infra-change.yml
name: Infrastructure Change with Feature Flag
on:
push:
paths: ['terraform/modules/**']
jobs:
deploy-flagged:
steps:
- name: Terraform Apply (flagged at 0%)
run: |
terraform apply -var="flag_percentage=0" -auto-approve
- name: Start Progressive Rollout
run: |
infra-flag advance \
--flag-id=${{ env.FLAG_ID }} \
--target-percentage=5 \
--bake-time=30m \
--auto-advance=true \
--rollback-on="error_rate > 1%"
- name: Monitor Rollout
run: |
infra-flag watch \
--flag-id=${{ env.FLAG_ID }} \
--timeout=4h \
--success-criteria="100% reached"
Key Takeaways
-
Treat infrastructure like code: If you would not deploy application code to 100% of users without a flag, do not deploy infrastructure changes to 100% of traffic without progressive rollout.
-
Not all changes are splittable: Some changes (like connection pool sizes) require instance-level canary patterns rather than traffic-level splitting.
-
Monitor-then-enforce for security rules: WAF rules, rate limits, and access controls should always deploy in monitor mode first, regardless of confidence level.
-
Automate the advancement: Manual flag progression introduces human delay and error. Automate advancement with metric gates and let rollback be the manual override.
-
Speed follows safety: Teams that trust their rollback mechanism deploy infrastructure changes 7x faster. Safety enables velocity, not the other way around.
Infrastructure feature flags bridged the gap between "infrastructure changes are scary" and "we ship infrastructure changes multiple times daily" — transforming our most conservative operational practice into our most agile.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.