Infrastructure Drift Detection and Self-Healing Remediation
Building an automated drift detection system that identifies infrastructure divergence from declared state and triggers self-healing remediation pipelines.

Infrastructure drift is the silent killer of reliability. Every manual console change, every hotfix applied directly to a running instance, every security group rule added during an incident — they all create invisible divergence between your declared state and reality. At scale, this drift compounds until your Terraform plans show hundreds of unexpected changes and nobody trusts terraform apply anymore.
We built an automated drift detection and remediation system across 340 AWS accounts that reduced our drift surface by 94% in four months. Here is the architecture.
The Problem: Drift at Enterprise Scale
An audit across our AWS organization revealed the scope of the problem:
- 47% of security groups had rules not declared in Terraform
- 23% of IAM policies contained hand-attached inline policies
- 31% of S3 buckets had configurations divergent from IaC
- Average drift age: 67 days (some drift was 18 months old)
- Compliance violations from drift: 89 per quarter
The business impact was severe. Drift caused 3 production incidents per month on average — resources that engineers believed existed in one configuration were actually running in another. Disaster recovery plans were unreliable because the declared state no longer matched production.
Architecture: Continuous Drift Detection Pipeline
Our system runs on a continuous loop: detect, classify, notify, and remediate.
Detection Layer
The detection system runs scheduled Terraform plans across all workspaces without applying changes:
# drift_detector/scanner.py
import subprocess
import json
from dataclasses import dataclass
from enum import Enum
class DriftSeverity(Enum):
CRITICAL = "critical" # Security-impacting changes
HIGH = "high" # Production resource modifications
MEDIUM = "medium" # Non-production changes
LOW = "low" # Cosmetic or tag-only drift
@dataclass
class DriftItem:
resource_address: str
resource_type: str
change_type: str # update, delete, create
attributes_changed: list[str]
severity: DriftSeverity
account_id: str
workspace: str
class DriftScanner:
def __init__(self, workspace_path: str):
self.workspace_path = workspace_path
def detect_drift(self) -> list[DriftItem]:
result = subprocess.run(
["terraform", "plan", "-detailed-exitcode", "-json", "-refresh-only"],
capture_output=True,
text=True,
cwd=self.workspace_path,
timeout=600
)
if result.returncode == 0:
return [] # No drift detected
if result.returncode == 2:
return self._parse_plan_output(result.stdout)
raise DriftScanError(f"Terraform plan failed: {result.stderr}")
def _parse_plan_output(self, json_output: str) -> list[DriftItem]:
drift_items = []
for line in json_output.strip().split('\n'):
event = json.loads(line)
if event.get('type') == 'resource_drift':
change = event['change']
drift_items.append(DriftItem(
resource_address=change['resource']['addr'],
resource_type=change['resource']['resource_type'],
change_type=change['action'],
attributes_changed=list(change.get('before_sensitive', {}).keys()),
severity=self._classify_severity(change),
account_id=self._get_account_id(),
workspace=self.workspace_path
))
return drift_items
def _classify_severity(self, change: dict) -> DriftSeverity:
resource_type = change['resource']['resource_type']
security_resources = {
'aws_security_group', 'aws_security_group_rule',
'aws_iam_policy', 'aws_iam_role_policy',
'aws_kms_key', 'aws_s3_bucket_policy'
}
if resource_type in security_resources:
return DriftSeverity.CRITICAL
if 'prod' in change['resource'].get('addr', ''):
return DriftSeverity.HIGH
return DriftSeverity.MEDIUM
Classification and Risk Scoring
Not all drift is equal. We classify drift by risk to determine the remediation strategy:
# drift-policies/classification.yaml
rules:
- name: security-group-ingress-drift
resource_types: [aws_security_group_rule]
attributes: [ingress]
severity: critical
auto_remediate: true
notification: pagerduty
- name: iam-policy-drift
resource_types: [aws_iam_policy, aws_iam_role_policy_attachment]
severity: critical
auto_remediate: false # Requires human approval
notification: slack-security
- name: tag-only-drift
attributes: [tags, tags_all]
severity: low
auto_remediate: true
notification: none
- name: scaling-parameter-drift
resource_types: [aws_autoscaling_group]
attributes: [desired_capacity, min_size, max_size]
severity: medium
auto_remediate: false # May be intentional scaling
notification: slack-team
Self-Healing Remediation
For drift classified as auto-remediable, the system triggers a targeted terraform apply scoped to only the drifted resources:
#!/bin/bash
# remediation/auto_heal.sh
set -euo pipefail
WORKSPACE="$1"
RESOURCE_ADDRESS="$2"
DRIFT_ID="$3"
echo "Starting auto-remediation for $RESOURCE_ADDRESS in $WORKSPACE"
echo "Drift ID: $DRIFT_ID"
cd "$WORKSPACE"
# Create a targeted plan for only the drifted resource
terraform plan \
-target="$RESOURCE_ADDRESS" \
-refresh-only \
-out="remediation-${DRIFT_ID}.tfplan" \
-json > "plan-output-${DRIFT_ID}.json"
# Validate the plan only reverts drift (no new changes)
CHANGE_COUNT=$(jq '[.resource_changes[] | select(.change.actions != ["no-op"])] | length' \
"plan-output-${DRIFT_ID}.json")
if [ "$CHANGE_COUNT" -gt 1 ]; then
echo "ERROR: Remediation plan affects more than the target resource. Aborting."
exit 1
fi
# Apply the remediation
terraform apply \
-auto-approve \
"remediation-${DRIFT_ID}.tfplan"
# Record remediation event
aws dynamodb put-item \
--table-name drift-remediation-log \
--item "{
\"DriftId\": {\"S\": \"$DRIFT_ID\"},
\"Workspace\": {\"S\": \"$WORKSPACE\"},
\"Resource\": {\"S\": \"$RESOURCE_ADDRESS\"},
\"RemediatedAt\": {\"S\": \"$(date -u +%Y-%m-%dT%H:%M:%SZ)\"},
\"Status\": {\"S\": \"success\"}
}"
Preventive Controls: Stopping Drift at the Source
Detection and remediation are reactive. We also deployed preventive controls:
SCPs for Critical Resources
Service Control Policies that prevent manual modification of IaC-managed resources:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "PreventManualSecurityGroupChanges",
"Effect": "Deny",
"Action": [
"ec2:AuthorizeSecurityGroupIngress",
"ec2:RevokeSecurityGroupIngress",
"ec2:AuthorizeSecurityGroupEgress"
],
"Resource": "*",
"Condition": {
"StringNotLike": {
"aws:PrincipalArn": [
"arn:aws:iam::*:role/TerraformExecutionRole",
"arn:aws:iam::*:role/BreakGlassRole"
]
},
"StringEquals": {
"ec2:ResourceTag/ManagedBy": "terraform"
}
}
}
]
}
This ensures that only the Terraform execution role can modify tagged resources, while preserving a break-glass escape for genuine emergencies.
Metrics and Results
After 4 months of operation:
| Metric | Month 0 | Month 4 | Improvement |
|---|---|---|---|
| Drifted resources | 2,847 | 171 | 94% reduction |
| Security-critical drift items | 342 | 8 | 98% reduction |
| Drift-caused incidents/month | 3.1 | 0.2 | 94% reduction |
| Mean time to detect drift | 67 days | 15 minutes | 99.98% faster |
| Auto-remediation success rate | N/A | 97.3% | — |
| Compliance audit findings | 89/quarter | 6/quarter | 93% reduction |
Operational Lessons
False Positive Management
Not all detected drift requires remediation. Auto-scaling groups legitimately change desired_capacity. Blue-green deployments intentionally shift target group weights. We maintain an allowlist of expected transient drift:
# drift-policies/allowlist.yaml
transient_drift:
- resource_type: aws_autoscaling_group
attributes: [desired_capacity]
reason: "ASG scaling events are expected"
- resource_type: aws_lb_target_group_attachment
reason: "Blue-green deployments shift targets"
- resource_type: aws_ecs_service
attributes: [desired_count]
reason: "Auto-scaling adjusts task count"
Break-Glass Process
When engineers need to make emergency manual changes, the break-glass process creates a tracked exception:
- Engineer assumes the BreakGlassRole (triggers CloudTrail alert)
- Makes the necessary change
- Opens a ticket that auto-creates a PR to update the Terraform code
- Drift detector marks this drift as "acknowledged, pending IaC update"
Key Takeaways
-
Drift is a symptom, not a cause: The root cause is usually missing automation, slow IaC pipelines, or insufficient permissions for the Terraform role. Fix the workflow, not just the drift.
-
Classify before remediating: Auto-healing everything is dangerous. Security drift should auto-remediate. Scaling parameters should not.
-
Prevention beats detection: SCPs and RBAC that prevent manual changes to IaC-managed resources eliminate entire categories of drift.
-
Track drift age, not just count: A single drifted resource that has been divergent for 6 months is more dangerous than 20 resources that drifted yesterday.
-
Maintain escape hatches: Break-glass processes must exist for genuine emergencies. The goal is controlled drift, not zero human access.
Infrastructure drift is solvable, but only with a system that treats it as a continuous operational concern rather than a periodic audit finding.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.