Infrastructure Drift Detection and Self-Healing Remediation

Building an automated drift detection system that identifies infrastructure divergence from declared state and triggers self-healing remediation pipelines.

#infrastructure-as-code#drift-detection#compliance#aws
Cover image for the article: Infrastructure Drift Detection and Self-Healing Remediation

Infrastructure drift is the silent killer of reliability. Every manual console change, every hotfix applied directly to a running instance, every security group rule added during an incident — they all create invisible divergence between your declared state and reality. At scale, this drift compounds until your Terraform plans show hundreds of unexpected changes and nobody trusts terraform apply anymore.

We built an automated drift detection and remediation system across 340 AWS accounts that reduced our drift surface by 94% in four months. Here is the architecture.

The Problem: Drift at Enterprise Scale

An audit across our AWS organization revealed the scope of the problem:

  • 47% of security groups had rules not declared in Terraform
  • 23% of IAM policies contained hand-attached inline policies
  • 31% of S3 buckets had configurations divergent from IaC
  • Average drift age: 67 days (some drift was 18 months old)
  • Compliance violations from drift: 89 per quarter

The business impact was severe. Drift caused 3 production incidents per month on average — resources that engineers believed existed in one configuration were actually running in another. Disaster recovery plans were unreliable because the declared state no longer matched production.

Architecture: Continuous Drift Detection Pipeline

Our system runs on a continuous loop: detect, classify, notify, and remediate.

Drift Detection Architecture

Detection Layer

The detection system runs scheduled Terraform plans across all workspaces without applying changes:

# drift_detector/scanner.py
import subprocess
import json
from dataclasses import dataclass
from enum import Enum

class DriftSeverity(Enum):
    CRITICAL = "critical"  # Security-impacting changes
    HIGH = "high"          # Production resource modifications
    MEDIUM = "medium"      # Non-production changes
    LOW = "low"            # Cosmetic or tag-only drift

@dataclass
class DriftItem:
    resource_address: str
    resource_type: str
    change_type: str  # update, delete, create
    attributes_changed: list[str]
    severity: DriftSeverity
    account_id: str
    workspace: str

class DriftScanner:
    def __init__(self, workspace_path: str):
        self.workspace_path = workspace_path

    def detect_drift(self) -> list[DriftItem]:
        result = subprocess.run(
            ["terraform", "plan", "-detailed-exitcode", "-json", "-refresh-only"],
            capture_output=True,
            text=True,
            cwd=self.workspace_path,
            timeout=600
        )

        if result.returncode == 0:
            return []  # No drift detected

        if result.returncode == 2:
            return self._parse_plan_output(result.stdout)

        raise DriftScanError(f"Terraform plan failed: {result.stderr}")

    def _parse_plan_output(self, json_output: str) -> list[DriftItem]:
        drift_items = []
        for line in json_output.strip().split('\n'):
            event = json.loads(line)
            if event.get('type') == 'resource_drift':
                change = event['change']
                drift_items.append(DriftItem(
                    resource_address=change['resource']['addr'],
                    resource_type=change['resource']['resource_type'],
                    change_type=change['action'],
                    attributes_changed=list(change.get('before_sensitive', {}).keys()),
                    severity=self._classify_severity(change),
                    account_id=self._get_account_id(),
                    workspace=self.workspace_path
                ))
        return drift_items

    def _classify_severity(self, change: dict) -> DriftSeverity:
        resource_type = change['resource']['resource_type']
        security_resources = {
            'aws_security_group', 'aws_security_group_rule',
            'aws_iam_policy', 'aws_iam_role_policy',
            'aws_kms_key', 'aws_s3_bucket_policy'
        }
        if resource_type in security_resources:
            return DriftSeverity.CRITICAL
        if 'prod' in change['resource'].get('addr', ''):
            return DriftSeverity.HIGH
        return DriftSeverity.MEDIUM

Classification and Risk Scoring

Not all drift is equal. We classify drift by risk to determine the remediation strategy:

# drift-policies/classification.yaml
rules:
  - name: security-group-ingress-drift
    resource_types: [aws_security_group_rule]
    attributes: [ingress]
    severity: critical
    auto_remediate: true
    notification: pagerduty

  - name: iam-policy-drift
    resource_types: [aws_iam_policy, aws_iam_role_policy_attachment]
    severity: critical
    auto_remediate: false  # Requires human approval
    notification: slack-security

  - name: tag-only-drift
    attributes: [tags, tags_all]
    severity: low
    auto_remediate: true
    notification: none

  - name: scaling-parameter-drift
    resource_types: [aws_autoscaling_group]
    attributes: [desired_capacity, min_size, max_size]
    severity: medium
    auto_remediate: false  # May be intentional scaling
    notification: slack-team

Self-Healing Remediation

For drift classified as auto-remediable, the system triggers a targeted terraform apply scoped to only the drifted resources:

#!/bin/bash
# remediation/auto_heal.sh

set -euo pipefail

WORKSPACE="$1"
RESOURCE_ADDRESS="$2"
DRIFT_ID="$3"

echo "Starting auto-remediation for $RESOURCE_ADDRESS in $WORKSPACE"
echo "Drift ID: $DRIFT_ID"

cd "$WORKSPACE"

# Create a targeted plan for only the drifted resource
terraform plan \
  -target="$RESOURCE_ADDRESS" \
  -refresh-only \
  -out="remediation-${DRIFT_ID}.tfplan" \
  -json > "plan-output-${DRIFT_ID}.json"

# Validate the plan only reverts drift (no new changes)
CHANGE_COUNT=$(jq '[.resource_changes[] | select(.change.actions != ["no-op"])] | length' \
  "plan-output-${DRIFT_ID}.json")

if [ "$CHANGE_COUNT" -gt 1 ]; then
  echo "ERROR: Remediation plan affects more than the target resource. Aborting."
  exit 1
fi

# Apply the remediation
terraform apply \
  -auto-approve \
  "remediation-${DRIFT_ID}.tfplan"

# Record remediation event
aws dynamodb put-item \
  --table-name drift-remediation-log \
  --item "{
    \"DriftId\": {\"S\": \"$DRIFT_ID\"},
    \"Workspace\": {\"S\": \"$WORKSPACE\"},
    \"Resource\": {\"S\": \"$RESOURCE_ADDRESS\"},
    \"RemediatedAt\": {\"S\": \"$(date -u +%Y-%m-%dT%H:%M:%SZ)\"},
    \"Status\": {\"S\": \"success\"}
  }"

Preventive Controls: Stopping Drift at the Source

Detection and remediation are reactive. We also deployed preventive controls:

SCPs for Critical Resources

Service Control Policies that prevent manual modification of IaC-managed resources:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "PreventManualSecurityGroupChanges",
      "Effect": "Deny",
      "Action": [
        "ec2:AuthorizeSecurityGroupIngress",
        "ec2:RevokeSecurityGroupIngress",
        "ec2:AuthorizeSecurityGroupEgress"
      ],
      "Resource": "*",
      "Condition": {
        "StringNotLike": {
          "aws:PrincipalArn": [
            "arn:aws:iam::*:role/TerraformExecutionRole",
            "arn:aws:iam::*:role/BreakGlassRole"
          ]
        },
        "StringEquals": {
          "ec2:ResourceTag/ManagedBy": "terraform"
        }
      }
    }
  ]
}

This ensures that only the Terraform execution role can modify tagged resources, while preserving a break-glass escape for genuine emergencies.

Metrics and Results

After 4 months of operation:

Drift Reduction Over Time

MetricMonth 0Month 4Improvement
Drifted resources2,84717194% reduction
Security-critical drift items342898% reduction
Drift-caused incidents/month3.10.294% reduction
Mean time to detect drift67 days15 minutes99.98% faster
Auto-remediation success rateN/A97.3%—
Compliance audit findings89/quarter6/quarter93% reduction

Operational Lessons

False Positive Management

Not all detected drift requires remediation. Auto-scaling groups legitimately change desired_capacity. Blue-green deployments intentionally shift target group weights. We maintain an allowlist of expected transient drift:

# drift-policies/allowlist.yaml
transient_drift:
  - resource_type: aws_autoscaling_group
    attributes: [desired_capacity]
    reason: "ASG scaling events are expected"

  - resource_type: aws_lb_target_group_attachment
    reason: "Blue-green deployments shift targets"

  - resource_type: aws_ecs_service
    attributes: [desired_count]
    reason: "Auto-scaling adjusts task count"

Break-Glass Process

When engineers need to make emergency manual changes, the break-glass process creates a tracked exception:

  1. Engineer assumes the BreakGlassRole (triggers CloudTrail alert)
  2. Makes the necessary change
  3. Opens a ticket that auto-creates a PR to update the Terraform code
  4. Drift detector marks this drift as "acknowledged, pending IaC update"

Key Takeaways

  1. Drift is a symptom, not a cause: The root cause is usually missing automation, slow IaC pipelines, or insufficient permissions for the Terraform role. Fix the workflow, not just the drift.

  2. Classify before remediating: Auto-healing everything is dangerous. Security drift should auto-remediate. Scaling parameters should not.

  3. Prevention beats detection: SCPs and RBAC that prevent manual changes to IaC-managed resources eliminate entire categories of drift.

  4. Track drift age, not just count: A single drifted resource that has been divergent for 6 months is more dangerous than 20 resources that drifted yesterday.

  5. Maintain escape hatches: Break-glass processes must exist for genuine emergencies. The goal is controlled drift, not zero human access.

Infrastructure drift is solvable, but only with a system that treats it as a continuous operational concern rather than a periodic audit finding.

Comments

    No comments yet. Be the first to share your thoughts.