AWS Cost Anomaly Detection: Catching Runaway Spend Before It Hits Your Bill

How we configured Cost Anomaly Detection to identify unexpected spending within 6 hours, saving $47,000 in a single quarter from early detection of misconfigurations.

#aws#cost-optimization#finops#monitoring
Cover image for the article: AWS Cost Anomaly Detection: Catching Runaway Spend Before It Hits Your Bill

A developer provisioned a p4d.24xlarge instance for a machine learning experiment on a Friday afternoon. They forgot to terminate it. By Monday morning, the instance had accumulated $2,400 in charges running idle. This was not a one-time incident. Over the previous quarter, we identified $47,000 in wasted spend from forgotten resources, misconfigured auto-scaling, and runaway data transfer — all caught after the monthly bill arrived.

We deployed AWS Cost Anomaly Detection with custom monitors and alert thresholds that catch spending deviations within 6 hours. Combined with automated remediation workflows, we now stop bleeding spend before it compounds.

The Problem: Bills Are Retrospective, Anomalies Are Real-Time

The default AWS billing experience is fundamentally backwards for cost control:

  • Billing dashboard updates with 12-24 hour delay
  • Budget alerts fire when you have already exceeded the threshold
  • Cost Explorer requires manual investigation
  • Monthly bills reveal problems 30 days after they started

In a team of 45 engineers across 12 AWS accounts, resources get provisioned daily. Without real-time anomaly detection, you are always discovering last month's mistakes.

Cost Anomaly Detection Architecture

Architecture: Layered Anomaly Monitors

AWS Cost Anomaly Detection uses machine learning to establish spending baselines and alert on deviations. We configured three layers of monitors:

Account-level monitors — Catch broad spending spikes across any service Service-level monitors — Detect unusual activity within specific high-cost services (EC2, RDS, Bedrock) Linked account monitors — Track per-team spending against historical patterns

Each layer has different sensitivity thresholds because the blast radius differs. An account-level anomaly at $500 is critical. A service-level anomaly at $50 might indicate a new feature rollout.

Implementation: Custom Monitor Configuration

import {
  CostExplorerClient,
  CreateAnomalyMonitorCommand,
  CreateAnomalySubscriptionCommand,
} from '@aws-sdk/client-cost-explorer';

const client = new CostExplorerClient({ region: 'us-east-1' });

// Service-level monitor for high-cost services
const serviceMonitor = new CreateAnomalyMonitorCommand({
  AnomalyMonitor: {
    MonitorName: 'HighCostServiceMonitor',
    MonitorType: 'DIMENSIONAL',
    MonitorDimension: 'SERVICE',
  },
});

const { MonitorArn: serviceMonitorArn } = await client.send(serviceMonitor);

// Custom monitor for specific cost allocation tags
const teamMonitor = new CreateAnomalyMonitorCommand({
  AnomalyMonitor: {
    MonitorName: 'TeamCostMonitor',
    MonitorType: 'CUSTOM',
    MonitorSpecification: JSON.stringify({
      Tags: {
        Key: 'team',
        Values: ['ml-platform', 'data-engineering', 'product-backend'],
        MatchOptions: ['EQUALS'],
      },
    }),
  },
});

const { MonitorArn: teamMonitorArn } = await client.send(teamMonitor);

// Subscription with tiered alerting
const subscription = new CreateAnomalySubscriptionCommand({
  AnomalySubscription: {
    SubscriptionName: 'CostAnomalyAlerts',
    MonitorArnList: [serviceMonitorArn!, teamMonitorArn!],
    Subscribers: [
      {
        Address: 'arn:aws:sns:us-east-1:123456789012:cost-anomaly-alerts',
        Type: 'SNS',
        Status: 'CONFIRMED',
      },
    ],
    Frequency: 'IMMEDIATE',
    ThresholdExpression: {
      Or: [
        {
          Dimensions: {
            Key: 'ANOMALY_TOTAL_IMPACT_ABSOLUTE',
            MatchOptions: ['GREATER_THAN_OR_EQUAL'],
            Values: ['100'],
          },
        },
        {
          Dimensions: {
            Key: 'ANOMALY_TOTAL_IMPACT_PERCENTAGE',
            MatchOptions: ['GREATER_THAN_OR_EQUAL'],
            Values: ['30'],
          },
        },
      ],
    },
  },
});

The threshold expression triggers on either an absolute spend increase of $100+ OR a percentage increase of 30%+ over baseline. This catches both large absolute deviations on high-spend services and proportionally large deviations on low-spend services.

Automated Remediation Workflow

Alerts alone are not enough if nobody acts on them at 2 AM. We built an automated triage and remediation workflow:

import { SNSEvent } from 'aws-lambda';
import { EC2Client, StopInstancesCommand } from '@aws-sdk/client-ec2';
import { RDSClient, StopDBInstanceCommand } from '@aws-sdk/client-rds';

interface AnomalyAlert {
  anomalyId: string;
  monitorArn: string;
  anomalyScore: number;
  impact: {
    maxImpact: number;
    totalImpact: number;
  };
  rootCauses: {
    service: string;
    region: string;
    linkedAccount: string;
    usageType: string;
  }[];
  anomalyStartDate: string;
  anomalyEndDate: string;
}

export async function handler(event: SNSEvent): Promise<void> {
  for (const record of event.Records) {
    const alert: AnomalyAlert = JSON.parse(record.Sns.Message);

    // Classify severity
    const severity = classifySeverity(alert);

    // Auto-remediate known patterns
    if (severity === 'auto-remediate') {
      await autoRemediate(alert);
    }

    // Enrich and route to appropriate channel
    await routeAlert(alert, severity);
  }
}

function classifySeverity(alert: AnomalyAlert): string {
  const { totalImpact } = alert.impact;
  const rootService = alert.rootCauses[0]?.service;

  // GPU instances running outside business hours — likely forgotten
  if (
    rootService === 'Amazon Elastic Compute Cloud' &&
    alert.rootCauses[0]?.usageType?.includes('p4d')
  ) {
    return 'auto-remediate';
  }

  if (totalImpact > 5000) return 'critical';
  if (totalImpact > 1000) return 'high';
  if (totalImpact > 100) return 'medium';
  return 'low';
}

async function autoRemediate(alert: AnomalyAlert): Promise<void> {
  const ec2 = new EC2Client({ region: alert.rootCauses[0].region });

  // Find and stop idle GPU instances tagged as 'experiment'
  // Only stop, never terminate — data preservation
  // Team gets notified and must explicitly restart
  console.log(`Auto-remediating anomaly ${alert.anomalyId}`);
  // Implementation: query for instances matching the anomaly pattern
  // Stop instances tagged with 'environment: experiment'
}

Auto-remediation only triggers for well-understood patterns (GPU instances tagged as experiments). Everything else routes to human review with enriched context.

Alert Configuration Patterns

We defined threshold patterns based on 6 months of historical data:

Monitor ScopeAbsolute ThresholdPercentage ThresholdAlert FrequencyAction
Account total$1,000/day50% over baselineImmediatePage on-call
EC2 compute$200/day40% over baselineImmediateAuto-investigate
RDS$100/day30% over baselineDaily digestTeam review
Data Transfer$50/day100% over baselineImmediateAuto-investigate
Bedrock/AI$500/day25% over baselineImmediatePage ML team
S3 Storage$30/day50% over baselineWeekly digestQuarterly review

Detection Results: Q1 2026

Over the first quarter, anomaly detection identified 34 actionable anomalies:

CategoryCountTotal SavingsAvg Detection Time
Forgotten GPU instances8$18,4004.2 hours
Misconfigured auto-scaling5$12,3006.1 hours
Data transfer explosion7$8,9003.8 hours
Unused NAT Gateways6$4,20012 hours
Bedrock token overuse4$2,8005.4 hours
S3 lifecycle policy gaps4$1,10024 hours
Total34$47,7006.2 hours avg

Without anomaly detection, these would have appeared on the monthly bill 2-4 weeks later. Average savings per anomaly: $1,400. Average detection time: 6.2 hours from spend occurrence.

Cost Anomaly Detection Savings

Lessons Learned

Percentage thresholds catch more than absolute thresholds. A $50/day service jumping to $150/day is a 200% increase and almost certainly a misconfiguration. But $50 absolute would not trigger an account-level alert. Use both.

Suppress known spikes proactively. Before running load tests, deploying new environments, or launching ML training jobs, create temporary anomaly suppressions. Otherwise your team gets alert fatigue from expected spikes.

Tag everything or anomalies are useless. Anomaly detection tells you which service spiked. Without cost allocation tags, you cannot determine which team, project, or environment caused it. We enforce tagging via SCP: resources without team and environment tags cannot be created.

Weekly cost review cadence reduces anomaly count over time. When teams review their spending weekly (not monthly), they catch drift before it becomes an anomaly. Our anomaly count dropped from 34 in Q1 to 19 in Q2 as teams developed better resource lifecycle habits.

Conclusion

AWS Cost Anomaly Detection transforms cloud cost management from reactive invoice review to proactive spend monitoring. The $47,700 we saved in Q1 represents resources that would have run unchecked for weeks without detection. The configuration itself costs nothing — it is included with Cost Explorer. The investment is in building the remediation workflow and establishing the operational cadence to act on alerts. For any organization spending more than $10,000/month on AWS, this is table-stakes infrastructure that pays for the engineering effort within the first detected anomaly.

Comments

    No comments yet. Be the first to share your thoughts.