AWS Cost Anomaly Detection: Catching Runaway Spend Before It Hits Your Bill
How we configured Cost Anomaly Detection to identify unexpected spending within 6 hours, saving $47,000 in a single quarter from early detection of misconfigurations.

A developer provisioned a p4d.24xlarge instance for a machine learning experiment on a Friday afternoon. They forgot to terminate it. By Monday morning, the instance had accumulated $2,400 in charges running idle. This was not a one-time incident. Over the previous quarter, we identified $47,000 in wasted spend from forgotten resources, misconfigured auto-scaling, and runaway data transfer — all caught after the monthly bill arrived.
We deployed AWS Cost Anomaly Detection with custom monitors and alert thresholds that catch spending deviations within 6 hours. Combined with automated remediation workflows, we now stop bleeding spend before it compounds.
The Problem: Bills Are Retrospective, Anomalies Are Real-Time
The default AWS billing experience is fundamentally backwards for cost control:
- Billing dashboard updates with 12-24 hour delay
- Budget alerts fire when you have already exceeded the threshold
- Cost Explorer requires manual investigation
- Monthly bills reveal problems 30 days after they started
In a team of 45 engineers across 12 AWS accounts, resources get provisioned daily. Without real-time anomaly detection, you are always discovering last month's mistakes.
Architecture: Layered Anomaly Monitors
AWS Cost Anomaly Detection uses machine learning to establish spending baselines and alert on deviations. We configured three layers of monitors:
Account-level monitors — Catch broad spending spikes across any service Service-level monitors — Detect unusual activity within specific high-cost services (EC2, RDS, Bedrock) Linked account monitors — Track per-team spending against historical patterns
Each layer has different sensitivity thresholds because the blast radius differs. An account-level anomaly at $500 is critical. A service-level anomaly at $50 might indicate a new feature rollout.
Implementation: Custom Monitor Configuration
import {
CostExplorerClient,
CreateAnomalyMonitorCommand,
CreateAnomalySubscriptionCommand,
} from '@aws-sdk/client-cost-explorer';
const client = new CostExplorerClient({ region: 'us-east-1' });
// Service-level monitor for high-cost services
const serviceMonitor = new CreateAnomalyMonitorCommand({
AnomalyMonitor: {
MonitorName: 'HighCostServiceMonitor',
MonitorType: 'DIMENSIONAL',
MonitorDimension: 'SERVICE',
},
});
const { MonitorArn: serviceMonitorArn } = await client.send(serviceMonitor);
// Custom monitor for specific cost allocation tags
const teamMonitor = new CreateAnomalyMonitorCommand({
AnomalyMonitor: {
MonitorName: 'TeamCostMonitor',
MonitorType: 'CUSTOM',
MonitorSpecification: JSON.stringify({
Tags: {
Key: 'team',
Values: ['ml-platform', 'data-engineering', 'product-backend'],
MatchOptions: ['EQUALS'],
},
}),
},
});
const { MonitorArn: teamMonitorArn } = await client.send(teamMonitor);
// Subscription with tiered alerting
const subscription = new CreateAnomalySubscriptionCommand({
AnomalySubscription: {
SubscriptionName: 'CostAnomalyAlerts',
MonitorArnList: [serviceMonitorArn!, teamMonitorArn!],
Subscribers: [
{
Address: 'arn:aws:sns:us-east-1:123456789012:cost-anomaly-alerts',
Type: 'SNS',
Status: 'CONFIRMED',
},
],
Frequency: 'IMMEDIATE',
ThresholdExpression: {
Or: [
{
Dimensions: {
Key: 'ANOMALY_TOTAL_IMPACT_ABSOLUTE',
MatchOptions: ['GREATER_THAN_OR_EQUAL'],
Values: ['100'],
},
},
{
Dimensions: {
Key: 'ANOMALY_TOTAL_IMPACT_PERCENTAGE',
MatchOptions: ['GREATER_THAN_OR_EQUAL'],
Values: ['30'],
},
},
],
},
},
});
The threshold expression triggers on either an absolute spend increase of $100+ OR a percentage increase of 30%+ over baseline. This catches both large absolute deviations on high-spend services and proportionally large deviations on low-spend services.
Automated Remediation Workflow
Alerts alone are not enough if nobody acts on them at 2 AM. We built an automated triage and remediation workflow:
import { SNSEvent } from 'aws-lambda';
import { EC2Client, StopInstancesCommand } from '@aws-sdk/client-ec2';
import { RDSClient, StopDBInstanceCommand } from '@aws-sdk/client-rds';
interface AnomalyAlert {
anomalyId: string;
monitorArn: string;
anomalyScore: number;
impact: {
maxImpact: number;
totalImpact: number;
};
rootCauses: {
service: string;
region: string;
linkedAccount: string;
usageType: string;
}[];
anomalyStartDate: string;
anomalyEndDate: string;
}
export async function handler(event: SNSEvent): Promise<void> {
for (const record of event.Records) {
const alert: AnomalyAlert = JSON.parse(record.Sns.Message);
// Classify severity
const severity = classifySeverity(alert);
// Auto-remediate known patterns
if (severity === 'auto-remediate') {
await autoRemediate(alert);
}
// Enrich and route to appropriate channel
await routeAlert(alert, severity);
}
}
function classifySeverity(alert: AnomalyAlert): string {
const { totalImpact } = alert.impact;
const rootService = alert.rootCauses[0]?.service;
// GPU instances running outside business hours — likely forgotten
if (
rootService === 'Amazon Elastic Compute Cloud' &&
alert.rootCauses[0]?.usageType?.includes('p4d')
) {
return 'auto-remediate';
}
if (totalImpact > 5000) return 'critical';
if (totalImpact > 1000) return 'high';
if (totalImpact > 100) return 'medium';
return 'low';
}
async function autoRemediate(alert: AnomalyAlert): Promise<void> {
const ec2 = new EC2Client({ region: alert.rootCauses[0].region });
// Find and stop idle GPU instances tagged as 'experiment'
// Only stop, never terminate — data preservation
// Team gets notified and must explicitly restart
console.log(`Auto-remediating anomaly ${alert.anomalyId}`);
// Implementation: query for instances matching the anomaly pattern
// Stop instances tagged with 'environment: experiment'
}
Auto-remediation only triggers for well-understood patterns (GPU instances tagged as experiments). Everything else routes to human review with enriched context.
Alert Configuration Patterns
We defined threshold patterns based on 6 months of historical data:
| Monitor Scope | Absolute Threshold | Percentage Threshold | Alert Frequency | Action |
|---|---|---|---|---|
| Account total | $1,000/day | 50% over baseline | Immediate | Page on-call |
| EC2 compute | $200/day | 40% over baseline | Immediate | Auto-investigate |
| RDS | $100/day | 30% over baseline | Daily digest | Team review |
| Data Transfer | $50/day | 100% over baseline | Immediate | Auto-investigate |
| Bedrock/AI | $500/day | 25% over baseline | Immediate | Page ML team |
| S3 Storage | $30/day | 50% over baseline | Weekly digest | Quarterly review |
Detection Results: Q1 2026
Over the first quarter, anomaly detection identified 34 actionable anomalies:
| Category | Count | Total Savings | Avg Detection Time |
|---|---|---|---|
| Forgotten GPU instances | 8 | $18,400 | 4.2 hours |
| Misconfigured auto-scaling | 5 | $12,300 | 6.1 hours |
| Data transfer explosion | 7 | $8,900 | 3.8 hours |
| Unused NAT Gateways | 6 | $4,200 | 12 hours |
| Bedrock token overuse | 4 | $2,800 | 5.4 hours |
| S3 lifecycle policy gaps | 4 | $1,100 | 24 hours |
| Total | 34 | $47,700 | 6.2 hours avg |
Without anomaly detection, these would have appeared on the monthly bill 2-4 weeks later. Average savings per anomaly: $1,400. Average detection time: 6.2 hours from spend occurrence.
Lessons Learned
Percentage thresholds catch more than absolute thresholds. A $50/day service jumping to $150/day is a 200% increase and almost certainly a misconfiguration. But $50 absolute would not trigger an account-level alert. Use both.
Suppress known spikes proactively. Before running load tests, deploying new environments, or launching ML training jobs, create temporary anomaly suppressions. Otherwise your team gets alert fatigue from expected spikes.
Tag everything or anomalies are useless. Anomaly detection tells you which service spiked. Without cost allocation tags, you cannot determine which team, project, or environment caused it. We enforce tagging via SCP: resources without team and environment tags cannot be created.
Weekly cost review cadence reduces anomaly count over time. When teams review their spending weekly (not monthly), they catch drift before it becomes an anomaly. Our anomaly count dropped from 34 in Q1 to 19 in Q2 as teams developed better resource lifecycle habits.
Conclusion
AWS Cost Anomaly Detection transforms cloud cost management from reactive invoice review to proactive spend monitoring. The $47,700 we saved in Q1 represents resources that would have run unchecked for weeks without detection. The configuration itself costs nothing — it is included with Cost Explorer. The investment is in building the remediation workflow and establishing the operational cadence to act on alerts. For any organization spending more than $10,000/month on AWS, this is table-stakes infrastructure that pays for the engineering effort within the first detected anomaly.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.