Running 70% of Batch Workloads on Fargate Spot: Our Cost Optimization Playbook
How we cut ECS compute costs by 62% by strategically routing batch workloads to Fargate Spot with graceful interruption handling and capacity provider strategies.

The Problem: Fargate Bills That Scale Faster Than Revenue
Our ECS Fargate bill hit $34,000/month and was growing 15% month-over-month. The workloads driving this cost fell into two buckets: long-running API services (30% of spend) and batch/async processing tasks (70% of spend). The batch workloads — data pipelines, report generation, image processing, ML inference jobs — were all running on on-demand Fargate capacity because "it's simpler."
Simpler, yes. But paying full price for interruptible workloads is a luxury that compounds badly at scale.
Architecture: Capacity Provider Strategy
The key architectural decision isn't "use Spot" — it's designing your capacity provider strategy to route the right workloads to Spot while keeping latency-sensitive services on on-demand.
We use a two-cluster architecture:
- Services Cluster — On-demand only. Runs API services, WebSocket handlers, and anything user-facing.
- Workers Cluster — Mixed capacity providers. 70% Spot base with on-demand fallback for capacity-unavailable scenarios.
import * as ecs from 'aws-cdk-lib/aws-ecs';
const workersCluster = new ecs.Cluster(this, 'WorkersCluster', {
vpc,
capacityProviders: ['FARGATE', 'FARGATE_SPOT'],
defaultCapacityProviderStrategy: [
{
capacityProvider: 'FARGATE_SPOT',
weight: 7,
base: 0,
},
{
capacityProvider: 'FARGATE',
weight: 3,
base: 1, // Always keep 1 on-demand task as baseline
},
],
});
The base: 1 on the FARGATE provider ensures at least one task always runs on on-demand — this acts as our canary. If Spot capacity completely evaporates, we always have minimum processing capability.
Graceful Interruption Handling
Fargate Spot tasks receive a SIGTERM 30 seconds before termination. Most teams ignore this signal, letting tasks die mid-execution. We implemented a checkpoint pattern for our data pipeline tasks:
import { setTimeout } from 'timers/promises';
interface CheckpointState {
processedRecords: number;
lastProcessedId: string;
batchId: string;
startedAt: string;
}
class SpotAwareProcessor {
private interrupted = false;
private checkpoint: CheckpointState;
private readonly checkpointStore: DynamoDBClient;
constructor(batchId: string) {
this.checkpoint = {
processedRecords: 0,
lastProcessedId: '',
batchId,
startedAt: new Date().toISOString(),
};
// Listen for SIGTERM (Spot interruption)
process.on('SIGTERM', async () => {
console.log('SIGTERM received — Spot interruption. Saving checkpoint...');
this.interrupted = true;
await this.saveCheckpoint();
process.exit(0);
});
}
async process(records: Record[]): Promise<void> {
// Resume from checkpoint if exists
const existing = await this.loadCheckpoint();
const startIndex = existing ? existing.processedRecords : 0;
for (let i = startIndex; i < records.length; i++) {
if (this.interrupted) break;
await this.processRecord(records[i]);
this.checkpoint.processedRecords = i + 1;
this.checkpoint.lastProcessedId = records[i].id;
// Checkpoint every 100 records
if (i % 100 === 0) {
await this.saveCheckpoint();
}
}
}
}
This pattern ensures that interrupted tasks don't lose progress. When the replacement task starts (either Spot or on-demand), it resumes from the last checkpoint.
Benchmarks: 90-Day Cost Analysis
After 90 days of running our mixed capacity strategy:
| Metric | Before (All On-Demand) | After (70/30 Spot/OD) | Savings |
|---|---|---|---|
| Monthly compute cost | $34,200 | $12,996 | -62% |
| Average task interruption rate | 0% | 4.7% | Expected |
| Task completion time (p50) | 12 min | 12 min | No change |
| Task completion time (p99) | 18 min | 24 min | +33% (retries) |
| Failed tasks (no recovery) | 0.2% | 0.3% | Negligible |
The 62% cost reduction comes from Fargate Spot's 70% discount off on-demand pricing, applied to 70% of our workload. The math: 0.70 * 0.70 = 0.49 savings on the batch portion, plus unchanged on-demand for services.
The p99 latency increase is real but acceptable for batch workloads — interrupted tasks checkpoint and restart, adding one retry cycle to the worst case.
Workload Classification: What Goes on Spot
Not every workload belongs on Spot. Our classification framework:
Spot-safe workloads:
- Data pipeline stages (checkpointable)
- Report generation (idempotent, restartable)
- Image/video processing (chunked, resumable)
- ML batch inference (stateless per-item)
- Log aggregation and ETL
- Scheduled maintenance tasks
On-demand only:
- Real-time API services
- WebSocket connection handlers
- Tasks with external side effects that aren't idempotent
- Tasks with startup times > 5 minutes (amortization issue)
- Anything with strict SLA on completion time
Advanced Pattern: Spot Capacity Monitoring
We built a CloudWatch dashboard that tracks Spot availability and automatically adjusts the capacity provider weights when interruption rates spike:
import { ECSClient, UpdateServiceCommand } from '@aws-sdk/client-ecs';
import { CloudWatchClient, GetMetricDataCommand } from '@aws-sdk/client-cloudwatch';
async function adjustCapacityStrategy(interruptionRate: number) {
const ecs = new ECSClient({});
// If interruption rate exceeds 10%, shift more to on-demand
const spotWeight = interruptionRate > 0.10 ? 5 : interruptionRate > 0.05 ? 6 : 7;
const onDemandWeight = 10 - spotWeight;
await ecs.send(new UpdateServiceCommand({
cluster: 'workers-cluster',
service: 'batch-processor',
capacityProviderStrategy: [
{ capacityProvider: 'FARGATE_SPOT', weight: spotWeight, base: 0 },
{ capacityProvider: 'FARGATE', weight: onDemandWeight, base: 1 },
],
}));
console.log(`Adjusted strategy: Spot=${spotWeight}, OnDemand=${onDemandWeight} (interruption rate: ${(interruptionRate * 100).toFixed(1)}%)`);
}
In practice, we've seen interruption spikes during major AWS events (re:Invent week, ironically) and regional capacity crunches. The auto-adjustment ensures we maintain processing throughput even when Spot capacity is scarce.
Operational Lessons
-
Checkpoint granularity matters — Checkpointing every record is too expensive (DynamoDB writes). Checkpointing every 10,000 records risks too much rework. We landed on every 100 records as the sweet spot.
-
Set stop timeout to 30 seconds — Match your ECS task stop timeout to the Spot warning period. This gives your SIGTERM handler the full 30 seconds to checkpoint.
-
Monitor the FARGATE capacity provider — If your on-demand base tasks are consistently running, it means Spot capacity is unavailable and your costs are higher than expected.
-
Don't use Spot for singleton tasks — If only one instance of a task should run at a time (leader election patterns), Spot interruptions cause unnecessary complexity.
-
Separate Spot and on-demand task definitions — Use different task definitions with appropriate resource limits. Spot tasks can be slightly over-provisioned (you're paying 70% less anyway) to improve checkpoint speed.
Conclusion
Fargate Spot is the highest-ROI cost optimization available to ECS teams — a 70% discount on compute with minimal architectural changes. The key is being deliberate about workload classification and investing in graceful interruption handling.
The 62% cost reduction we achieved wasn't a one-time win — it compounds as workloads grow. Every new batch workload automatically routes to Spot via the capacity provider strategy, so our cost efficiency improves as we scale rather than degrading.
If your ECS bill is dominated by batch processing and you haven't implemented Spot, you're leaving money on the table. Start with your most idempotent workload, add checkpoint logic, and let the savings compound.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.