Running 70% of Batch Workloads on Fargate Spot: Our Cost Optimization Playbook

How we cut ECS compute costs by 62% by strategically routing batch workloads to Fargate Spot with graceful interruption handling and capacity provider strategies.

#aws#fargate#spot#cost-optimization
Cover image for the article: Running 70% of Batch Workloads on Fargate Spot: Our Cost Optimization Playbook

The Problem: Fargate Bills That Scale Faster Than Revenue

Our ECS Fargate bill hit $34,000/month and was growing 15% month-over-month. The workloads driving this cost fell into two buckets: long-running API services (30% of spend) and batch/async processing tasks (70% of spend). The batch workloads — data pipelines, report generation, image processing, ML inference jobs — were all running on on-demand Fargate capacity because "it's simpler."

Simpler, yes. But paying full price for interruptible workloads is a luxury that compounds badly at scale.

Fargate Spot Architecture

Architecture: Capacity Provider Strategy

The key architectural decision isn't "use Spot" — it's designing your capacity provider strategy to route the right workloads to Spot while keeping latency-sensitive services on on-demand.

We use a two-cluster architecture:

  1. Services Cluster — On-demand only. Runs API services, WebSocket handlers, and anything user-facing.
  2. Workers Cluster — Mixed capacity providers. 70% Spot base with on-demand fallback for capacity-unavailable scenarios.
import * as ecs from 'aws-cdk-lib/aws-ecs';

const workersCluster = new ecs.Cluster(this, 'WorkersCluster', {
  vpc,
  capacityProviders: ['FARGATE', 'FARGATE_SPOT'],
  defaultCapacityProviderStrategy: [
    {
      capacityProvider: 'FARGATE_SPOT',
      weight: 7,
      base: 0,
    },
    {
      capacityProvider: 'FARGATE',
      weight: 3,
      base: 1, // Always keep 1 on-demand task as baseline
    },
  ],
});

The base: 1 on the FARGATE provider ensures at least one task always runs on on-demand — this acts as our canary. If Spot capacity completely evaporates, we always have minimum processing capability.

Graceful Interruption Handling

Fargate Spot tasks receive a SIGTERM 30 seconds before termination. Most teams ignore this signal, letting tasks die mid-execution. We implemented a checkpoint pattern for our data pipeline tasks:

import { setTimeout } from 'timers/promises';

interface CheckpointState {
  processedRecords: number;
  lastProcessedId: string;
  batchId: string;
  startedAt: string;
}

class SpotAwareProcessor {
  private interrupted = false;
  private checkpoint: CheckpointState;
  private readonly checkpointStore: DynamoDBClient;

  constructor(batchId: string) {
    this.checkpoint = {
      processedRecords: 0,
      lastProcessedId: '',
      batchId,
      startedAt: new Date().toISOString(),
    };

    // Listen for SIGTERM (Spot interruption)
    process.on('SIGTERM', async () => {
      console.log('SIGTERM received — Spot interruption. Saving checkpoint...');
      this.interrupted = true;
      await this.saveCheckpoint();
      process.exit(0);
    });
  }

  async process(records: Record[]): Promise<void> {
    // Resume from checkpoint if exists
    const existing = await this.loadCheckpoint();
    const startIndex = existing ? existing.processedRecords : 0;

    for (let i = startIndex; i < records.length; i++) {
      if (this.interrupted) break;

      await this.processRecord(records[i]);
      this.checkpoint.processedRecords = i + 1;
      this.checkpoint.lastProcessedId = records[i].id;

      // Checkpoint every 100 records
      if (i % 100 === 0) {
        await this.saveCheckpoint();
      }
    }
  }
}

This pattern ensures that interrupted tasks don't lose progress. When the replacement task starts (either Spot or on-demand), it resumes from the last checkpoint.

Benchmarks: 90-Day Cost Analysis

After 90 days of running our mixed capacity strategy:

MetricBefore (All On-Demand)After (70/30 Spot/OD)Savings
Monthly compute cost$34,200$12,996-62%
Average task interruption rate0%4.7%Expected
Task completion time (p50)12 min12 minNo change
Task completion time (p99)18 min24 min+33% (retries)
Failed tasks (no recovery)0.2%0.3%Negligible

The 62% cost reduction comes from Fargate Spot's 70% discount off on-demand pricing, applied to 70% of our workload. The math: 0.70 * 0.70 = 0.49 savings on the batch portion, plus unchanged on-demand for services.

The p99 latency increase is real but acceptable for batch workloads — interrupted tasks checkpoint and restart, adding one retry cycle to the worst case.

Workload Classification: What Goes on Spot

Not every workload belongs on Spot. Our classification framework:

Spot-safe workloads:

  • Data pipeline stages (checkpointable)
  • Report generation (idempotent, restartable)
  • Image/video processing (chunked, resumable)
  • ML batch inference (stateless per-item)
  • Log aggregation and ETL
  • Scheduled maintenance tasks

On-demand only:

  • Real-time API services
  • WebSocket connection handlers
  • Tasks with external side effects that aren't idempotent
  • Tasks with startup times > 5 minutes (amortization issue)
  • Anything with strict SLA on completion time

Advanced Pattern: Spot Capacity Monitoring

We built a CloudWatch dashboard that tracks Spot availability and automatically adjusts the capacity provider weights when interruption rates spike:

import { ECSClient, UpdateServiceCommand } from '@aws-sdk/client-ecs';
import { CloudWatchClient, GetMetricDataCommand } from '@aws-sdk/client-cloudwatch';

async function adjustCapacityStrategy(interruptionRate: number) {
  const ecs = new ECSClient({});

  // If interruption rate exceeds 10%, shift more to on-demand
  const spotWeight = interruptionRate > 0.10 ? 5 : interruptionRate > 0.05 ? 6 : 7;
  const onDemandWeight = 10 - spotWeight;

  await ecs.send(new UpdateServiceCommand({
    cluster: 'workers-cluster',
    service: 'batch-processor',
    capacityProviderStrategy: [
      { capacityProvider: 'FARGATE_SPOT', weight: spotWeight, base: 0 },
      { capacityProvider: 'FARGATE', weight: onDemandWeight, base: 1 },
    ],
  }));

  console.log(`Adjusted strategy: Spot=${spotWeight}, OnDemand=${onDemandWeight} (interruption rate: ${(interruptionRate * 100).toFixed(1)}%)`);
}

In practice, we've seen interruption spikes during major AWS events (re:Invent week, ironically) and regional capacity crunches. The auto-adjustment ensures we maintain processing throughput even when Spot capacity is scarce.

Operational Lessons

  1. Checkpoint granularity matters — Checkpointing every record is too expensive (DynamoDB writes). Checkpointing every 10,000 records risks too much rework. We landed on every 100 records as the sweet spot.

  2. Set stop timeout to 30 seconds — Match your ECS task stop timeout to the Spot warning period. This gives your SIGTERM handler the full 30 seconds to checkpoint.

  3. Monitor the FARGATE capacity provider — If your on-demand base tasks are consistently running, it means Spot capacity is unavailable and your costs are higher than expected.

  4. Don't use Spot for singleton tasks — If only one instance of a task should run at a time (leader election patterns), Spot interruptions cause unnecessary complexity.

  5. Separate Spot and on-demand task definitions — Use different task definitions with appropriate resource limits. Spot tasks can be slightly over-provisioned (you're paying 70% less anyway) to improve checkpoint speed.

Conclusion

Fargate Spot is the highest-ROI cost optimization available to ECS teams — a 70% discount on compute with minimal architectural changes. The key is being deliberate about workload classification and investing in graceful interruption handling.

The 62% cost reduction we achieved wasn't a one-time win — it compounds as workloads grow. Every new batch workload automatically routes to Spot via the capacity provider strategy, so our cost efficiency improves as we scale rather than degrading.

If your ECS bill is dominated by batch processing and you haven't implemented Spot, you're leaving money on the table. Start with your most idempotent workload, add checkpoint logic, and let the savings compound.

Comments

    No comments yet. Be the first to share your thoughts.