AWS SQS FIFO Queues: Achieving Exactly-Once Processing at Scale

How we implemented exactly-once message processing using SQS FIFO queues with deduplication and ordering guarantees handling 4.2 million messages daily.

#aws#sqs#messaging#reliability
Cover image for the article: AWS SQS FIFO Queues: Achieving Exactly-Once Processing at Scale

Our payment reconciliation system had a subtle but expensive bug. During network retries, the same payment confirmation message would occasionally be processed twice, resulting in double credits to merchant accounts. Over three months, we identified $23,000 in duplicate payments that required manual clawback. The root cause: standard SQS queues provide at-least-once delivery, and our consumer was not idempotent.

We migrated the payment pipeline to SQS FIFO queues with message deduplication and group-based ordering. Duplicate processing dropped to zero, message ordering guarantees eliminated race conditions in balance calculations, and we maintained throughput at 4.2 million messages per day.

The Problem: At-Least-Once Is Not Enough for Financial Operations

Standard SQS queues make two guarantees: messages will be delivered, and they might be delivered more than once. For most workloads, building idempotent consumers handles the duplicate delivery. But financial operations have additional requirements:

  • Exactly-once processing — A payment credit must apply exactly once
  • Ordering guarantees — Balance operations for the same account must process sequentially
  • Deduplication at source — Retry storms from upstream services must not create duplicate messages

Our standard queue setup had no answer for ordering. Two messages for the same merchant account could be processed concurrently by different Lambda invocations, creating race conditions in balance calculations.

SQS FIFO Architecture

Architecture: FIFO Queues with Message Groups

SQS FIFO queues provide two critical features:

Message Deduplication — Messages with the same deduplication ID within a 5-minute window are delivered only once, regardless of how many times the producer sends them.

Message Grouping — Messages with the same group ID are processed in strict FIFO order. Only one consumer processes messages from a given group at any time.

We designed our system with merchant_id as the message group ID. This guarantees all payment operations for a single merchant process sequentially while different merchants process in parallel.

Implementation: Producer with Deduplication

import { SQSClient, SendMessageCommand } from '@aws-sdk/client-sqs';
import { createHash } from 'crypto';

interface PaymentEvent {
  paymentId: string;
  merchantId: string;
  amount: number;
  currency: string;
  type: 'credit' | 'debit' | 'refund';
  idempotencyKey: string;
  timestamp: string;
}

class PaymentQueueProducer {
  private readonly client: SQSClient;
  private readonly queueUrl: string;

  constructor(queueUrl: string) {
    this.client = new SQSClient({ region: 'us-east-1' });
    this.queueUrl = queueUrl;
  }

  async publishPaymentEvent(event: PaymentEvent): Promise<string> {
    // Deduplication ID from the payment's idempotency key
    // This ensures retries from the payment gateway don't create duplicates
    const deduplicationId = this.generateDeduplicationId(event);

    const command = new SendMessageCommand({
      QueueUrl: this.queueUrl,
      MessageBody: JSON.stringify(event),
      MessageGroupId: event.merchantId,
      MessageDeduplicationId: deduplicationId,
      MessageAttributes: {
        EventType: {
          DataType: 'String',
          StringValue: event.type,
        },
        Priority: {
          DataType: 'Number',
          StringValue: event.type === 'refund' ? '1' : '5',
        },
      },
    });

    const response = await this.client.send(command);
    return response.MessageId!;
  }

  private generateDeduplicationId(event: PaymentEvent): string {
    // Combine idempotency key with payment type to ensure
    // a credit and refund for the same payment are distinct messages
    const composite = `${event.idempotencyKey}:${event.type}:${event.amount}`;
    return createHash('sha256').update(composite).digest('hex').slice(0, 128);
  }
}

The deduplication ID design is critical. We hash the payment's idempotency key combined with the operation type. This means:

  • Retries of the same credit produce the same dedup ID (deduplicated)
  • A credit and subsequent refund for the same payment produce different dedup IDs (both processed)
  • Amount is included to catch edge cases where a retry has a modified amount

Consumer: Ordered Processing with Concurrency Control

import { SQSEvent, SQSRecord } from 'aws-lambda';

interface ProcessingResult {
  success: boolean;
  paymentId: string;
  newBalance?: number;
  error?: string;
}

export async function handler(event: SQSEvent): Promise<{
  batchItemFailures: { itemIdentifier: string }[];
}> {
  const failures: { itemIdentifier: string }[] = [];

  // FIFO queues deliver messages in order within a message group
  // Lambda receives messages from the SAME group in order
  // Process sequentially to maintain ordering guarantees
  for (const record of event.Records) {
    const result = await processPaymentRecord(record);

    if (!result.success) {
      // Report failure — SQS will retry this message and all subsequent
      // messages in the same group (maintains ordering)
      failures.push({ itemIdentifier: record.messageId });

      // Once a message fails, all subsequent messages in the batch
      // from the same group must also be reported as failures
      // to maintain strict ordering
      const failedGroupId = record.attributes.MessageGroupId;
      for (const remaining of event.Records) {
        if (
          remaining.attributes.MessageGroupId === failedGroupId &&
          remaining.messageId !== record.messageId
        ) {
          failures.push({ itemIdentifier: remaining.messageId });
        }
      }
      break;
    }
  }

  return { batchItemFailures: failures };
}

async function processPaymentRecord(
  record: SQSRecord
): Promise<ProcessingResult> {
  const event: PaymentEvent = JSON.parse(record.body);

  // Additional application-level idempotency check
  // Belt and suspenders — FIFO dedup handles the 5-min window,
  // this handles replays beyond that window
  const alreadyProcessed = await checkIdempotencyStore(event.idempotencyKey);
  if (alreadyProcessed) {
    return { success: true, paymentId: event.paymentId };
  }

  try {
    const newBalance = await applyPaymentToBalance(event);
    await recordIdempotencyKey(event.idempotencyKey, newBalance);

    return {
      success: true,
      paymentId: event.paymentId,
      newBalance,
    };
  } catch (error) {
    return {
      success: false,
      paymentId: event.paymentId,
      error: (error as Error).message,
    };
  }
}

The partial batch failure response (batchItemFailures) is essential. When a message fails, we must report all subsequent messages from the same group as failures too. Otherwise SQS might acknowledge later messages, breaking the ordering guarantee.

Throughput: Overcoming the 300 TPS Limit

FIFO queues have a base limit of 300 messages per second (3,000 with high throughput mode). Our 4.2 million daily messages average 48 messages/second but peak at 850 during reconciliation windows. High throughput FIFO mode was necessary:

ConfigurationThroughput LimitOur Usage
Standard FIFO300 msg/sInsufficient
High Throughput FIFO3,000 msg/s per group prefixSufficient
Batched sends (10 msg/batch)30,000 msg/s effectiveHeadroom

High throughput mode partitions message groups internally. The tradeoff: messages across different group IDs may not maintain strict global ordering (only per-group ordering is guaranteed). This is acceptable for our use case since we only need ordering within a single merchant's operations.

Deduplication Window: The 5-Minute Limitation

FIFO deduplication works within a 5-minute window. If the same message is sent 6 minutes apart, it will be processed twice. For payment retries that span longer windows, we added application-level deduplication:

import { DynamoDBClient, PutItemCommand } from '@aws-sdk/client-dynamodb';

const IDEMPOTENCY_TTL_HOURS = 24;

async function checkIdempotencyStore(key: string): Promise<boolean> {
  const client = new DynamoDBClient({ region: 'us-east-1' });
  try {
    await client.send(
      new PutItemCommand({
        TableName: 'PaymentIdempotency',
        Item: {
          idempotencyKey: { S: key },
          processedAt: { N: String(Date.now()) },
          ttl: {
            N: String(
              Math.floor(Date.now() / 1000) + IDEMPOTENCY_TTL_HOURS * 3600
            ),
          },
        },
        ConditionExpression: 'attribute_not_exists(idempotencyKey)',
      })
    );
    return false; // New key, not yet processed
  } catch (error: any) {
    if (error.name === 'ConditionalCheckFailedException') {
      return true; // Already processed
    }
    throw error;
  }
}

This DynamoDB-backed idempotency store provides a 24-hour deduplication window, extending far beyond SQS's 5-minute built-in window.

Results After 90 Days

MetricStandard SQSFIFO SQSChange
Duplicate payment processing34 incidents/quarter0Eliminated
Financial impact of duplicates$23,000/quarter$0Eliminated
Message ordering violations~200/day0Eliminated
Daily message volume4.2M4.2MNo change
P99 processing latency145ms168ms+15.9%
Monthly queue cost$12.40$24.802x
Manual reconciliation hours12 hrs/week0.5 hrs/week-95.8%

The 15.9% latency increase is the cost of ordering guarantees. Messages within the same group wait for the previous message to complete processing. For payment operations requiring consistency, this tradeoff is obvious.

FIFO Queue Processing Metrics

Lessons Learned

Message group ID cardinality matters. Too few groups create hot partitions (one slow merchant blocks all merchants in that group). Too many groups reduce the benefit of high throughput mode. We use merchant_id which gives us ~8,000 unique groups — enough parallelism without fragmentation.

Dead letter queues need special handling for FIFO. When a message goes to the DLQ after max retries, subsequent messages in that group continue processing. This can violate ordering assumptions. We pause the entire group when a DLQ redrive is detected until manual review completes.

Batch size affects ordering semantics. With batch size > 1, Lambda receives multiple messages from the same group. They arrive in order, but you must process them sequentially — not in parallel. We set batch size to 5 and process within a group sequentially.

Content-based deduplication is a trap for complex events. SQS can auto-generate dedup IDs from message body hash. But if your event includes a timestamp that changes on retry, every retry looks unique. Always use explicit deduplication IDs based on business-meaningful keys.

Conclusion

SQS FIFO queues solved our duplicate payment processing problem completely. The combination of message deduplication (eliminating duplicates at the queue level), message grouping (guaranteeing per-entity ordering), and application-level idempotency (handling edge cases beyond the 5-minute window) creates a system where exactly-once semantics are achievable in practice. The $12.40/month cost increase and 16% latency overhead are trivial compared to eliminating $23,000/quarter in duplicate payments and 12 hours/week of manual reconciliation.

Comments

    No comments yet. Be the first to share your thoughts.