Zero-Downtime MongoDB to AWS DocumentDB Migration

A battle-tested playbook for migrating 2TB MongoDB clusters to DocumentDB with zero downtime — CDC replication, compatibility testing, and cutover orchestration.

#aws#documentdb#mongodb#migration
Cover image for the article: Zero-Downtime MongoDB to AWS DocumentDB Migration

Migrating a production MongoDB cluster to AWS DocumentDB sounds straightforward until you realize your application uses 47 different MongoDB features, three of which DocumentDB doesn't support, and your SLA requires zero downtime during the cutover. We migrated a 2TB, 85K ops/sec MongoDB deployment serving a fintech platform — and learned every edge case the documentation glosses over.

This is the complete playbook: compatibility auditing, CDC-based replication, dual-write verification, and the 90-second cutover window.

The Problem: Self-Managed MongoDB at Scale

Our MongoDB cluster ran on three r6g.4xlarge EC2 instances in a replica set configuration. The operational burden was significant:

Operational TaskFrequencyEngineering Hours/Month
Patch managementMonthly8h
Backup verificationWeekly4h
Performance tuningOngoing12h
Capacity planningQuarterly6h
Security auditsMonthly4h
Incident responseVariable8-20h
Total42-54h/month

At $180/hour fully loaded engineering cost, we were spending $7,500-$9,700/month just keeping MongoDB running — before infrastructure costs. DocumentDB promised to eliminate most of this operational overhead.

Phase 1: Compatibility Audit

DocumentDB implements the MongoDB 5.0 wire protocol but not every feature. Before writing migration code, we ran a comprehensive compatibility scan:

// DocumentDB compatibility checker
interface CompatibilityResult {
  feature: string;
  supported: boolean;
  workaround?: string;
  affectedCollections: string[];
  queryCount: number; // from profiler data
}

async function auditCompatibility(
  db: Db,
  profilerData: ProfilerEntry[]
): Promise<CompatibilityResult[]> {
  const results: CompatibilityResult[] = [];

  // Check for unsupported index types
  const collections = await db.listCollections().toArray();
  for (const col of collections) {
    const indexes = await db.collection(col.name).indexes();
    for (const idx of indexes) {
      if (idx.key && Object.values(idx.key).includes('text')) {
        results.push({
          feature: 'Text indexes',
          supported: false,
          workaround: 'Migrate to OpenSearch for full-text search',
          affectedCollections: [col.name],
          queryCount: countQueriesUsing(profilerData, col.name, '$text'),
        });
      }
    }
  }

  // Check for $graphLookup usage
  const graphLookupQueries = profilerData.filter(
    (p) => JSON.stringify(p.command).includes('$graphLookup')
  );
  if (graphLookupQueries.length > 0) {
    results.push({
      feature: '$graphLookup aggregation stage',
      supported: false,
      workaround: 'Rewrite as application-level recursive queries or use Neptune',
      affectedCollections: [...new Set(graphLookupQueries.map((q) => q.ns))],
      queryCount: graphLookupQueries.length,
    });
  }

  // Check for client-side field level encryption
  // Check for change stream resume tokens format differences
  // Check for collation-specific behaviors
  // ... 23 more checks

  return results;
}

Our audit revealed three incompatibilities requiring code changes before migration:

FeatureUsageResolutionEffort
$text indexesFull-text search on 2 collectionsMigrated to OpenSearch3 days
$graphLookupOrg hierarchy queriesRewrote as recursive application queries2 days
Capped collectionsEvent log (1 collection)Replaced with TTL index on timestamp1 day

Phase 2: CDC Replication Pipeline

Migration Architecture

We used AWS Database Migration Service (DMS) for continuous replication from MongoDB to DocumentDB. The architecture maintains both databases in sync during the validation period:

// DMS task configuration for MongoDB → DocumentDB CDC
const replicationTask = {
  ReplicationTaskIdentifier: 'mongo-to-docdb-cdc',
  SourceEndpointArn: mongoEndpoint.arn,
  TargetEndpointArn: docdbEndpoint.arn,
  MigrationType: 'full-load-and-cdc',
  TableMappings: JSON.stringify({
    rules: [
      {
        'rule-type': 'selection',
        'rule-id': '1',
        'rule-name': 'include-all',
        'object-locator': {
          'schema-name': 'production',
          'table-name': '%',
        },
        'rule-action': 'include',
      },
      {
        'rule-type': 'selection',
        'rule-id': '2',
        'rule-name': 'exclude-temp',
        'object-locator': {
          'schema-name': 'production',
          'table-name': 'tmp_%',
        },
        'rule-action': 'exclude',
      },
    ],
  }),
  ReplicationTaskSettings: JSON.stringify({
    TargetMetadata: {
      ParallelLoadThreads: 8,
      ParallelLoadBufferSize: 500,
    },
    FullLoadSettings: {
      TargetTablePrepMode: 'DO_NOTHING',
      MaxFullLoadSubTasks: 8,
    },
    Logging: {
      EnableLogging: true,
      LogComponents: [
        { Id: 'SOURCE_UNLOAD', Severity: 'LOGGER_SEVERITY_DEFAULT' },
        { Id: 'TARGET_LOAD', Severity: 'LOGGER_SEVERITY_DEFAULT' },
        { Id: 'CDC', Severity: 'LOGGER_SEVERITY_DEFAULT' },
      ],
    },
  }),
};

Full Load Performance

CollectionDocument CountSizeLoad TimeThroughput
transactions840M1.2TB4.2 hours55K docs/s
accounts12M48GB18 min11K docs/s
audit_logs2.1B680GB3.8 hours153K docs/s
sessions45M12GB4 min187K docs/s

Total full load: 8.5 hours for 2TB across 28 collections.

Phase 3: Dual-Write Verification

After CDC replication stabilized (lag consistently under 500ms), we deployed a dual-write proxy that validated every write against both databases:

// Dual-write verification proxy
class DualWriteProxy {
  private source: MongoClient;
  private target: MongoClient;
  private discrepancies: DiscrepancyLog;

  async write(
    collection: string,
    operation: 'insert' | 'update' | 'delete',
    query: Document,
    doc?: Document
  ): Promise<WriteResult> {
    // Primary write to MongoDB (source of truth)
    const sourceResult = await this.executeWrite(this.source, collection, operation, query, doc);

    // Shadow write to DocumentDB (async, non-blocking)
    setImmediate(async () => {
      try {
        const targetResult = await this.executeWrite(
          this.target, collection, operation, query, doc
        );
        await this.compareResults(collection, operation, sourceResult, targetResult);
      } catch (error) {
        this.discrepancies.log({
          collection,
          operation,
          query,
          error: error.message,
          timestamp: new Date(),
        });
      }
    });

    return sourceResult;
  }

  private async compareResults(
    collection: string,
    operation: string,
    source: WriteResult,
    target: WriteResult
  ): Promise<void> {
    if (source.modifiedCount !== target.modifiedCount) {
      this.discrepancies.log({
        type: 'MODIFIED_COUNT_MISMATCH',
        collection,
        operation,
        source: source.modifiedCount,
        target: target.modifiedCount,
      });
    }
  }
}

We ran dual-write verification for 14 days. The discrepancy rate dropped from 0.003% to 0.0001% after fixing three edge cases related to upsert behavior differences.

Migration Verification Metrics

Phase 4: The 90-Second Cutover

The actual cutover was the shortest phase but required the most coordination:

  1. T-0s: Pause application writes (feature flag)
  2. T-5s: Verify CDC lag is 0
  3. T-10s: Update connection strings in Parameter Store
  4. T-15s: Rolling restart of application pods (30 pods, 3 at a time)
  5. T-60s: All pods connected to DocumentDB
  6. T-75s: Resume writes
  7. T-90s: Verify write success rate > 99.99%
#!/bin/bash
# cutover.sh - Orchestrates the 90-second migration cutover

set -euo pipefail

echo "[$(date)] Starting cutover..."

# Step 1: Enable write pause
aws ssm put-parameter --name "/app/write-pause" --value "true" --overwrite
sleep 5

# Step 2: Wait for CDC lag to reach 0
while true; do
  LAG=$(aws dms describe-replication-tasks \
    --filters Name=replication-task-id,Values=mongo-to-docdb-cdc \
    --query 'ReplicationTasks[0].ReplicationTaskStats.CDCLatencySource' \
    --output text)
  [ "$LAG" = "0" ] && break
  sleep 1
done
echo "[$(date)] CDC lag is 0, proceeding..."

# Step 3: Switch connection string
aws ssm put-parameter --name "/app/mongodb-uri" \
  --value "mongodb://docdb-cluster.cluster-xyz.us-east-1.docdb.amazonaws.com:27017" \
  --overwrite

# Step 4: Rolling restart
kubectl rollout restart deployment/api-server -n production
kubectl rollout status deployment/api-server -n production --timeout=60s

# Step 5: Resume writes
aws ssm put-parameter --name "/app/write-pause" --value "false" --overwrite
echo "[$(date)] Cutover complete!"

Post-Migration Results

MetricMongoDB (Self-Managed)DocumentDBChange
P99 read latency8ms5ms-37%
P99 write latency12ms9ms-25%
Operational hours/month48h6h-87%
Monthly infrastructure cost$6,200$4,800-23%
Monthly ops cost$8,640$1,080-87%
Recovery time (failover)30-60s10-15s-75%

Lessons Learned

Test aggregation pipelines exhaustively. DocumentDB's aggregation engine has subtle differences in sort stability, numeric precision, and NULL handling. We found 4 queries producing different results only under specific data conditions that our unit tests didn't cover.

Index builds behave differently. DocumentDB builds indexes in the foreground by default (background builds available in 5.0 compat mode). Plan index creation during maintenance windows.

Connection handling matters. DocumentDB has stricter connection limits per instance. We needed to tune our connection pool from 100 to 50 connections per pod and enable connection multiplexing.

Key Takeaways

  1. Run the compatibility audit first — it determines whether migration is feasible and scopes the pre-work.
  2. DMS CDC replication is reliable — but monitor for replication lag spikes during high-write periods.
  3. 14 days of dual-write verification catches edge cases that synthetic tests miss.
  4. The cutover itself is the easy part — preparation and verification consume 95% of the effort.
  5. Calculate total cost including ops — DocumentDB's infrastructure cost is only slightly cheaper, but eliminating 42 engineering hours/month is the real ROI.

Zero-downtime migrations aren't magic — they're engineering discipline applied to a well-understood sequence of steps. The key is making the cutover reversible and keeping the old system running until you're confident the new one is production-ready.

Comments

    No comments yet. Be the first to share your thoughts.