AWS MemoryDB for Redis: Durable Caching with Microsecond Reads

Replacing Redis + DynamoDB dual-write patterns with MemoryDB — achieving microsecond read latency with full durability and Multi-AZ replication.

#aws#memorydb#redis#durability
Cover image for the article: AWS MemoryDB for Redis: Durable Caching with Microsecond Reads

The classic Redis problem: blazing-fast reads but data evaporates on restart. The classic workaround: dual-write to Redis and a durable store (DynamoDB, PostgreSQL), then hydrate Redis on cold start. This pattern works until it doesn't — write amplification doubles your costs, consistency bugs emerge between stores, and cache warming after failover takes minutes while users experience cold-cache latency.

AWS MemoryDB for Redis eliminates this trade-off. It's Redis with a Multi-AZ transactional log that makes every write durable without sacrificing read performance. We replaced our Redis + DynamoDB dual-write architecture and the results changed how we think about caching versus primary storage.

The Problem: Dual-Write Complexity

Our e-commerce platform maintained shopping carts, session state, and real-time inventory counts in Redis — all data that must survive restarts. The dual-write pattern created cascading problems:

IssueImpactFrequency
Write ordering inconsistencyStale data served~200 events/day
DynamoDB write throttlingRedis ahead of truthDuring sales events
Cache warming after failover3-5 min degraded latencyMonthly failover tests
Dual-write code complexityBug surface areaEvery feature change
Cost of write amplification2x write costsContinuous

The consistency bugs were the worst. A customer adds item to cart (Redis write succeeds), DynamoDB write fails silently, Redis node fails over, cart rehydrates from DynamoDB — item is gone. These edge cases eroded user trust.

Architecture: MemoryDB as Primary Data Store

MemoryDB Architecture

MemoryDB changes the mental model: it's not a cache in front of a database — it IS the database for latency-sensitive workloads. Every write is committed to a Multi-AZ transaction log before acknowledgment, providing the same durability guarantees as RDS.

// Before: Dual-write pattern (error-prone)
class DualWriteCartService {
  async addItem(userId: string, item: CartItem): Promise<void> {
    // Write to Redis (fast)
    await this.redis.hset(`cart:${userId}`, item.sku, JSON.stringify(item));

    // Write to DynamoDB (durable) — what if this fails?
    try {
      await this.dynamo.put({
        TableName: 'carts',
        Item: { userId, sku: item.sku, ...item },
      });
    } catch (error) {
      // Now Redis and DynamoDB are inconsistent
      // Do we rollback Redis? Retry DynamoDB? Queue for later?
      logger.error('DynamoDB write failed, data inconsistency', { userId, item });
      await this.inconsistencyQueue.send({ userId, item, error });
    }
  }
}

// After: MemoryDB as single source of truth
class MemoryDBCartService {
  async addItem(userId: string, item: CartItem): Promise<void> {
    // Single write — durable AND fast
    await this.memorydb.hset(`cart:${userId}`, item.sku, JSON.stringify(item));
    // That's it. Write is durable across AZs. No reconciliation needed.
  }

  async getCart(userId: string): Promise<CartItem[]> {
    const items = await this.memorydb.hgetall(`cart:${userId}`);
    return Object.values(items).map((i) => JSON.parse(i));
  }
}

Cluster Configuration

resource "aws_memorydb_cluster" "primary" {
  name                   = "prod-primary-store"
  node_type              = "db.r7g.xlarge"
  num_shards             = 8
  num_replicas_per_shard = 2
  acl_name               = aws_memorydb_acl.app.name

  tls_enabled            = true
  auto_minor_version_upgrade = true

  snapshot_retention_limit = 7
  snapshot_window          = "03:00-04:00"
  maintenance_window       = "sun:05:00-sun:06:00"

  subnet_group_name  = aws_memorydb_subnet_group.private.name
  security_group_ids = [aws_security_group.memorydb.id]

  parameter_group_name = aws_memorydb_parameter_group.optimized.name
}

resource "aws_memorydb_parameter_group" "optimized" {
  name   = "prod-optimized"
  family = "memorydb_redis7"

  parameter {
    name  = "maxmemory-policy"
    value = "noeviction"  # Critical: no eviction for primary data store
  }

  parameter {
    name  = "activedefrag"
    value = "yes"
  }

  parameter {
    name  = "hz"
    value = "100"  # Higher frequency for faster expired key cleanup
  }
}

Note maxmemory-policy = noeviction — because MemoryDB is our primary store, we never want data silently evicted. Memory exhaustion should trigger an alert, not data loss.

Performance Benchmarks

We benchmarked MemoryDB against our previous ElastiCache + DynamoDB architecture using production traffic patterns:

OperationElastiCache (Read)DynamoDB (Read)MemoryDB (Read)MemoryDB (Write)
Single key GET0.2ms4ms0.2ms1.8ms
Multi-key MGET (10)0.4ms12ms (batch)0.4msN/A
Hash HGETALL0.3ms8ms0.3ms2.1ms
Sorted set range0.5ms15ms (query)0.5ms2.4ms

Latency Comparison

Read latency is identical to ElastiCache — MemoryDB serves reads from memory with no transaction log overhead. Write latency is ~1.5ms higher than ElastiCache (which is non-durable), but dramatically lower than the dual-write roundtrip to DynamoDB.

Throughput Under Load

Concurrent ConnectionsElastiCache Reads/sMemoryDB Reads/sMemoryDB Writes/s
100450K445K120K
5001.2M1.18M280K
1,0001.8M1.75M380K
2,0002.1M2.05M420K

Read throughput difference is within noise. Write throughput is bounded by the transaction log commit — still 420K writes/sec, far exceeding our 85K ops/sec requirement.

Cost Comparison

ComponentBefore (ElastiCache + DynamoDB)After (MemoryDB)
Compute/Memory$2,400/mo (ElastiCache r6g.xlarge × 6)$3,800/mo (MemoryDB r7g.xlarge × 8 shards × 3)
DynamoDB writes$1,800/mo (85K WCU)$0
DynamoDB reads$600/mo (warm-up + fallback)$0
DynamoDB storage$250/mo$0
Reconciliation Lambda$120/mo$0
Total$5,170/mo$3,800/mo
Engineering overhead12h/month (consistency bugs)2h/month (monitoring)

Cost Breakdown

Net savings: $1,370/month in infrastructure plus ~10 engineering hours/month no longer spent debugging consistency issues.

Use Cases That Fit MemoryDB

Not everything should move to MemoryDB. Here's our decision framework:

Use CaseMemoryDB?Reasoning
Shopping cartsYesMust survive restarts, latency-critical
Session stateYesDurable sessions eliminate re-auth on failover
Real-time inventoryYesSingle source of truth, no dual-write
Rate limitingNo (ElastiCache)Ephemeral by nature, eviction is acceptable
Page cachingNo (ElastiCache)Recomputable, cost-sensitive
LeaderboardsYesSorted sets with durability
Pub/Sub messagingNo (ElastiCache)Messages are ephemeral

Operational Considerations

Failover behavior: MemoryDB promotes a replica to primary in under 10 seconds. Unlike ElastiCache where failover means cold cache, MemoryDB replicas have full data — zero warm-up time.

Backup and restore: Point-in-time snapshots to S3 enable cross-region disaster recovery. Restore creates a new cluster from any snapshot within the retention window.

Scaling writes: Adding shards redistributes hash slots automatically. Unlike ElastiCache cluster mode, MemoryDB's slot migration maintains durability guarantees throughout the rebalancing process.

Key Takeaways

  1. MemoryDB eliminates dual-write patterns — single writes that are both fast and durable simplify your architecture dramatically.
  2. Read performance is identical to ElastiCache — the durability layer only affects write latency (adds ~1.5ms).
  3. Use noeviction policy — MemoryDB is a primary store, not a cache. Data loss from eviction is unacceptable.
  4. Cost savings come from removing the durable backend — MemoryDB itself costs more than ElastiCache, but eliminating DynamoDB and reconciliation logic saves net.
  5. Not a replacement for all Redis use cases — ephemeral data (rate limits, page cache) should stay on ElastiCache where eviction is acceptable and costs are lower.

MemoryDB represents a category shift: Redis-compatible, microsecond-read, durable storage. For any workload where you're currently running Redis + a durable backend in parallel, MemoryDB collapses that into a single system with fewer failure modes and lower total cost.

Comments

    No comments yet. Be the first to share your thoughts.