AWS MemoryDB for Redis: Durable Caching with Microsecond Reads
Replacing Redis + DynamoDB dual-write patterns with MemoryDB — achieving microsecond read latency with full durability and Multi-AZ replication.

The classic Redis problem: blazing-fast reads but data evaporates on restart. The classic workaround: dual-write to Redis and a durable store (DynamoDB, PostgreSQL), then hydrate Redis on cold start. This pattern works until it doesn't — write amplification doubles your costs, consistency bugs emerge between stores, and cache warming after failover takes minutes while users experience cold-cache latency.
AWS MemoryDB for Redis eliminates this trade-off. It's Redis with a Multi-AZ transactional log that makes every write durable without sacrificing read performance. We replaced our Redis + DynamoDB dual-write architecture and the results changed how we think about caching versus primary storage.
The Problem: Dual-Write Complexity
Our e-commerce platform maintained shopping carts, session state, and real-time inventory counts in Redis — all data that must survive restarts. The dual-write pattern created cascading problems:
| Issue | Impact | Frequency |
|---|---|---|
| Write ordering inconsistency | Stale data served | ~200 events/day |
| DynamoDB write throttling | Redis ahead of truth | During sales events |
| Cache warming after failover | 3-5 min degraded latency | Monthly failover tests |
| Dual-write code complexity | Bug surface area | Every feature change |
| Cost of write amplification | 2x write costs | Continuous |
The consistency bugs were the worst. A customer adds item to cart (Redis write succeeds), DynamoDB write fails silently, Redis node fails over, cart rehydrates from DynamoDB — item is gone. These edge cases eroded user trust.
Architecture: MemoryDB as Primary Data Store
MemoryDB changes the mental model: it's not a cache in front of a database — it IS the database for latency-sensitive workloads. Every write is committed to a Multi-AZ transaction log before acknowledgment, providing the same durability guarantees as RDS.
// Before: Dual-write pattern (error-prone)
class DualWriteCartService {
async addItem(userId: string, item: CartItem): Promise<void> {
// Write to Redis (fast)
await this.redis.hset(`cart:${userId}`, item.sku, JSON.stringify(item));
// Write to DynamoDB (durable) — what if this fails?
try {
await this.dynamo.put({
TableName: 'carts',
Item: { userId, sku: item.sku, ...item },
});
} catch (error) {
// Now Redis and DynamoDB are inconsistent
// Do we rollback Redis? Retry DynamoDB? Queue for later?
logger.error('DynamoDB write failed, data inconsistency', { userId, item });
await this.inconsistencyQueue.send({ userId, item, error });
}
}
}
// After: MemoryDB as single source of truth
class MemoryDBCartService {
async addItem(userId: string, item: CartItem): Promise<void> {
// Single write — durable AND fast
await this.memorydb.hset(`cart:${userId}`, item.sku, JSON.stringify(item));
// That's it. Write is durable across AZs. No reconciliation needed.
}
async getCart(userId: string): Promise<CartItem[]> {
const items = await this.memorydb.hgetall(`cart:${userId}`);
return Object.values(items).map((i) => JSON.parse(i));
}
}
Cluster Configuration
resource "aws_memorydb_cluster" "primary" {
name = "prod-primary-store"
node_type = "db.r7g.xlarge"
num_shards = 8
num_replicas_per_shard = 2
acl_name = aws_memorydb_acl.app.name
tls_enabled = true
auto_minor_version_upgrade = true
snapshot_retention_limit = 7
snapshot_window = "03:00-04:00"
maintenance_window = "sun:05:00-sun:06:00"
subnet_group_name = aws_memorydb_subnet_group.private.name
security_group_ids = [aws_security_group.memorydb.id]
parameter_group_name = aws_memorydb_parameter_group.optimized.name
}
resource "aws_memorydb_parameter_group" "optimized" {
name = "prod-optimized"
family = "memorydb_redis7"
parameter {
name = "maxmemory-policy"
value = "noeviction" # Critical: no eviction for primary data store
}
parameter {
name = "activedefrag"
value = "yes"
}
parameter {
name = "hz"
value = "100" # Higher frequency for faster expired key cleanup
}
}
Note maxmemory-policy = noeviction — because MemoryDB is our primary store, we never want data silently evicted. Memory exhaustion should trigger an alert, not data loss.
Performance Benchmarks
We benchmarked MemoryDB against our previous ElastiCache + DynamoDB architecture using production traffic patterns:
| Operation | ElastiCache (Read) | DynamoDB (Read) | MemoryDB (Read) | MemoryDB (Write) |
|---|---|---|---|---|
| Single key GET | 0.2ms | 4ms | 0.2ms | 1.8ms |
| Multi-key MGET (10) | 0.4ms | 12ms (batch) | 0.4ms | N/A |
| Hash HGETALL | 0.3ms | 8ms | 0.3ms | 2.1ms |
| Sorted set range | 0.5ms | 15ms (query) | 0.5ms | 2.4ms |
Read latency is identical to ElastiCache — MemoryDB serves reads from memory with no transaction log overhead. Write latency is ~1.5ms higher than ElastiCache (which is non-durable), but dramatically lower than the dual-write roundtrip to DynamoDB.
Throughput Under Load
| Concurrent Connections | ElastiCache Reads/s | MemoryDB Reads/s | MemoryDB Writes/s |
|---|---|---|---|
| 100 | 450K | 445K | 120K |
| 500 | 1.2M | 1.18M | 280K |
| 1,000 | 1.8M | 1.75M | 380K |
| 2,000 | 2.1M | 2.05M | 420K |
Read throughput difference is within noise. Write throughput is bounded by the transaction log commit — still 420K writes/sec, far exceeding our 85K ops/sec requirement.
Cost Comparison
| Component | Before (ElastiCache + DynamoDB) | After (MemoryDB) |
|---|---|---|
| Compute/Memory | $2,400/mo (ElastiCache r6g.xlarge × 6) | $3,800/mo (MemoryDB r7g.xlarge × 8 shards × 3) |
| DynamoDB writes | $1,800/mo (85K WCU) | $0 |
| DynamoDB reads | $600/mo (warm-up + fallback) | $0 |
| DynamoDB storage | $250/mo | $0 |
| Reconciliation Lambda | $120/mo | $0 |
| Total | $5,170/mo | $3,800/mo |
| Engineering overhead | 12h/month (consistency bugs) | 2h/month (monitoring) |
Net savings: $1,370/month in infrastructure plus ~10 engineering hours/month no longer spent debugging consistency issues.
Use Cases That Fit MemoryDB
Not everything should move to MemoryDB. Here's our decision framework:
| Use Case | MemoryDB? | Reasoning |
|---|---|---|
| Shopping carts | Yes | Must survive restarts, latency-critical |
| Session state | Yes | Durable sessions eliminate re-auth on failover |
| Real-time inventory | Yes | Single source of truth, no dual-write |
| Rate limiting | No (ElastiCache) | Ephemeral by nature, eviction is acceptable |
| Page caching | No (ElastiCache) | Recomputable, cost-sensitive |
| Leaderboards | Yes | Sorted sets with durability |
| Pub/Sub messaging | No (ElastiCache) | Messages are ephemeral |
Operational Considerations
Failover behavior: MemoryDB promotes a replica to primary in under 10 seconds. Unlike ElastiCache where failover means cold cache, MemoryDB replicas have full data — zero warm-up time.
Backup and restore: Point-in-time snapshots to S3 enable cross-region disaster recovery. Restore creates a new cluster from any snapshot within the retention window.
Scaling writes: Adding shards redistributes hash slots automatically. Unlike ElastiCache cluster mode, MemoryDB's slot migration maintains durability guarantees throughout the rebalancing process.
Key Takeaways
- MemoryDB eliminates dual-write patterns — single writes that are both fast and durable simplify your architecture dramatically.
- Read performance is identical to ElastiCache — the durability layer only affects write latency (adds ~1.5ms).
- Use
noevictionpolicy — MemoryDB is a primary store, not a cache. Data loss from eviction is unacceptable. - Cost savings come from removing the durable backend — MemoryDB itself costs more than ElastiCache, but eliminating DynamoDB and reconciliation logic saves net.
- Not a replacement for all Redis use cases — ephemeral data (rate limits, page cache) should stay on ElastiCache where eviction is acceptable and costs are lower.
MemoryDB represents a category shift: Redis-compatible, microsecond-read, durable storage. For any workload where you're currently running Redis + a durable backend in parallel, MemoryDB collapses that into a single system with fewer failure modes and lower total cost.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.