AWS ElastiCache Redis Cluster Mode: Multi-Tenant Caching at Scale

How we scaled Redis cluster mode to serve 200+ tenants with sub-millisecond latency while reducing cache infrastructure costs by 40%.

#aws#elasticache#redis#caching
Cover image for the article: AWS ElastiCache Redis Cluster Mode: Multi-Tenant Caching at Scale

When your multi-tenant SaaS platform handles 2 million requests per second and every tenant expects sub-millisecond cache responses, single-node Redis becomes a liability. We hit this wall at 850K requests/second — connection limits exhausted, memory fragmentation climbing, and failovers causing tenant-visible latency spikes.

This is how we migrated to ElastiCache Redis Cluster Mode and built a caching architecture that serves 200+ tenants with predictable performance characteristics.

The Problem: Single-Node Redis at Breaking Point

Our initial architecture ran six r6g.2xlarge Redis nodes in a replica configuration. Each tenant's data lived in a dedicated key prefix, and our application layer handled routing. This worked until it didn't.

MetricBefore (Single Node)Breaking PointAfter (Cluster Mode)
Max throughput850K req/sN/A3.2M req/s
P99 latency2.8ms12ms+0.9ms
Memory utilization78%92% (OOM risk)45% avg across shards
Failover time15-30s45s+<1s (automatic)
Monthly cost$4,200N/A$2,520

The breaking point came during a flash sale event when three large tenants simultaneously spiked their cache usage. The primary node hit 92% memory, triggering swap usage that cascaded into latency spikes across all tenants.

Architecture: Cluster Mode with Tenant-Aware Sharding

Redis Cluster Architecture

Redis Cluster Mode distributes data across up to 500 shards, each with its own primary and replica nodes. The critical design decision was our sharding strategy — we needed tenant isolation without sacrificing the benefits of data locality.

Hash Tag Strategy for Tenant Isolation

Redis Cluster uses hash slots (0-16383) to determine data placement. By embedding tenant IDs in hash tags, we ensure all data for a single tenant lives on the same shard:

// Hash tag strategy ensures tenant data locality
class TenantCacheKeyBuilder {
  private readonly tenantId: string;

  constructor(tenantId: string) {
    this.tenantId = tenantId;
  }

  // {tenant:abc123} forces all keys to same hash slot
  buildKey(namespace: string, identifier: string): string {
    return `{tenant:${this.tenantId}}:${namespace}:${identifier}`;
  }

  buildSessionKey(sessionId: string): string {
    return `{tenant:${this.tenantId}}:session:${sessionId}`;
  }

  buildRateLimitKey(endpoint: string, window: string): string {
    return `{tenant:${this.tenantId}}:rl:${endpoint}:${window}`;
  }
}

// Usage ensures all operations for a tenant hit the same shard
const cache = new TenantCacheKeyBuilder('acme-corp');
const userKey = cache.buildKey('user', 'u-789');
const sessionKey = cache.buildSessionKey('sess-456');
// Both keys: {tenant:acme-corp}:... → same hash slot → same shard

This approach gives us multi-key operations (MGET, pipelines, Lua scripts) within a tenant without cross-slot errors, while distributing tenants evenly across shards.

Cluster Configuration with Terraform

resource "aws_elasticache_replication_group" "multi_tenant" {
  replication_group_id       = "prod-mt-cache"
  description                = "Multi-tenant Redis cluster"
  node_type                  = "cache.r7g.xlarge"
  num_node_groups            = 16  # 16 shards
  replicas_per_node_group    = 2   # 2 replicas per shard

  automatic_failover_enabled = true
  multi_az_enabled           = true
  at_rest_encryption_enabled = true
  transit_encryption_enabled = true

  parameter_group_name = aws_elasticache_parameter_group.cluster.name

  subnet_group_name  = aws_elasticache_subnet_group.private.name
  security_group_ids = [aws_security_group.redis_cluster.id]

  # Enable cluster mode
  cluster_mode {
    num_node_groups         = 16
    replicas_per_node_group = 2
  }

  log_delivery_configuration {
    destination      = aws_cloudwatch_log_group.redis_slow.name
    destination_type = "cloudwatch-logs"
    log_format       = "json"
    log_type         = "slow-log"
  }
}

resource "aws_elasticache_parameter_group" "cluster" {
  name   = "prod-mt-cluster-params"
  family = "redis7"

  parameter {
    name  = "cluster-enabled"
    value = "yes"
  }

  parameter {
    name  = "maxmemory-policy"
    value = "volatile-lru"
  }

  parameter {
    name  = "notify-keyspace-events"
    value = "Ex"  # Expired events for TTL monitoring
  }
}

Tenant Isolation: Preventing Noisy Neighbors

The hash tag approach solves data locality, but doesn't prevent a single tenant from consuming disproportionate resources. We implemented a three-layer isolation strategy:

Tenant Isolation Layers

Layer 1: Memory Quotas via Lua Scripts

-- enforce_quota.lua: Atomic quota check + set
local tenant_key = KEYS[1]
local quota_key = KEYS[2]
local value = ARGV[1]
local ttl = tonumber(ARGV[2])
local max_bytes = tonumber(ARGV[3])

-- Check current tenant memory usage
local current_usage = tonumber(redis.call('GET', quota_key) or '0')
local value_size = #value

if (current_usage + value_size) > max_bytes then
  return redis.error_reply('QUOTA_EXCEEDED:' .. current_usage .. ':' .. max_bytes)
end

-- Set the value and update quota tracking
redis.call('SET', tenant_key, value, 'EX', ttl)
redis.call('INCRBY', quota_key, value_size)
redis.call('EXPIRE', quota_key, 3600)  -- Reset quota window hourly

return redis.status_reply('OK')

Layer 2: Connection Pooling per Tenant Tier

We segment connection pools by tenant tier to prevent connection starvation:

Tenant TierMax ConnectionsPipeline DepthTimeout
Enterprise50128500ms
Business2064300ms
Starter532200ms

Layer 3: Shard Allocation Strategy

Large enterprise tenants get dedicated shards. Smaller tenants are bin-packed across shared shards using a consistent hashing algorithm that maintains balance as tenants are added or removed.

Performance Results

After migrating 200+ tenants over a two-week rolling deployment, the results exceeded our targets:

MetricTargetAchieved
P50 latency<0.5ms0.3ms
P99 latency<2ms0.9ms
P99.9 latency<5ms2.1ms
Throughput ceiling2M req/s3.2M req/s
Cross-tenant impact<5%<1% measured
Cost reduction30%40%

Latency Distribution

Operational Lessons

Online resharding is not free. While ElastiCache supports adding shards without downtime, the slot migration process temporarily increases latency for affected keys. Schedule resharding during low-traffic windows and monitor migrate_cached_sockets metrics.

Cluster topology propagation takes time. After a failover, client libraries need 1-3 seconds to refresh their slot maps. Use aggressive retry policies with exponential backoff during this window.

Memory fragmentation differs per shard. Unlike single-node Redis where you monitor one mem_fragmentation_ratio, cluster mode requires per-shard monitoring. We alert at 1.5 ratio per shard, not cluster-wide averages.

Key Takeaways

  1. Hash tags are your tenant isolation primitive — they give you data locality without sacrificing distribution.
  2. Quota enforcement must be atomic — Lua scripts running on the same shard as the data eliminate race conditions.
  3. Connection pooling by tier prevents noisy-neighbor effects at the network layer.
  4. Plan for 3x your current peak — cluster mode's horizontal scaling makes this economically feasible.
  5. Monitor per-shard, not per-cluster — aggregate metrics hide problems until they cascade.

The migration from standalone Redis to Cluster Mode took three weeks of engineering time and paid for itself in infrastructure savings within two months. More importantly, it gave us a caching layer that scales linearly with tenant count — something our sales team appreciates as much as our engineers.

Comments

    No comments yet. Be the first to share your thoughts.