AWS Keyspaces: Serverless Cassandra for High-Write Workloads

Migrating from self-managed Cassandra to AWS Keyspaces for event streaming workloads — throughput benchmarks, CQL compatibility, and cost modeling at 500K writes/sec.

#aws#keyspaces#cassandra#serverless
Cover image for the article: AWS Keyspaces: Serverless Cassandra for High-Write Workloads

Self-managed Cassandra clusters are operationally expensive. Compaction tuning, gossip protocol issues, JVM garbage collection pauses, tombstone accumulation — each requires specialized knowledge that most engineering teams don't have. When your write-heavy workload demands Cassandra's throughput characteristics but your team can't justify a dedicated database engineering function, AWS Keyspaces offers a compelling alternative.

We migrated a 500K writes/second event streaming platform from a 24-node self-managed Cassandra cluster to Keyspaces. The operational burden dropped from 60 engineering hours/month to near zero. The cost story is more nuanced.

The Problem: Cassandra Operations at Scale

Our event streaming platform ingests clickstream, transaction, and system events from 200+ microservices. The access pattern is write-heavy (95% writes, 5% reads) with queries scoped by partition key (tenant_id + event_date):

Operational IssueMonthly ImpactSeverity
Compaction storms (STCS → LCS migration)16h response timeHigh
GC pauses causing coordinator timeouts4-8 incidents/monthMedium
Node replacement (disk failures)8h each, ~2/monthHigh
Schema changes (adding new event types)4h coordination per changeMedium
Repair cycles (anti-entropy)Runs continuously, impacts latencyLow
Capacity planning (adding nodes)2 days planning + executionMedium
Total engineering time~60 hours/month

At $180/hour, operations alone cost $10,800/month — more than the infrastructure itself.

Architecture: Keyspaces for Event Streaming

Keyspaces Architecture

Keyspaces provides a CQL-compatible interface backed by AWS-managed infrastructure. No nodes to manage, no compaction to tune, no repairs to schedule. The architecture simplifies from "application → client → coordinator → replicas → compaction" to "application → Keyspaces endpoint."

Table Design for High-Write Throughput

-- Event streaming table optimized for write throughput
CREATE TABLE events.raw_events (
    tenant_id       TEXT,
    event_date      DATE,
    event_id        TIMEUUID,
    event_type      TEXT,
    source_service  TEXT,
    payload         BLOB,
    metadata        MAP<TEXT, TEXT>,
    processed       BOOLEAN,
    created_at      TIMESTAMP,
    PRIMARY KEY ((tenant_id, event_date), event_id)
) WITH CLUSTERING ORDER BY (event_id DESC)
  AND default_time_to_live = 7776000  -- 90-day retention
  AND CUSTOM_PROPERTIES = {
    'capacity_mode': {
      'throughput_mode': 'PAY_PER_REQUEST'
    }
  };

-- Materialized view for event-type queries (Keyspaces supports these)
CREATE TABLE events.events_by_type (
    tenant_id       TEXT,
    event_type      TEXT,
    event_date      DATE,
    event_id        TIMEUUID,
    source_service  TEXT,
    created_at      TIMESTAMP,
    PRIMARY KEY ((tenant_id, event_type), event_date, event_id)
) WITH CLUSTERING ORDER BY (event_date DESC, event_id DESC)
  AND default_time_to_live = 7776000;

Application-Level Write Client

import { Client, policies, types } from 'cassandra-driver';

// Keyspaces-optimized client configuration
const client = new Client({
  contactPoints: ['cassandra.us-east-1.amazonaws.com'],
  localDataCenter: 'us-east-1',
  port: 9142,
  authProvider: new policies.auth.SigV4AuthProvider({
    region: 'us-east-1',
    accessKeyId: process.env.AWS_ACCESS_KEY_ID!,
    secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY!,
  }),
  sslOptions: {
    rejectUnauthorized: true,
  },
  protocolOptions: { maxVersion: 4 },
  policies: {
    retry: new policies.retry.ExponentialBackoffRetry(5, 100, 5000),
  },
  pooling: {
    coreConnectionsPerHost: {
      [types.distance.local]: 4,
      [types.distance.remote]: 1,
    },
    maxRequestsPerConnection: 32768,
  },
});

// High-throughput write with batching
class EventWriter {
  private buffer: EventRecord[] = [];
  private readonly FLUSH_SIZE = 25; // Keyspaces batch limit
  private readonly FLUSH_INTERVAL_MS = 100;

  private readonly insertQuery = `
    INSERT INTO events.raw_events
      (tenant_id, event_date, event_id, event_type, source_service, payload, metadata, processed, created_at)
    VALUES (?, ?, now(), ?, ?, ?, ?, false, toTimestamp(now()))
  `;

  async write(event: EventRecord): Promise<void> {
    this.buffer.push(event);
    if (this.buffer.length >= this.FLUSH_SIZE) {
      await this.flush();
    }
  }

  private async flush(): Promise<void> {
    const batch = this.buffer.splice(0, this.FLUSH_SIZE);
    if (batch.length === 0) return;

    // Use unlogged batch for same-partition writes (single partition key)
    const partitionGroups = this.groupByPartition(batch);

    const writePromises = Array.from(partitionGroups.entries()).map(
      async ([_, events]) => {
        if (events.length === 1) {
          return client.execute(this.insertQuery, this.toParams(events[0]), {
            prepare: true,
            consistency: types.consistencies.localQuorum,
          });
        }

        // Batch only within same partition (Keyspaces requirement)
        const queries = events.map((e) => ({
          query: this.insertQuery,
          params: this.toParams(e),
        }));
        return client.batch(queries, {
          prepare: true,
          logged: false,
          consistency: types.consistencies.localQuorum,
        });
      }
    );

    await Promise.allSettled(writePromises);
  }

  private groupByPartition(events: EventRecord[]): Map<string, EventRecord[]> {
    const groups = new Map<string, EventRecord[]>();
    for (const event of events) {
      const key = `${event.tenantId}:${event.eventDate}`;
      const group = groups.get(key) || [];
      group.push(event);
      groups.set(key, group);
    }
    return groups;
  }
}

Throughput Benchmarks

We benchmarked Keyspaces against our self-managed Cassandra cluster using production-equivalent traffic patterns:

MetricSelf-Managed (24 nodes)Keyspaces (On-Demand)Keyspaces (Provisioned)
Sustained writes/sec520K480K500K (provisioned cap)
Burst writes/sec520K (no headroom)850K (auto-scales)500K (hard cap)
Write latency P502.1ms4.8ms3.9ms
Write latency P9912ms18ms14ms
Read latency P503.2ms6.1ms5.4ms
Read latency P9925ms32ms28ms

Throughput Comparison

Key observation: Keyspaces latency is 2-3x higher than self-managed Cassandra. For our event streaming use case (async writes, batch reads), this is acceptable. For latency-sensitive request-path operations, this trade-off needs careful evaluation.

CQL Compatibility: What Works and What Doesn't

Keyspaces supports a subset of CQL. We documented every compatibility issue during migration:

FeatureCassandraKeyspacesWorkaround
Logged batchesYesSame-partition onlyGroup by partition key
ALLOW FILTERINGYesNoRedesign queries/add tables
Materialized ViewsYesNo (use client-side)Maintain denormalized tables
User-Defined FunctionsYesNoMove logic to application
COUNTER columnsYesYesWorks as expected
LWT (IF NOT EXISTS)YesYesWorks (higher latency)
TTLYesYesPer-row and table-level
Static columnsYesYesWorks as expected
Collections (SET, LIST, MAP)YesYes (size limits)Cap at 1MB per collection

The biggest migration effort was eliminating ALLOW FILTERING queries. In self-managed Cassandra, teams had grown lazy with partition scans. Keyspaces forces proper data modeling.

Cost Model: On-Demand vs. Provisioned

Capacity ModeWrite Cost (500K/s)Read Cost (25K/s)Storage (50TB)Total/Month
On-Demand$8,100$405$1,250$9,755
Provisioned (reserved)$5,400$270$1,250$6,920
Self-Managed (24 nodes)N/AN/AN/A$7,200 (infra only)

Cost Comparison

The raw infrastructure comparison is misleading. Adding the $10,800/month operational cost to self-managed Cassandra:

  • Self-managed total: $7,200 (infra) + $10,800 (ops) = $18,000/month
  • Keyspaces provisioned: $6,920/month
  • Net savings: $11,080/month (62%)

Migration Strategy

We used a dual-write approach with progressive traffic shift:

// Progressive traffic migration controller
class MigrationController {
  private keyspacesWeight: number = 0; // 0-100

  async handleWrite(event: EventRecord): Promise<void> {
    // Always write to Cassandra during migration
    await this.cassandraWriter.write(event);

    // Progressively write to Keyspaces
    if (Math.random() * 100 < this.keyspacesWeight) {
      try {
        await this.keyspacesWriter.write(event);
      } catch (error) {
        // Log but don't fail — Keyspaces is shadow during migration
        metrics.increment('keyspaces.shadow_write_error');
      }
    }
  }

  async handleRead(query: ReadQuery): Promise<EventRecord[]> {
    if (this.keyspacesWeight >= 100) {
      return this.keyspacesReader.execute(query);
    }
    return this.cassandraReader.execute(query);
  }

  // Called by feature flag system
  setWeight(weight: number): void {
    this.keyspacesWeight = Math.min(100, Math.max(0, weight));
    metrics.gauge('migration.keyspaces_weight', this.keyspacesWeight);
  }
}

Migration timeline:

  • Week 1: 10% shadow writes (validate throughput)
  • Week 2: 50% shadow writes (validate consistency)
  • Week 3: 100% writes to both (validate durability)
  • Week 4: Flip reads to Keyspaces, decommission Cassandra

Key Takeaways

  1. Keyspaces latency is higher than self-managed — 2-3x on P50. Acceptable for async workloads, problematic for request-path queries under 5ms SLA.
  2. The real savings come from eliminating operations — infrastructure costs are similar, but removing 60 engineering hours/month changes the economics completely.
  3. CQL compatibility gaps force better data modeling — no ALLOW FILTERING means you design proper partition keys or don't migrate.
  4. On-demand mode for unpredictable traffic — if your writes spike 3x during events, provisioned mode either wastes capacity or throttles.
  5. Batch writes must be same-partition — cross-partition batches (common in Cassandra) fail silently or error in Keyspaces.

Keyspaces isn't a drop-in replacement for Cassandra. It's a serverless wide-column store that speaks CQL. If your workload fits its constraints — write-heavy, partition-key-scoped queries, tolerance for higher latency — the operational savings are substantial. If you need sub-5ms reads or complex query patterns, keep self-managing.

Comments

    No comments yet. Be the first to share your thoughts.