AWS Keyspaces: Serverless Cassandra for High-Write Workloads
Migrating from self-managed Cassandra to AWS Keyspaces for event streaming workloads — throughput benchmarks, CQL compatibility, and cost modeling at 500K writes/sec.

Self-managed Cassandra clusters are operationally expensive. Compaction tuning, gossip protocol issues, JVM garbage collection pauses, tombstone accumulation — each requires specialized knowledge that most engineering teams don't have. When your write-heavy workload demands Cassandra's throughput characteristics but your team can't justify a dedicated database engineering function, AWS Keyspaces offers a compelling alternative.
We migrated a 500K writes/second event streaming platform from a 24-node self-managed Cassandra cluster to Keyspaces. The operational burden dropped from 60 engineering hours/month to near zero. The cost story is more nuanced.
The Problem: Cassandra Operations at Scale
Our event streaming platform ingests clickstream, transaction, and system events from 200+ microservices. The access pattern is write-heavy (95% writes, 5% reads) with queries scoped by partition key (tenant_id + event_date):
| Operational Issue | Monthly Impact | Severity |
|---|---|---|
| Compaction storms (STCS → LCS migration) | 16h response time | High |
| GC pauses causing coordinator timeouts | 4-8 incidents/month | Medium |
| Node replacement (disk failures) | 8h each, ~2/month | High |
| Schema changes (adding new event types) | 4h coordination per change | Medium |
| Repair cycles (anti-entropy) | Runs continuously, impacts latency | Low |
| Capacity planning (adding nodes) | 2 days planning + execution | Medium |
| Total engineering time | ~60 hours/month |
At $180/hour, operations alone cost $10,800/month — more than the infrastructure itself.
Architecture: Keyspaces for Event Streaming
Keyspaces provides a CQL-compatible interface backed by AWS-managed infrastructure. No nodes to manage, no compaction to tune, no repairs to schedule. The architecture simplifies from "application → client → coordinator → replicas → compaction" to "application → Keyspaces endpoint."
Table Design for High-Write Throughput
-- Event streaming table optimized for write throughput
CREATE TABLE events.raw_events (
tenant_id TEXT,
event_date DATE,
event_id TIMEUUID,
event_type TEXT,
source_service TEXT,
payload BLOB,
metadata MAP<TEXT, TEXT>,
processed BOOLEAN,
created_at TIMESTAMP,
PRIMARY KEY ((tenant_id, event_date), event_id)
) WITH CLUSTERING ORDER BY (event_id DESC)
AND default_time_to_live = 7776000 -- 90-day retention
AND CUSTOM_PROPERTIES = {
'capacity_mode': {
'throughput_mode': 'PAY_PER_REQUEST'
}
};
-- Materialized view for event-type queries (Keyspaces supports these)
CREATE TABLE events.events_by_type (
tenant_id TEXT,
event_type TEXT,
event_date DATE,
event_id TIMEUUID,
source_service TEXT,
created_at TIMESTAMP,
PRIMARY KEY ((tenant_id, event_type), event_date, event_id)
) WITH CLUSTERING ORDER BY (event_date DESC, event_id DESC)
AND default_time_to_live = 7776000;
Application-Level Write Client
import { Client, policies, types } from 'cassandra-driver';
// Keyspaces-optimized client configuration
const client = new Client({
contactPoints: ['cassandra.us-east-1.amazonaws.com'],
localDataCenter: 'us-east-1',
port: 9142,
authProvider: new policies.auth.SigV4AuthProvider({
region: 'us-east-1',
accessKeyId: process.env.AWS_ACCESS_KEY_ID!,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY!,
}),
sslOptions: {
rejectUnauthorized: true,
},
protocolOptions: { maxVersion: 4 },
policies: {
retry: new policies.retry.ExponentialBackoffRetry(5, 100, 5000),
},
pooling: {
coreConnectionsPerHost: {
[types.distance.local]: 4,
[types.distance.remote]: 1,
},
maxRequestsPerConnection: 32768,
},
});
// High-throughput write with batching
class EventWriter {
private buffer: EventRecord[] = [];
private readonly FLUSH_SIZE = 25; // Keyspaces batch limit
private readonly FLUSH_INTERVAL_MS = 100;
private readonly insertQuery = `
INSERT INTO events.raw_events
(tenant_id, event_date, event_id, event_type, source_service, payload, metadata, processed, created_at)
VALUES (?, ?, now(), ?, ?, ?, ?, false, toTimestamp(now()))
`;
async write(event: EventRecord): Promise<void> {
this.buffer.push(event);
if (this.buffer.length >= this.FLUSH_SIZE) {
await this.flush();
}
}
private async flush(): Promise<void> {
const batch = this.buffer.splice(0, this.FLUSH_SIZE);
if (batch.length === 0) return;
// Use unlogged batch for same-partition writes (single partition key)
const partitionGroups = this.groupByPartition(batch);
const writePromises = Array.from(partitionGroups.entries()).map(
async ([_, events]) => {
if (events.length === 1) {
return client.execute(this.insertQuery, this.toParams(events[0]), {
prepare: true,
consistency: types.consistencies.localQuorum,
});
}
// Batch only within same partition (Keyspaces requirement)
const queries = events.map((e) => ({
query: this.insertQuery,
params: this.toParams(e),
}));
return client.batch(queries, {
prepare: true,
logged: false,
consistency: types.consistencies.localQuorum,
});
}
);
await Promise.allSettled(writePromises);
}
private groupByPartition(events: EventRecord[]): Map<string, EventRecord[]> {
const groups = new Map<string, EventRecord[]>();
for (const event of events) {
const key = `${event.tenantId}:${event.eventDate}`;
const group = groups.get(key) || [];
group.push(event);
groups.set(key, group);
}
return groups;
}
}
Throughput Benchmarks
We benchmarked Keyspaces against our self-managed Cassandra cluster using production-equivalent traffic patterns:
| Metric | Self-Managed (24 nodes) | Keyspaces (On-Demand) | Keyspaces (Provisioned) |
|---|---|---|---|
| Sustained writes/sec | 520K | 480K | 500K (provisioned cap) |
| Burst writes/sec | 520K (no headroom) | 850K (auto-scales) | 500K (hard cap) |
| Write latency P50 | 2.1ms | 4.8ms | 3.9ms |
| Write latency P99 | 12ms | 18ms | 14ms |
| Read latency P50 | 3.2ms | 6.1ms | 5.4ms |
| Read latency P99 | 25ms | 32ms | 28ms |
Key observation: Keyspaces latency is 2-3x higher than self-managed Cassandra. For our event streaming use case (async writes, batch reads), this is acceptable. For latency-sensitive request-path operations, this trade-off needs careful evaluation.
CQL Compatibility: What Works and What Doesn't
Keyspaces supports a subset of CQL. We documented every compatibility issue during migration:
| Feature | Cassandra | Keyspaces | Workaround |
|---|---|---|---|
| Logged batches | Yes | Same-partition only | Group by partition key |
| ALLOW FILTERING | Yes | No | Redesign queries/add tables |
| Materialized Views | Yes | No (use client-side) | Maintain denormalized tables |
| User-Defined Functions | Yes | No | Move logic to application |
| COUNTER columns | Yes | Yes | Works as expected |
| LWT (IF NOT EXISTS) | Yes | Yes | Works (higher latency) |
| TTL | Yes | Yes | Per-row and table-level |
| Static columns | Yes | Yes | Works as expected |
| Collections (SET, LIST, MAP) | Yes | Yes (size limits) | Cap at 1MB per collection |
The biggest migration effort was eliminating ALLOW FILTERING queries. In self-managed Cassandra, teams had grown lazy with partition scans. Keyspaces forces proper data modeling.
Cost Model: On-Demand vs. Provisioned
| Capacity Mode | Write Cost (500K/s) | Read Cost (25K/s) | Storage (50TB) | Total/Month |
|---|---|---|---|---|
| On-Demand | $8,100 | $405 | $1,250 | $9,755 |
| Provisioned (reserved) | $5,400 | $270 | $1,250 | $6,920 |
| Self-Managed (24 nodes) | N/A | N/A | N/A | $7,200 (infra only) |
The raw infrastructure comparison is misleading. Adding the $10,800/month operational cost to self-managed Cassandra:
- Self-managed total: $7,200 (infra) + $10,800 (ops) = $18,000/month
- Keyspaces provisioned: $6,920/month
- Net savings: $11,080/month (62%)
Migration Strategy
We used a dual-write approach with progressive traffic shift:
// Progressive traffic migration controller
class MigrationController {
private keyspacesWeight: number = 0; // 0-100
async handleWrite(event: EventRecord): Promise<void> {
// Always write to Cassandra during migration
await this.cassandraWriter.write(event);
// Progressively write to Keyspaces
if (Math.random() * 100 < this.keyspacesWeight) {
try {
await this.keyspacesWriter.write(event);
} catch (error) {
// Log but don't fail — Keyspaces is shadow during migration
metrics.increment('keyspaces.shadow_write_error');
}
}
}
async handleRead(query: ReadQuery): Promise<EventRecord[]> {
if (this.keyspacesWeight >= 100) {
return this.keyspacesReader.execute(query);
}
return this.cassandraReader.execute(query);
}
// Called by feature flag system
setWeight(weight: number): void {
this.keyspacesWeight = Math.min(100, Math.max(0, weight));
metrics.gauge('migration.keyspaces_weight', this.keyspacesWeight);
}
}
Migration timeline:
- Week 1: 10% shadow writes (validate throughput)
- Week 2: 50% shadow writes (validate consistency)
- Week 3: 100% writes to both (validate durability)
- Week 4: Flip reads to Keyspaces, decommission Cassandra
Key Takeaways
- Keyspaces latency is higher than self-managed — 2-3x on P50. Acceptable for async workloads, problematic for request-path queries under 5ms SLA.
- The real savings come from eliminating operations — infrastructure costs are similar, but removing 60 engineering hours/month changes the economics completely.
- CQL compatibility gaps force better data modeling — no
ALLOW FILTERINGmeans you design proper partition keys or don't migrate. - On-demand mode for unpredictable traffic — if your writes spike 3x during events, provisioned mode either wastes capacity or throttles.
- Batch writes must be same-partition — cross-partition batches (common in Cassandra) fail silently or error in Keyspaces.
Keyspaces isn't a drop-in replacement for Cassandra. It's a serverless wide-column store that speaks CQL. If your workload fits its constraints — write-heavy, partition-key-scoped queries, tolerance for higher latency — the operational savings are substantial. If you need sub-5ms reads or complex query patterns, keep self-managing.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.