Firestore vs Bigtable: A Decision Matrix for High-Scale GCP Applications

When to choose Firestore over Bigtable (and vice versa) with a practical decision matrix based on access patterns, consistency needs, and cost at scale.

#gcp#firestore#bigtable#database
Cover image for the article: Firestore vs Bigtable: A Decision Matrix for High-Scale GCP Applications

"Should we use Firestore or Bigtable?" is the GCP equivalent of "should we use DynamoDB or Cassandra?" — and the answer is almost never obvious from the documentation alone. Both handle massive scale. Both are NoSQL. Both are fully managed. But they're architecturally different systems optimized for different access patterns.

After running both in production — Firestore for our customer-facing application layer, Bigtable for our analytics and time-series pipeline — I've built a decision framework that cuts through the marketing overlap.

The Core Architectural Difference

Firestore is a document database with strong consistency, real-time subscriptions, and a hierarchical data model. Think of it as a managed, serverless replacement for MongoDB with transactions.

Bigtable is a wide-column store optimized for sequential reads and writes at extreme throughput. Think of it as managed HBase — designed for time-series, analytics, and workloads measured in millions of operations per second.

Firestore vs Bigtable Architecture

The Decision Matrix

DimensionFirestore WinsBigtable Wins
ConsistencyStrong consistency requiredEventual consistency acceptable
Access PatternRandom point reads by document IDSequential range scans, time-series
Query FlexibilityNeed secondary indexes, compound queriesSingle row key design, no secondary indexes
Throughput< 100K ops/sec> 100K ops/sec, scaling to millions
Document SizeVariable-size documents with nested dataFixed-width rows, mostly uniform
Real-timeNeed real-time listeners/subscriptionsBatch or poll-based reads
TransactionsMulti-document transactions requiredNo transaction support needed
Latency Target< 10ms single-doc reads< 5ms single-row reads at any scale
Cost at Scale< 50GB data, moderate ops> 100GB data, high ops/sec
Team ExpertiseApplication developers, mobile/webData engineers, infrastructure teams

Access Pattern Deep Dive

When Firestore Is the Right Choice

Firestore excels at application-layer data with complex query requirements:

// User profile with real-time subscription
import { getFirestore, doc, onSnapshot, collection, query, where, orderBy } from 'firebase/firestore';

const db = getFirestore();

// Real-time listener - Firestore's killer feature
const unsubscribe = onSnapshot(
  doc(db, 'users', userId),
  (snapshot) => {
    const userData = snapshot.data();
    updateUI(userData);
  }
);

// Compound query with multiple conditions
// (Requires composite index - created automatically or via CLI)
const recentOrders = query(
  collection(db, 'orders'),
  where('userId', '==', userId),
  where('status', 'in', ['pending', 'processing']),
  where('createdAt', '>=', thirtyDaysAgo),
  orderBy('createdAt', 'desc')
);

// Multi-document transaction
import { runTransaction } from 'firebase/firestore';

await runTransaction(db, async (transaction) => {
  const accountRef = doc(db, 'accounts', accountId);
  const accountSnap = await transaction.get(accountRef);
  const currentBalance = accountSnap.data().balance;

  if (currentBalance &#x3C; amount) {
    throw new Error('Insufficient funds');
  }

  transaction.update(accountRef, { balance: currentBalance - amount });
  transaction.set(doc(collection(db, 'transactions')), {
    accountId,
    amount: -amount,
    type: 'debit',
    timestamp: new Date(),
  });
});

When Bigtable Is the Right Choice

Bigtable excels at high-throughput sequential workloads:

# Time-series ingestion - Bigtable's sweet spot
from google.cloud import bigtable
from google.cloud.bigtable import row_filters
import struct
import time

client = bigtable.Client(project="my-project", admin=True)
instance = client.instance("analytics-prod")
table = instance.table("events")

# Row key design: reverse timestamp for recent-first scans
# Format: {entity_id}#{reverse_timestamp}#{event_type}
def create_row_key(entity_id: str, timestamp: float, event_type: str) -> str:
    reverse_ts = str(int(9999999999 - timestamp)).zfill(10)
    return f"{entity_id}#{reverse_ts}#{event_type}"


def write_event_batch(events: list[dict]) -> None:
    """Write a batch of events with microsecond precision."""
    rows = []
    for event in events:
        row_key = create_row_key(
            event["entity_id"],
            event["timestamp"],
            event["event_type"]
        )
        row = table.direct_row(row_key)
        row.set_cell(
            "event_data",
            "payload",
            event["payload"].encode("utf-8"),
            timestamp=datetime.datetime.utcnow()
        )
        row.set_cell(
            "event_data",
            "source",
            event["source"].encode("utf-8")
        )
        row.set_cell(
            "metadata",
            "size_bytes",
            struct.pack(">i", len(event["payload"]))
        )
        rows.append(row)

    # Batch write - Bigtable handles millions of these per second
    table.mutate_rows(rows)


def read_recent_events(entity_id: str, limit: int = 100) -> list:
    """Read most recent events for an entity using row key prefix scan."""
    prefix = f"{entity_id}#"
    end_key = f"{entity_id}$"  # '$' is lexicographically after '#'

    partial_rows = table.read_rows(
        start_key=prefix.encode(),
        end_key=end_key.encode(),
        limit=limit,
        filter_=row_filters.CellsColumnLimitFilter(1)  # Latest version only
    )

    events = []
    for row in partial_rows:
        events.append(parse_row(row))
    return events

Performance Benchmarks at Scale

I benchmarked both services with equivalent workloads (adjusted for their natural access patterns):

Point Read Performance

Concurrent OpsFirestore P50Firestore P99Bigtable P50Bigtable P99
1,0004.2ms18ms3.1ms8ms
10,0005.8ms45ms3.2ms9ms
50,0008.1ms120ms3.4ms12ms
100,00012.4ms280ms3.5ms14ms

Bigtable's latency stays nearly flat regardless of throughput. Firestore degrades under extreme concurrency — its strong consistency guarantee has a cost.

Range Scan Performance (1000 rows/documents)

Data SizeFirestoreBigtableRatio
10GB45ms12ms3.75x
100GB52ms13ms4.0x
1TB68ms14ms4.9x
10TB89ms15ms5.9x

Bigtable's advantage grows with data size because its storage engine is optimized for sequential disk access patterns.

Cost Comparison at Scale

Scenario 1: 10M reads/day, 1M writes/day, 50GB storage

ComponentFirestoreBigtable (3-node cluster)
Reads$3.60/dayIncluded in node cost
Writes$1.80/dayIncluded in node cost
Storage$0.90/day$0.85/day
ComputeN/A (serverless)$14.04/day
Daily Total$6.30$14.89

At moderate scale, Firestore's serverless pricing wins.

Scenario 2: 1B reads/day, 100M writes/day, 5TB storage

ComponentFirestoreBigtable (20-node cluster)
Reads$360/dayIncluded in node cost
Writes$180/dayIncluded in node cost
Storage$90/day$85/day
ComputeN/A$93.60/day
Daily Total$630$178.60

At high scale, Bigtable's flat node-based pricing dramatically wins.

Cost Crossover Point

The crossover point is approximately 50M operations/day. Below that, Firestore is cheaper. Above that, Bigtable's economics improve linearly while Firestore's costs grow proportionally with operations.

Hybrid Architecture: Using Both

In practice, most large GCP applications use both. Our architecture:

┌─────────────────────────────────────────────────┐
│                 Application Layer                 │
├─────────────────────────────────────────────────┤
│                                                   │
│   User Profiles ──────► Firestore                │
│   Orders/Transactions ──► Firestore              │
│   Real-time Chat ──────► Firestore               │
│   Feature Flags ───────► Firestore               │
│                                                   │
│   Event Stream ────────► Bigtable                │
│   Analytics Data ──────► Bigtable                │
│   Time-series Metrics ─► Bigtable                │
│   Audit Logs ──────────► Bigtable                │
│                                                   │
└─────────────────────────────────────────────────┘

The event flow bridges both: Firestore triggers Cloud Functions on document changes, which write denormalized events to Bigtable for analytical queries.

Common Anti-Patterns

Using Firestore for time-series data. Firestore's document model and pricing make it expensive for append-heavy workloads. A sensor writing once per second generates 2.6M writes/month per sensor — at $0.18 per 100K writes, that's $4.68/sensor/month just for writes.

Using Bigtable for < 1TB of data. The minimum Bigtable cluster (3 nodes) costs ~$1,400/month. If your data fits comfortably in Firestore's pricing model, Bigtable is overkill.

Expecting Bigtable to handle ad-hoc queries. Bigtable has no secondary indexes, no query language, no filtering beyond row key prefix and column family. If you need flexible querying, you need Firestore or BigQuery.

Ignoring Firestore's 500-write-per-second-per-collection limit. This isn't documented prominently but it's real. High-throughput write workloads hitting a single collection need sharding strategies.

Decision Flowchart

  1. Do you need real-time subscriptions? → Firestore
  2. Do you need multi-document transactions? → Firestore
  3. Is your access pattern primarily time-range scans? → Bigtable
  4. Will you exceed 100K ops/second sustained? → Bigtable
  5. Is your data > 1TB? → Bigtable (cost-wise)
  6. Do you need secondary indexes or compound queries? → Firestore
  7. Is latency consistency critical under variable load? → Bigtable

Conclusion

Firestore and Bigtable aren't competing products — they're complementary tools for different layers of your architecture. Firestore handles application state with strong consistency and flexible querying. Bigtable handles high-throughput analytical workloads with predictable latency at any scale. The decision isn't which one to use — it's where the boundary sits between them in your data architecture.

Comments

    No comments yet. Be the first to share your thoughts.