Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Eight months ago, we migrated our development and staging databases from RDS Postgres to Neon. Three months later, we moved two production workloads. The result: our database bill dropped from $14,200/month to $3,800/month, preview environments spin up in 3 seconds instead of 12 minutes, and our developers stopped asking me for database access because they can branch production data instantly. Here is the full story — the wins, the gotchas, and the workloads we intentionally kept on RDS.
Why We Looked Beyond RDS
Our infrastructure serves a B2B SaaS platform with 2,400 active tenants. The database landscape before migration:
- Production: 2× db.r6g.2xlarge Multi-AZ ($4,800/month each)
- Staging: 1× db.r6g.xlarge ($2,400/month)
- Developer databases: 8× db.t3.medium ($1,600/month total) — always running
- Total: $14,200/month
The waste was obvious. Developer databases ran 24/7 but were used 8 hours/day, 5 days/week. Staging was idle 18 hours/day. Even production ran at 15% CPU utilization 90% of the time, with spikes only during business hours.
Neon's Architecture: Why It Enables Serverless
Traditional Postgres ties compute and storage together. Neon separates them:
- Compute: Postgres instances that scale up/down independently (even to zero)
- Storage: Distributed page server that stores data as an LSM tree of WAL records
- Branching: Copy-on-write semantics — branching a 500GB database is instantaneous because it shares pages with the parent until writes diverge
This architecture means you pay for compute only when queries are running, and branching does not duplicate storage.
Migration Path: RDS to Neon
We used pg_dump/pg_restore for the initial migration, with logical replication for the cutover window:
#!/bin/bash
# migrate-to-neon.sh - RDS to Neon migration with minimal downtime
set -euo pipefail
SOURCE_HOST="prod-db.xxx.us-east-1.rds.amazonaws.com"
SOURCE_DB="app_production"
NEON_HOST="ep-xyz.us-east-1.aws.neon.tech"
NEON_DB="app_production"
echo "=== Phase 1: Schema + data dump ==="
pg_dump \
--host="$SOURCE_HOST" \
--dbname="$SOURCE_DB" \
--format=directory \
--jobs=8 \
--no-owner \
--no-privileges \
--compress=zstd:6 \
--file=/tmp/migration_dump
echo "=== Phase 2: Restore to Neon ==="
pg_restore \
--host="$NEON_HOST" \
--dbname="$NEON_DB" \
--format=directory \
--jobs=8 \
--no-owner \
--no-privileges \
/tmp/migration_dump
echo "=== Phase 3: Set up logical replication for delta sync ==="
# On RDS source (requires rds.logical_replication = 1)
psql --host="$SOURCE_HOST" --dbname="$SOURCE_DB" -c "
CREATE PUBLICATION neon_migration FOR ALL TABLES;
"
# On Neon target
psql --host="$NEON_HOST" --dbname="$NEON_DB" -c "
CREATE SUBSCRIPTION neon_sub
CONNECTION 'host=$SOURCE_HOST dbname=$SOURCE_DB'
PUBLICATION neon_migration
WITH (copy_data = false);
"
echo "=== Phase 4: Monitor replication lag ==="
echo "Run: SELECT * FROM pg_stat_subscription;"
echo "When lag = 0, switch application connection strings."
The 500GB production database took 4 hours to dump/restore. Logical replication caught up the delta in 12 minutes. Total cutover downtime: 8 seconds (DNS TTL for the connection string swap).
Database Branching: The Killer Feature
Every pull request in our CI/CD pipeline gets its own database branch. This is not a copy — it is an instant, copy-on-write fork of production data:
// scripts/create-preview-branch.ts
import { createClient } from '@neondatabase/api-client';
interface PreviewBranch {
branchId: string;
host: string;
connectionString: string;
}
async function createPreviewBranch(prNumber: number): Promise<PreviewBranch> {
const neon = createClient({ apiKey: process.env.NEON_API_KEY! });
// Branch from production — instant, regardless of DB size
const { data: branch } = await neon.createProjectBranch(
process.env.NEON_PROJECT_ID!,
{
branch: {
name: `preview/pr-${prNumber}`,
parent_id: 'br-production-main', // Fork production data
},
endpoints: [{
type: 'read_write',
autoscaling_limit_min_cu: 0.25, // Scale to near-zero when idle
autoscaling_limit_max_cu: 2, // Cap at 2 CU for previews
suspend_timeout_seconds: 300, // Sleep after 5 min idle
}],
}
);
const endpoint = branch.endpoints![0];
// Run migrations on the branch
const connectionString = `postgresql://${endpoint.host}/app_production?sslmode=require`;
return {
branchId: branch.branch!.id,
host: endpoint.host,
connectionString,
};
}
async function deletePreviewBranch(prNumber: number): Promise<void> {
const neon = createClient({ apiKey: process.env.NEON_API_KEY! });
const { data: branches } = await neon.listProjectBranches(
process.env.NEON_PROJECT_ID!
);
const branch = branches.branches.find(
b => b.name === `preview/pr-${prNumber}`
);
if (branch) {
await neon.deleteProjectBranch(
process.env.NEON_PROJECT_ID!,
branch.id
);
}
}
The developer experience transformation:
| Metric | Before (RDS) | After (Neon) |
|---|---|---|
| New environment database setup | 12 minutes | 3 seconds |
| Data freshness in previews | Weekly snapshot | Real-time fork |
| Storage cost per preview | $80/month (full copy) | $0.03/month (CoW delta) |
| Max concurrent preview DBs | 4 (cost-limited) | Unlimited |
| Cleanup on PR close | Manual | Automated webhook |
Scale-to-Zero Economics
Neon compute suspends after configurable idle timeout. For our workloads:
Production (always-on):
- Min: 4 CU, Max: 16 CU
- Autoscales based on connection load
- Never suspends (suspend_timeout = 0)
Staging:
- Min: 0.25 CU, Max: 8 CU
- Suspends after 10 min idle
- Active ~10 hours/day = 42% compute savings
Developer branches:
- Min: 0.25 CU, Max: 2 CU
- Suspends after 5 min idle
- Active ~3 hours/day = 87% compute savings
The cold start penalty when a suspended compute resumes is 400-800ms for the first query. For developer and staging environments, this is imperceptible. For production, we keep compute always warm.
Connection Pooling: The Critical Configuration
Neon uses a built-in connection pooler based on PgBouncer. For serverless applications (Lambda, Vercel Functions) that create many short-lived connections, this is essential:
// Correct: Use pooled connection string for serverless
const pooledUrl = 'postgresql://user:pass@ep-xyz-pooler.us-east-1.aws.neon.tech/db?sslmode=require';
// Direct connection for migrations and long-running queries
const directUrl = 'postgresql://user:pass@ep-xyz.us-east-1.aws.neon.tech/db?sslmode=require';
// Prisma configuration example
// schema.prisma
// datasource db {
// provider = "postgresql"
// url = env("DATABASE_URL") // Pooled for application queries
// directUrl = env("DIRECT_DATABASE_URL") // Direct for migrations
// }
Without the pooler, serverless functions opening 500+ connections during traffic spikes will exhaust Postgres's max_connections. The built-in pooler handles 10,000+ concurrent connections routing to ~100 backend Postgres connections.
What We Kept on RDS
Not everything migrated. Two workloads stayed on RDS:
1. High-write OLTP with strict latency SLAs (<5ms P99)
Neon adds ~2ms network latency due to the separated storage architecture. For our payment processing service requiring <5ms P99 query latency, that overhead pushes us over budget. RDS with local NVMe storage delivers consistent 1.2ms P99.
2. PostGIS-heavy workloads
Neon supports PostGIS, but spatial index performance on separated storage shows 40% regression compared to local-disk RDS for complex geospatial queries. Our logistics service running polygon intersection queries stays on RDS.
Production Benchmarks: Neon vs RDS
Testing our primary application workload (mixed read/write, 70/30 split):
| Metric | RDS r6g.2xlarge | Neon 8 CU | Notes |
|---|---|---|---|
| Simple SELECT P50 | 1.1ms | 2.8ms | Storage round-trip overhead |
| Simple SELECT P99 | 3.2ms | 6.1ms | Consistent ~3ms delta |
| Complex JOIN P50 | 12ms | 14ms | Difference narrows with query complexity |
| INSERT P50 | 1.4ms | 3.1ms | WAL write to remote storage |
| Bulk INSERT (10K rows) | 180ms | 210ms | Nearly equivalent at batch scale |
| Branch creation | N/A | 2.8s | Instant vs. hours for RDS snapshot |
| Connection establish | 45ms | 52ms | Via pooler, negligible difference |
Cost Comparison: 8-Month Reality
| Category | RDS (before) | Neon (after) | Savings |
|---|---|---|---|
| Production compute | $9,600/mo | $2,400/mo | -75% |
| Staging compute | $2,400/mo | $340/mo | -86% |
| Developer databases | $1,600/mo | $60/mo | -96% |
| Storage | $600/mo | $1,000/mo | +67% |
| Total | $14,200/mo | $3,800/mo | -73% |
Storage is more expensive on Neon because of the WAL-based architecture (storing history enables branching). But compute savings overwhelm the storage premium.
Gotchas We Hit in Production
Cold start after suspension is user-visible. If you configure production to suspend, the first request after idle will see 400-800ms latency. We learned this the hard way during a 3 AM monitoring gap. Solution: keep production compute always warm.
Logical replication FROM Neon is limited. We needed to replicate data to our analytics warehouse. Neon supports logical replication as a publisher, but with caveats around WAL retention during storage scale events. We switched to CDC via Debezium pointing at Neon's WAL.
Extension availability differs from RDS. Most common extensions work (pg_stat_statements, PostGIS, pgvector). But some proprietary RDS extensions (pg_cron with RDS event scheduling, aws_s3 for direct S3 imports) have no equivalent. Check your extension list before migrating.
Connection string format matters. The pooled endpoint (-pooler suffix) uses transaction-mode pooling. Prepared statements and advisory locks do not work in transaction mode. Use direct connections for migrations and background jobs that need session-level features.
Conclusion
Neon is not a universal RDS replacement — it is a fundamentally different database architecture that excels at specific workload patterns: variable traffic, developer environments, preview deployments, and cost-sensitive applications where compute utilization is low. The branching capability alone justified our migration by transforming how developers interact with data.
For latency-critical OLTP workloads where every millisecond matters, keep RDS with local storage. For everything else — and especially for the 80% of database instances that sit idle most of the time — serverless Postgres eliminates the provisioning tax that has plagued database operations for decades. The future of database infrastructure is not bigger instances; it is smarter allocation of compute to the moments when queries actually run.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Synthetic Monitors From 12 Regions Catching Issues Before Users
How to implement global synthetic monitoring that detects availability and performance degradation from every major region—catching issues minutes before real users are impacted.

Comments
No comments yet. Be the first to share your thoughts.