GCP Cloud SQL High Availability: Failover Behavior, RTO/RPO, and Production Lessons
Deep analysis of Cloud SQL high availability failover mechanics with measured RTO/RPO data, connection handling, and production incident lessons.

Cloud SQL's high availability configuration promises automatic failover with zero data loss. After experiencing 7 failover events in production (3 planned, 4 unplanned) over 14 months, I can tell you exactly what happens during a failover, how long it takes, and where applications break despite Google's promises.
The Problem: Database Availability in Production
Our core transaction database (PostgreSQL 15 on Cloud SQL) serves 4,200 queries per second across 8 application services. A single minute of database downtime costs approximately $12,000 in lost transactions and requires 15 minutes of queue processing to recover. We needed HA, but we needed to understand exactly what "HA" means in Cloud SQL's implementation.
Cloud SQL HA Architecture
Cloud SQL HA uses synchronous replication between a primary instance and a standby instance in a different zone within the same region.
Key architectural details:
- Synchronous replication: Every write commits to both primary and standby before acknowledging
- Zone-level isolation: Primary and standby are in different availability zones
- Shared IP: A single IP address floats between primary and standby
- Heartbeat monitoring: Google's internal health checks detect primary failure
The synchronous replication is what makes RPO (Recovery Point Objective) equal to zero — no committed transaction is ever lost during failover.
Measured Failover Performance
We instrumented our application to measure exact downtime during each failover event:
| Event | Type | RTO (Total Downtime) | RPO (Data Loss) | Cause |
|---|---|---|---|---|
| 1 | Planned | 28s | 0 | Maintenance window |
| 2 | Planned | 32s | 0 | Instance resize |
| 3 | Unplanned | 47s | 0 | Zone failure |
| 4 | Unplanned | 62s | 0 | Instance crash |
| 5 | Planned | 25s | 0 | OS patch |
| 6 | Unplanned | 53s | 0 | Storage issue |
| 7 | Unplanned | 118s | 0 | Network partition |
Average RTO: 52 seconds. Maximum RTO: 118 seconds. RPO: Always zero.
The documented target is "under 60 seconds." Our data shows that's true 71% of the time, but network partitions can extend failover significantly.
What Happens During Those 52 Seconds
The failover sequence:
- Detection (5-15s): Health check fails, Google's control plane confirms the failure
- Decision (2-5s): Control plane initiates failover
- Promotion (10-20s): Standby is promoted to primary, WAL replay completes
- DNS/IP flip (5-10s): The shared IP is reassigned to the new primary
- Connection acceptance (3-8s): New primary begins accepting connections
During this window, applications experience:
- Existing connections: Broken (TCP reset or timeout)
- New connections: Refused (until IP flip completes)
- In-flight transactions: Rolled back
Application-Level Resilience
The database failing over is predictable. Applications handling it correctly is not. Here's our connection configuration:
package database
import (
"database/sql"
"time"
_ "github.com/jackc/pgx/v5/stdlib"
)
func NewConnectionPool(dsn string) (*sql.DB, error) {
db, err := sql.Open("pgx", dsn)
if err != nil {
return nil, err
}
// Connection pool settings optimized for failover resilience
db.SetMaxOpenConns(50)
db.SetMaxIdleConns(10)
db.SetConnMaxLifetime(5 * time.Minute) // Recycle connections frequently
db.SetConnMaxIdleTime(1 * time.Minute) // Don't hold stale connections
return db, nil
}
The critical settings are ConnMaxLifetime and ConnMaxIdleTime. Short lifetimes ensure that after a failover, stale connections to the old primary are naturally replaced within minutes even if the application doesn't detect the failure immediately.
Retry Logic for Failover Resilience
package database
import (
"context"
"errors"
"fmt"
"time"
"github.com/jackc/pgx/v5/pgconn"
)
type RetryConfig struct {
MaxAttempts int
InitialBackoff time.Duration
MaxBackoff time.Duration
RetryableErrors []string
}
var DefaultRetryConfig = RetryConfig{
MaxAttempts: 5,
InitialBackoff: 200 * time.Millisecond,
MaxBackoff: 10 * time.Second,
RetryableErrors: []string{
"57P01", // admin_shutdown
"57P02", // crash_shutdown
"57P03", // cannot_connect_now
"08006", // connection_failure
"08001", // sqlclient_unable_to_establish_sqlconnection
"08004", // sqlserver_rejected_establishment_of_sqlconnection
},
}
func ExecuteWithRetry[T any](
ctx context.Context,
config RetryConfig,
operation func(ctx context.Context) (T, error),
) (T, error) {
var result T
var lastErr error
backoff := config.InitialBackoff
for attempt := 1; attempt <= config.MaxAttempts; attempt++ {
result, lastErr = operation(ctx)
if lastErr == nil {
return result, nil
}
if !isRetryableError(lastErr, config.RetryableErrors) {
return result, lastErr
}
if attempt == config.MaxAttempts {
break
}
// Exponential backoff with jitter
sleep := backoff + time.Duration(rand.Int63n(int64(backoff/2)))
select {
case <-ctx.Done():
return result, ctx.Err()
case <-time.After(sleep):
}
backoff = min(backoff*2, config.MaxBackoff)
}
return result, fmt.Errorf("operation failed after %d attempts: %w", config.MaxAttempts, lastErr)
}
func isRetryableError(err error, codes []string) bool {
var pgErr *pgconn.PgError
if errors.As(err, &pgErr) {
for _, code := range codes {
if pgErr.Code == code {
return true
}
}
}
// Also retry on connection-level errors
return errors.Is(err, context.DeadlineExceeded) ||
isConnectionError(err)
}
Connection Proxy Configuration
We use Cloud SQL Auth Proxy with connection draining:
# Cloud SQL Proxy sidecar configuration (Kubernetes)
apiVersion: v1
kind: Pod
spec:
containers:
- name: cloud-sql-proxy
image: gcr.io/cloud-sql-connectors/cloud-sql-proxy:2.8.0
args:
- "--structured-logs"
- "--port=5432"
- "--auto-iam-authn"
- "--max-sigterm-delay=30s" # Graceful shutdown
- "--health-check"
- "--http-port=9090"
- "project:us-central1:primary-db"
resources:
requests:
cpu: "100m"
memory: "128Mi"
livenessProbe:
httpGet:
path: /liveness
port: 9090
initialDelaySeconds: 5
periodSeconds: 10
readinessProbe:
httpGet:
path: /readiness
port: 9090
initialDelaySeconds: 5
periodSeconds: 5
The proxy handles TLS, IAM authentication, and connection management. The readiness probe integration means Kubernetes stops routing traffic to pods whose database connection is broken during failover.
Read Replica Strategy for HA
For read-heavy workloads, read replicas provide additional resilience:
# Create read replicas in different zones
gcloud sql instances create primary-db-read-1 \
--master-instance-name=primary-db \
--zone=us-central1-b \
--tier=db-custom-8-32768 \
--availability-type=zonal
gcloud sql instances create primary-db-read-2 \
--master-instance-name=primary-db \
--zone=us-central1-c \
--tier=db-custom-8-32768 \
--availability-type=zonal
During primary failover, read replicas continue serving read traffic with minimal disruption (1-3 seconds of replication lag catch-up).
Instance Configuration for HA
# Create HA instance with optimal settings
gcloud sql instances create primary-db \
--database-version=POSTGRES_15 \
--tier=db-custom-8-32768 \
--region=us-central1 \
--availability-type=REGIONAL \
--storage-type=SSD \
--storage-size=500GB \
--storage-auto-increase \
--backup-start-time=04:00 \
--enable-point-in-time-recovery \
--retained-backups-count=14 \
--maintenance-window-day=SUN \
--maintenance-window-hour=6 \
--deny-maintenance-period-start-date=2026-01-20 \
--deny-maintenance-period-end-date=2026-01-27 \
--insights-config-query-insights-enabled \
--insights-config-record-application-tags
Lessons From Production Incidents
Incident 1: Application used connection pooling with 30-minute ConnMaxLifetime. After failover, 80% of connections were stale for up to 30 minutes. Fix: reduced to 5 minutes.
Incident 2: Application retry logic retried non-idempotent writes (INSERT without ON CONFLICT). After failover, the in-flight transaction was rolled back, retried, and created a duplicate record. Fix: added idempotency keys to all write operations.
Incident 3: Maintenance window failover triggered during peak hours because we set the window to "any time Sunday" instead of "Sunday 6 AM UTC." Fix: explicit maintenance window with deny periods around peak traffic.
Key Takeaways
- Average RTO is 52 seconds, not "instant". Plan for 60-120 seconds of database unavailability during failover.
- RPO is genuinely zero. Synchronous replication means no committed data is lost. This is the strongest guarantee Cloud SQL offers.
- Applications break, not the database. Connection handling, retry logic, and idempotency determine whether a failover is a non-event or an incident.
- Short connection lifetimes save you. 5-minute
ConnMaxLifetimeensures natural connection rotation without depending on failure detection. - Test failover regularly. Run
gcloud sql instances failovermonthly in staging. The only way to validate your application's behavior is to trigger it. - Maintenance windows are failover events. Every patch, resize, or version upgrade triggers a failover. Schedule them deliberately.
The database will fail over correctly. The question is whether your application will handle those 52 seconds gracefully.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.