GCP Cloud SQL High Availability: Failover Behavior, RTO/RPO, and Production Lessons

Deep analysis of Cloud SQL high availability failover mechanics with measured RTO/RPO data, connection handling, and production incident lessons.

#gcp#cloud-sql#database#high-availability
Cover image for the article: GCP Cloud SQL High Availability: Failover Behavior, RTO/RPO, and Production Lessons

Cloud SQL's high availability configuration promises automatic failover with zero data loss. After experiencing 7 failover events in production (3 planned, 4 unplanned) over 14 months, I can tell you exactly what happens during a failover, how long it takes, and where applications break despite Google's promises.

The Problem: Database Availability in Production

Our core transaction database (PostgreSQL 15 on Cloud SQL) serves 4,200 queries per second across 8 application services. A single minute of database downtime costs approximately $12,000 in lost transactions and requires 15 minutes of queue processing to recover. We needed HA, but we needed to understand exactly what "HA" means in Cloud SQL's implementation.

Cloud SQL HA Architecture

Cloud SQL HA uses synchronous replication between a primary instance and a standby instance in a different zone within the same region.

Cloud SQL HA Failover Architecture

Key architectural details:

  • Synchronous replication: Every write commits to both primary and standby before acknowledging
  • Zone-level isolation: Primary and standby are in different availability zones
  • Shared IP: A single IP address floats between primary and standby
  • Heartbeat monitoring: Google's internal health checks detect primary failure

The synchronous replication is what makes RPO (Recovery Point Objective) equal to zero — no committed transaction is ever lost during failover.

Measured Failover Performance

We instrumented our application to measure exact downtime during each failover event:

EventTypeRTO (Total Downtime)RPO (Data Loss)Cause
1Planned28s0Maintenance window
2Planned32s0Instance resize
3Unplanned47s0Zone failure
4Unplanned62s0Instance crash
5Planned25s0OS patch
6Unplanned53s0Storage issue
7Unplanned118s0Network partition

Average RTO: 52 seconds. Maximum RTO: 118 seconds. RPO: Always zero.

The documented target is "under 60 seconds." Our data shows that's true 71% of the time, but network partitions can extend failover significantly.

What Happens During Those 52 Seconds

The failover sequence:

  1. Detection (5-15s): Health check fails, Google's control plane confirms the failure
  2. Decision (2-5s): Control plane initiates failover
  3. Promotion (10-20s): Standby is promoted to primary, WAL replay completes
  4. DNS/IP flip (5-10s): The shared IP is reassigned to the new primary
  5. Connection acceptance (3-8s): New primary begins accepting connections

During this window, applications experience:

  • Existing connections: Broken (TCP reset or timeout)
  • New connections: Refused (until IP flip completes)
  • In-flight transactions: Rolled back

Application-Level Resilience

The database failing over is predictable. Applications handling it correctly is not. Here's our connection configuration:

package database

import (
    "database/sql"
    "time"
    
    _ "github.com/jackc/pgx/v5/stdlib"
)

func NewConnectionPool(dsn string) (*sql.DB, error) {
    db, err := sql.Open("pgx", dsn)
    if err != nil {
        return nil, err
    }
    
    // Connection pool settings optimized for failover resilience
    db.SetMaxOpenConns(50)
    db.SetMaxIdleConns(10)
    db.SetConnMaxLifetime(5 * time.Minute)   // Recycle connections frequently
    db.SetConnMaxIdleTime(1 * time.Minute)   // Don't hold stale connections
    
    return db, nil
}

The critical settings are ConnMaxLifetime and ConnMaxIdleTime. Short lifetimes ensure that after a failover, stale connections to the old primary are naturally replaced within minutes even if the application doesn't detect the failure immediately.

Retry Logic for Failover Resilience

package database

import (
    "context"
    "errors"
    "fmt"
    "time"
    
    "github.com/jackc/pgx/v5/pgconn"
)

type RetryConfig struct {
    MaxAttempts     int
    InitialBackoff  time.Duration
    MaxBackoff      time.Duration
    RetryableErrors []string
}

var DefaultRetryConfig = RetryConfig{
    MaxAttempts:    5,
    InitialBackoff: 200 * time.Millisecond,
    MaxBackoff:     10 * time.Second,
    RetryableErrors: []string{
        "57P01", // admin_shutdown
        "57P02", // crash_shutdown  
        "57P03", // cannot_connect_now
        "08006", // connection_failure
        "08001", // sqlclient_unable_to_establish_sqlconnection
        "08004", // sqlserver_rejected_establishment_of_sqlconnection
    },
}

func ExecuteWithRetry[T any](
    ctx context.Context,
    config RetryConfig,
    operation func(ctx context.Context) (T, error),
) (T, error) {
    var result T
    var lastErr error
    
    backoff := config.InitialBackoff
    
    for attempt := 1; attempt <= config.MaxAttempts; attempt++ {
        result, lastErr = operation(ctx)
        if lastErr == nil {
            return result, nil
        }
        
        if !isRetryableError(lastErr, config.RetryableErrors) {
            return result, lastErr
        }
        
        if attempt == config.MaxAttempts {
            break
        }
        
        // Exponential backoff with jitter
        sleep := backoff + time.Duration(rand.Int63n(int64(backoff/2)))
        select {
        case <-ctx.Done():
            return result, ctx.Err()
        case <-time.After(sleep):
        }
        
        backoff = min(backoff*2, config.MaxBackoff)
    }
    
    return result, fmt.Errorf("operation failed after %d attempts: %w", config.MaxAttempts, lastErr)
}

func isRetryableError(err error, codes []string) bool {
    var pgErr *pgconn.PgError
    if errors.As(err, &pgErr) {
        for _, code := range codes {
            if pgErr.Code == code {
                return true
            }
        }
    }
    // Also retry on connection-level errors
    return errors.Is(err, context.DeadlineExceeded) ||
           isConnectionError(err)
}

Connection Proxy Configuration

We use Cloud SQL Auth Proxy with connection draining:

# Cloud SQL Proxy sidecar configuration (Kubernetes)
apiVersion: v1
kind: Pod
spec:
  containers:
    - name: cloud-sql-proxy
      image: gcr.io/cloud-sql-connectors/cloud-sql-proxy:2.8.0
      args:
        - "--structured-logs"
        - "--port=5432"
        - "--auto-iam-authn"
        - "--max-sigterm-delay=30s"    # Graceful shutdown
        - "--health-check"
        - "--http-port=9090"
        - "project:us-central1:primary-db"
      resources:
        requests:
          cpu: "100m"
          memory: "128Mi"
      livenessProbe:
        httpGet:
          path: /liveness
          port: 9090
        initialDelaySeconds: 5
        periodSeconds: 10
      readinessProbe:
        httpGet:
          path: /readiness
          port: 9090
        initialDelaySeconds: 5
        periodSeconds: 5

The proxy handles TLS, IAM authentication, and connection management. The readiness probe integration means Kubernetes stops routing traffic to pods whose database connection is broken during failover.

Read Replica Strategy for HA

For read-heavy workloads, read replicas provide additional resilience:

# Create read replicas in different zones
gcloud sql instances create primary-db-read-1 \
  --master-instance-name=primary-db \
  --zone=us-central1-b \
  --tier=db-custom-8-32768 \
  --availability-type=zonal

gcloud sql instances create primary-db-read-2 \
  --master-instance-name=primary-db \
  --zone=us-central1-c \
  --tier=db-custom-8-32768 \
  --availability-type=zonal

During primary failover, read replicas continue serving read traffic with minimal disruption (1-3 seconds of replication lag catch-up).

Cloud SQL Failover Timeline

Instance Configuration for HA

# Create HA instance with optimal settings
gcloud sql instances create primary-db \
  --database-version=POSTGRES_15 \
  --tier=db-custom-8-32768 \
  --region=us-central1 \
  --availability-type=REGIONAL \
  --storage-type=SSD \
  --storage-size=500GB \
  --storage-auto-increase \
  --backup-start-time=04:00 \
  --enable-point-in-time-recovery \
  --retained-backups-count=14 \
  --maintenance-window-day=SUN \
  --maintenance-window-hour=6 \
  --deny-maintenance-period-start-date=2026-01-20 \
  --deny-maintenance-period-end-date=2026-01-27 \
  --insights-config-query-insights-enabled \
  --insights-config-record-application-tags

Lessons From Production Incidents

Incident 1: Application used connection pooling with 30-minute ConnMaxLifetime. After failover, 80% of connections were stale for up to 30 minutes. Fix: reduced to 5 minutes.

Incident 2: Application retry logic retried non-idempotent writes (INSERT without ON CONFLICT). After failover, the in-flight transaction was rolled back, retried, and created a duplicate record. Fix: added idempotency keys to all write operations.

Incident 3: Maintenance window failover triggered during peak hours because we set the window to "any time Sunday" instead of "Sunday 6 AM UTC." Fix: explicit maintenance window with deny periods around peak traffic.

Key Takeaways

  1. Average RTO is 52 seconds, not "instant". Plan for 60-120 seconds of database unavailability during failover.
  2. RPO is genuinely zero. Synchronous replication means no committed data is lost. This is the strongest guarantee Cloud SQL offers.
  3. Applications break, not the database. Connection handling, retry logic, and idempotency determine whether a failover is a non-event or an incident.
  4. Short connection lifetimes save you. 5-minute ConnMaxLifetime ensures natural connection rotation without depending on failure detection.
  5. Test failover regularly. Run gcloud sql instances failover monthly in staging. The only way to validate your application's behavior is to trigger it.
  6. Maintenance windows are failover events. Every patch, resize, or version upgrade triggers a failover. Schedule them deliberately.

The database will fail over correctly. The question is whether your application will handle those 52 seconds gracefully.

Comments

    No comments yet. Be the first to share your thoughts.