GCP Global Load Balancer Routing Strategies for Multi-Region Apps
Implementing Google Cloud global load balancing with advanced routing rules, traffic splitting, and intelligent failover across regions

Introduction
Google Cloud's global load balancer is fundamentally different from traditional regional load balancers. It provides a single anycast IP address that routes traffic to the nearest healthy backend across all GCP regions. This architecture eliminates the need for DNS-based failover (and its inherent TTL delays) and provides sub-second failover between regions.
In production deployments, GCP's global load balancer achieves 99.99% availability with cross-region failover completing in under 3 seconds. This article covers the routing strategies, traffic management features, and configuration patterns that make this possible.
Global Load Balancer Architecture
GCP's global load balancer operates at Google's network edge through over 140 points of presence worldwide. Traffic enters Google's network at the closest edge location and is routed internally to the optimal backend.
Load Balancer Types Comparison
| Feature | External HTTP(S) | External TCP/UDP | Internal HTTP(S) | Internal TCP/UDP |
|---|---|---|---|---|
| Scope | Global | Global (Premium) | Regional | Regional |
| Protocol | HTTP/1.1, HTTP/2, gRPC | TCP, UDP | HTTP/1.1, HTTP/2, gRPC | TCP, UDP |
| Anycast IP | Yes | Yes | No | No |
| TLS Termination | Yes | No | Yes | No |
| URL Map | Yes | No | Yes | No |
| Cloud CDN | Yes | No | No | No |
| Cloud Armor | Yes | Yes | No | No |
| Health Checks | HTTP(S), TCP, gRPC | TCP, HTTP(S) | HTTP(S), TCP, gRPC | TCP, HTTP(S) |
Terraform: Full Global Load Balancer Setup
# Reserve global static IP
resource "google_compute_global_address" "default" {
name = "global-lb-ip"
}
# Health check
resource "google_compute_health_check" "default" {
name = "api-health-check"
check_interval_sec = 5
timeout_sec = 3
healthy_threshold = 2
unhealthy_threshold = 3
http_health_check {
port = 8080
request_path = "/healthz"
}
}
# Backend service with multiple region backends
resource "google_compute_backend_service" "api" {
name = "api-backend-service"
protocol = "HTTP"
port_name = "http"
timeout_sec = 30
health_checks = [google_compute_health_check.default.id]
load_balancing_scheme = "EXTERNAL_MANAGED"
locality_lb_policy = "ROUND_ROBIN"
backend {
group = google_compute_region_network_endpoint_group.us.id
balancing_mode = "RATE"
max_rate_per_endpoint = 100
capacity_scaler = 1.0
}
backend {
group = google_compute_region_network_endpoint_group.eu.id
balancing_mode = "RATE"
max_rate_per_endpoint = 100
capacity_scaler = 1.0
}
backend {
group = google_compute_region_network_endpoint_group.asia.id
balancing_mode = "RATE"
max_rate_per_endpoint = 100
capacity_scaler = 0.5 # Lower capacity in Asia
}
outlier_detection {
consecutive_errors = 5
interval {
seconds = 10
}
base_ejection_time {
seconds = 30
}
max_ejection_percent = 50
}
}
URL Map with Advanced Routing
URL maps enable content-based routing where different URL paths are directed to different backend services:
resource "google_compute_url_map" "default" {
name = "api-url-map"
default_service = google_compute_backend_service.api.id
host_rule {
hosts = ["api.example.com"]
path_matcher = "api-paths"
}
host_rule {
hosts = ["static.example.com"]
path_matcher = "static-paths"
}
path_matcher {
name = "api-paths"
default_service = google_compute_backend_service.api.id
route_rules {
priority = 1
match_rules {
prefix_match = "/v2/"
header_matches {
header_name = "X-API-Version"
exact_match = "2.0"
}
}
route_action {
weighted_backend_services {
backend_service = google_compute_backend_service.api_v2.id
weight = 100
}
}
}
route_rules {
priority = 2
match_rules {
prefix_match = "/api/"
}
route_action {
weighted_backend_services {
backend_service = google_compute_backend_service.api.id
weight = 90
}
weighted_backend_services {
backend_service = google_compute_backend_service.api_canary.id
weight = 10
}
retry_policy {
retry_conditions = ["5xx", "reset", "connect-failure"]
num_retries = 3
per_try_timeout {
seconds = 5
}
}
timeout {
seconds = 30
}
}
}
}
path_matcher {
name = "static-paths"
default_service = google_compute_backend_bucket.static.id
}
}
Traffic Splitting for Canary Deployments
Traffic splitting enables gradual rollouts with real-time traffic control:
# Update traffic split via gcloud
gcloud compute url-maps edit api-url-map \
--project=my-project
# Example: Route 5% traffic to canary
# In the URL map YAML:
# routeAction:
# weightedBackendServices:
# - backendService: projects/my-project/global/backendServices/api-stable
# weight: 95
# - backendService: projects/my-project/global/backendServices/api-canary
# weight: 5
Traffic Split Strategy by Risk Level
| Deployment Risk | Initial Split | Step Size | Observation Period | Full Rollout |
|---|---|---|---|---|
| Low (config change) | 10% | 25% | 5 min | 20 min |
| Medium (new feature) | 5% | 10% | 15 min | 2 hours |
| High (core refactor) | 1% | 5% | 30 min | 6 hours |
| Critical (data layer) | 0.1% | 1% | 1 hour | 24 hours |
Failover Performance Benchmarks
Google's global load balancer provides significantly faster failover than DNS-based approaches:
| Metric | GCP Global LB | DNS Failover (Route 53) | CDN Failover |
|---|---|---|---|
| Detection time | 5-15s | 10-30s | 10-30s |
| Failover execution | < 1s | 0-60s (TTL) | 5-30s |
| Total MTTR | 6-16s | 10-90s | 15-60s |
| False positive rate | < 0.01% | 0.02-0.05% | 0.01-0.03% |
| Cost (monthly) | ~$18 + traffic | ~$8 | Varies |
Outlier Detection and Circuit Breaking
Configure backend-level circuit breaking to prevent cascading failures:
resource "google_compute_backend_service" "api" {
# ... other config ...
circuit_breakers {
max_connections = 1000
max_pending_requests = 500
max_requests = 2000
max_retries = 3
max_requests_per_connection = 100
}
outlier_detection {
consecutive_errors = 5
interval { seconds = 10 }
base_ejection_time { seconds = 30 }
max_ejection_percent = 50
enforcing_consecutive_errors = 100
enforcing_success_rate = 100
success_rate_minimum_hosts = 3
success_rate_request_volume = 100
success_rate_stdev_factor = 1900
}
}
Cloud Armor Integration
Protect global endpoints with Cloud Armor security policies:
resource "google_compute_security_policy" "api_policy" {
name = "api-security-policy"
# Rate limiting
rule {
action = "throttle"
priority = 1000
match {
versioned_expr = "SRC_IPS_V1"
config {
src_ip_ranges = ["*"]
}
}
rate_limit_options {
conform_action = "allow"
exceed_action = "deny(429)"
rate_limit_threshold {
count = 100
interval_sec = 60
}
}
}
# Block known bad IPs
rule {
action = "deny(403)"
priority = 100
match {
versioned_expr = "SRC_IPS_V1"
config {
src_ip_ranges = ["192.0.2.0/24", "198.51.100.0/24"]
}
}
}
# Default allow
rule {
action = "allow"
priority = 2147483647
match {
versioned_expr = "SRC_IPS_V1"
config {
src_ip_ranges = ["*"]
}
}
}
}
Cost Optimization
| Component | Cost | Optimization |
|---|---|---|
| Forwarding rule | $0.025/hr ($18/mo) | Combine services behind one LB |
| Data processing | $0.008-0.012/GB | Use Cloud CDN for static content |
| Cloud Armor | $5/policy + $1/rule | Consolidate rules |
| SSL Certificate | Free (managed) | Use managed certificates |
| Health checks | Free | No additional cost |
Key Takeaways
- GCP's global load balancer provides sub-3-second failover without DNS propagation delays, making it significantly faster than DNS-based failover approaches.
- Use URL maps with weighted backend services for canary deployments with real-time traffic splitting and zero-downtime rollbacks.
- Configure outlier detection to automatically eject unhealthy backends and prevent cascading failures across regions.
- Combine Cloud Armor with global load balancing for DDoS protection and rate limiting at Google's network edge before traffic reaches your backends.
- Use capacity scalers on backend groups to control regional traffic distribution based on available capacity.
- Global load balancing costs approximately $18/month for the forwarding rule plus data processing fees, which is cost-effective for the reliability it provides.
- Circuit breakers with aggressive timeouts prevent resource exhaustion during backend failures and enable graceful degradation.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.