GCP Cloud CDN Multi-Region Strategy: Cache Hit Optimization and Global Performance
Architecting a multi-region content delivery strategy on Cloud CDN with cache hit analysis, origin shielding, and edge configuration for global latency reduction.

Cloud CDN doesn't get the spotlight that CloudFront or Cloudflare receives, but for GCP-native architectures it offers tight integration with Cloud Load Balancing, Cloud Storage, and Cloud Run that makes global content delivery remarkably simple to configure and operate. After optimizing our CDN layer to serve 800M requests per month across 42 countries, here's what actually matters for cache hit ratios and end-user latency.
The Problem: Inconsistent Global Latency
Our SaaS platform serves a mix of static assets (JS bundles, images, fonts) and dynamic API responses to users across North America, Europe, and Asia-Pacific. Without CDN optimization, users in Singapore experienced 380ms TTFB compared to 45ms for users in Virginia where our primary origin sits.
The goal: sub-100ms TTFB for static content globally, and sub-200ms for cacheable API responses.
Cloud CDN Architecture
Cloud CDN sits behind Cloud HTTP(S) Load Balancing. Every request that hits the load balancer is evaluated for cache eligibility before reaching the origin.
The architecture has three layers:
- Edge caches: 150+ points of presence globally, first layer to serve cached content
- Mid-tier caches: Regional aggregation points that reduce origin load
- Origin: Your backend (Cloud Storage, Cloud Run, GCE, or external)
Cache Hit Ratio Optimization
Our starting cache hit ratio was 34%. After 6 weeks of optimization, we reached 89%. Here's what moved the needle:
1. Cache Key Normalization
By default, Cloud CDN uses the full request URL including query parameters as the cache key. Marketing UTM parameters were creating unique cache entries for identical content:
# Before: Each URL variant creates a separate cache entry
/app.js?utm_source=google&utm_medium=cpc # Cache miss
/app.js?utm_source=twitter&utm_campaign=q4 # Cache miss
/app.js # Cache miss (different key!)
# Solution: Exclude query parameters from cache key for static assets
gcloud compute backend-services update static-assets-backend \
--cache-key-policy-include-query-string=false \
--global
For API responses where some query parameters affect content:
# Include only parameters that change the response
gcloud compute backend-services update api-backend \
--cache-key-policy-query-string-whitelist="page,limit,category" \
--global
This single change increased our cache hit ratio from 34% to 61%.
2. Cache-Control Headers
The origin must send proper cache headers. We configured different TTLs by content type:
# Cloud Run service response headers
location /static/ {
add_header Cache-Control "public, max-age=31536000, immutable";
add_header CDN-Cache-Control "public, max-age=31536000";
}
location /api/catalog/ {
add_header Cache-Control "public, max-age=300, s-maxage=600, stale-while-revalidate=86400";
add_header CDN-Cache-Control "max-age=600";
}
location /api/user/ {
add_header Cache-Control "private, no-store";
}
The CDN-Cache-Control header overrides Cache-Control for CDN behavior while allowing different browser caching:
func SetCacheHeaders(w http.ResponseWriter, contentType string) {
switch contentType {
case "static":
// Immutable assets (hashed filenames)
w.Header().Set("Cache-Control", "public, max-age=31536000, immutable")
w.Header().Set("CDN-Cache-Control", "public, max-age=31536000")
case "catalog":
// Product data: fresh for 5min in browser, 10min in CDN
w.Header().Set("Cache-Control", "public, max-age=300, stale-while-revalidate=3600")
w.Header().Set("CDN-Cache-Control", "max-age=600")
case "personalized":
// User-specific: never cache in CDN
w.Header().Set("Cache-Control", "private, no-store")
w.Header().Set("CDN-Cache-Control", "no-store")
}
}
3. Signed URLs for Private Content
For authenticated content that should still be cached at the edge:
# Create a signing key
gcloud compute backend-services add-signed-url-key api-backend \
--signed-url-key-name=cdn-key-v1 \
--signed-url-key-file=cdn-key.base64 \
--global
# Configure signed URL requirements
gcloud compute backend-services update api-backend \
--signed-url-cache-max-age=3600 \
--global
func GenerateSignedURL(baseURL, keyName string, key []byte, expiration time.Time) (string, error) {
urlToSign := fmt.Sprintf("%s?Expires=%d&KeyName=%s",
baseURL,
expiration.Unix(),
keyName,
)
mac := hmac.New(sha256.New, key)
mac.Write([]byte(urlToSign))
sig := base64.URLEncoding.EncodeToString(mac.Sum(nil))
return fmt.Sprintf("%s&Signature=%s", urlToSign, sig), nil
}
Multi-Region Origin Configuration
For dynamic API responses, we deployed Cloud Run services in three regions with the load balancer routing to the nearest healthy origin:
# Create NEGs for each region
gcloud compute network-endpoint-groups create api-neg-us \
--region=us-central1 \
--network-endpoint-type=serverless \
--cloud-run-service=api-service
gcloud compute network-endpoint-groups create api-neg-eu \
--region=europe-west1 \
--network-endpoint-type=serverless \
--cloud-run-service=api-service
gcloud compute network-endpoint-groups create api-neg-asia \
--region=asia-southeast1 \
--network-endpoint-type=serverless \
--cloud-run-service=api-service
# Attach all NEGs to the backend service
gcloud compute backend-services add-backend api-backend \
--network-endpoint-group=api-neg-us \
--network-endpoint-group-region=us-central1 \
--global
gcloud compute backend-services add-backend api-backend \
--network-endpoint-group=api-neg-eu \
--network-endpoint-group-region=europe-west1 \
--global
gcloud compute backend-services add-backend api-backend \
--network-endpoint-group=api-neg-asia \
--network-endpoint-group-region=asia-southeast1 \
--global
Performance Results
After full optimization, our latency numbers across regions:
| Region | Before (TTFB) | After (TTFB) | Cache Hit Ratio |
|---|---|---|---|
| US (Virginia) | 45ms | 12ms | 92% |
| EU (Frankfurt) | 180ms | 28ms | 88% |
| APAC (Singapore) | 380ms | 41ms | 86% |
| LATAM (Sao Paulo) | 290ms | 52ms | 84% |
| Global Average | 210ms | 31ms | 89% |
Cache Invalidation Strategy
Cache invalidation is where most CDN implementations break down. We use a versioned approach:
# Invalidate specific paths (completes in <5 minutes globally)
gcloud compute url-maps invalidate-cdn-cache web-map \
--path="/api/catalog/*" \
--global
# For immediate propagation: use versioned URLs for static assets
# /static/v20251228/app.js -> never invalidate, new versions get new URLs
For API responses, we rely on short TTLs (5-10 minutes) rather than active invalidation. The cost of serving slightly stale catalog data for 5 minutes is lower than the complexity of real-time invalidation.
Cost Analysis
Cloud CDN charges for cache egress ($0.02-0.08/GB depending on region) plus cache fill ($0.01/10K requests). At our scale:
| Cost Component | Monthly |
|---|---|
| Cache egress (6TB) | $420 |
| Cache fill requests | $180 |
| HTTP(S) LB forwarding | $18 |
| Origin compute savings | -$2,100 |
| Net savings | $1,482/month |
The CDN pays for itself three times over through reduced origin compute.
Key Takeaways
- Cache key normalization is the highest-ROI optimization. Query parameter exclusion alone can double your hit ratio.
- Use CDN-Cache-Control to decouple CDN and browser caching. Different TTLs for different layers gives you flexibility without breaking browser behavior.
- Multi-region origins eliminate the latency floor. CDN caches the content, but cache misses still hit the origin — make that origin close to the user.
- Short TTLs beat active invalidation for dynamic content. 5-minute staleness is acceptable for most use cases and eliminates invalidation complexity.
- Monitor cache hit ratio by path pattern. Aggregate numbers hide problems. A 90% global hit ratio might mask a critical API path with 20% hit ratio.
Cloud CDN's tight integration with GCP services makes it the path of least resistance for GCP-native architectures. The performance gains are substantial and the operational overhead is minimal once properly configured.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.