GCP Cloud Run Auto-Scaling in Production: A Deep Dive into Request-Based Metrics

Analyzing Cloud Run scaling behavior under production traffic with request-based metrics, concurrency tuning, and cold start mitigation strategies.

#gcp#cloud-run#serverless#auto-scaling
Cover image for the article: GCP Cloud Run Auto-Scaling in Production: A Deep Dive into Request-Based Metrics

After running Cloud Run workloads serving 50M+ requests per month across multiple services, I've developed a nuanced understanding of how its autoscaler behaves under real production conditions. The documentation tells you what it does. This article tells you what actually happens when traffic spikes hit at 3 AM.

The Problem: Unpredictable Scaling Under Bursty Traffic

Our payment processing service experienced a pattern that's common in production: traffic that doesn't arrive in smooth curves. We saw 10x spikes within 60-second windows during flash sales, and Cloud Run's default configuration wasn't keeping up. Request latency would spike from p50 of 45ms to p99 of 8.2 seconds during these bursts.

The root cause wasn't Cloud Run itself — it was our misunderstanding of how the autoscaler makes decisions.

How Cloud Run's Autoscaler Actually Works

Cloud Run uses a request-based autoscaling model built on top of Knative's autoscaler. The key metric isn't CPU or memory — it's concurrent requests per container instance.

Cloud Run Autoscaler Decision Flow

The formula is straightforward:

Target instances = Current requests / (Target utilization × Max concurrency)

With default settings (target utilization of 60%, max concurrency of 80):

Target instances = 1000 concurrent requests / (0.6 × 80) = 21 instances

But the autoscaler doesn't react instantaneously. It operates on a tick-based system with a 2-second observation window and a 60-second stable window before scaling down.

Production Scaling Data

Here's what we measured across our service fleet over 30 days:

MetricDefault ConfigTuned Config
Scale-up latency (0→1)2.1s avg0s (min instances)
Scale-up latency (1→10)4.8s avg3.2s avg
Scale-up latency (10→50)8.3s avg6.1s avg
Cold start duration1.8s avg0.9s avg
Request timeout rate0.3%0.02%
Monthly cost (per service)$847$1,120

The tuned configuration costs 32% more but eliminates virtually all timeout errors.

Concurrency Tuning Strategy

The default max concurrency of 80 is rarely optimal. For CPU-bound workloads, we found the sweet spot through systematic testing:

# cloud-run-service.yaml
apiVersion: serving.knative.dev/v1
kind: Service
metadata:
  name: payment-processor
spec:
  template:
    metadata:
      annotations:
        autoscaling.knative.dev/minScale: "3"
        autoscaling.knative.dev/maxScale: "100"
    spec:
      containerConcurrency: 25
      timeoutSeconds: 30
      containers:
        - image: gcr.io/project/payment-processor:v2.4.1
          resources:
            limits:
              cpu: "2"
              memory: "1Gi"

For I/O-bound services (database queries, external API calls), higher concurrency works:

spec:
  template:
    spec:
      containerConcurrency: 200
      containers:
        - image: gcr.io/project/api-gateway:v3.1.0
          resources:
            limits:
              cpu: "1"
              memory: "512Mi"

Cold Start Mitigation

Cold starts are the silent killer of serverless latency. We implemented a three-layer strategy:

Layer 1: Minimum Instances

Setting minScale to a baseline that handles your steady-state traffic eliminates cold starts for normal load. We use our p25 traffic level as the minimum.

Layer 2: CPU Boost

Cloud Run's startup CPU boost feature allocates additional CPU during container startup:

gcloud run services update payment-processor \
  --cpu-boost \
  --min-instances 3 \
  --region us-central1

This reduced our cold start from 1.8s to 0.9s — a 50% improvement.

Layer 3: Lazy Initialization

We restructured our application to defer non-critical initialization:

// Before: Everything initialized at startup
class PaymentService {
  private db: Database;
  private cache: Redis;
  private validator: SchemaValidator;

  constructor() {
    this.db = new Database(config.dbUrl);        // 400ms
    this.cache = new Redis(config.redisUrl);     // 200ms
    this.validator = new SchemaValidator(schemas); // 300ms
  }
}

// After: Critical path only at startup, rest lazy-loaded
class PaymentService {
  private _db?: Database;
  private _cache?: Redis;
  private validator: SchemaValidator;

  constructor() {
    // Only schema validation is needed for every request
    this.validator = new SchemaValidator(schemas);
  }

  private get db(): Database {
    if (!this._db) {
      this._db = new Database(config.dbUrl);
    }
    return this._db;
  }

  private get cache(): Redis {
    if (!this._cache) {
      this._cache = new Redis(config.redisUrl);
    }
    return this._cache;
  }
}

Request-Based Metrics Dashboard

We built a custom monitoring stack to understand scaling decisions in real-time:

# Query to track autoscaler behavior
gcloud logging read '
  resource.type="cloud_run_revision"
  AND resource.labels.service_name="payment-processor"
  AND jsonPayload.message=~"Autoscaler"
' --limit 100 --format json | \
  jq '[.[] | {
    timestamp: .timestamp,
    instances: .jsonPayload.desiredInstances,
    concurrent_requests: .jsonPayload.observedConcurrency
  }]'

Cloud Run Scaling Response Time

The 60-Second Stability Window

One critical behavior that isn't obvious from documentation: Cloud Run won't scale down until metrics have been stable for 60 seconds. This means after a traffic spike, you're paying for peak capacity for at least a minute after traffic normalizes.

For services with frequent short bursts, this creates a "ratcheting" effect where instance count stays elevated. Our solution: accept the cost during business hours (traffic is bursty) and set aggressive max-instances limits during off-peak via scheduled updates.

Key Takeaways

  1. Set concurrency based on workload type, not defaults. CPU-bound services need 10-30 concurrency; I/O-bound can handle 100-250.
  2. Minimum instances aren't waste — they're insurance. Calculate the cost of cold-start-induced timeouts vs. idle instance cost.
  3. Monitor the autoscaler itself, not just your application metrics. The decision-making lag is where latency hides.
  4. Cold start optimization is cumulative. CPU boost + lazy init + minimum instances together cut our cold starts by 85%.
  5. Budget for the 60-second stability window. Your cost model needs to account for post-spike tail costs.

Cloud Run's autoscaler is remarkably good once you understand its decision model. The gap between "it works" and "it works well in production" is entirely in the tuning.

Comments

    No comments yet. Be the first to share your thoughts.