GCP Cloud Run Auto-Scaling in Production: A Deep Dive into Request-Based Metrics
Analyzing Cloud Run scaling behavior under production traffic with request-based metrics, concurrency tuning, and cold start mitigation strategies.

After running Cloud Run workloads serving 50M+ requests per month across multiple services, I've developed a nuanced understanding of how its autoscaler behaves under real production conditions. The documentation tells you what it does. This article tells you what actually happens when traffic spikes hit at 3 AM.
The Problem: Unpredictable Scaling Under Bursty Traffic
Our payment processing service experienced a pattern that's common in production: traffic that doesn't arrive in smooth curves. We saw 10x spikes within 60-second windows during flash sales, and Cloud Run's default configuration wasn't keeping up. Request latency would spike from p50 of 45ms to p99 of 8.2 seconds during these bursts.
The root cause wasn't Cloud Run itself — it was our misunderstanding of how the autoscaler makes decisions.
How Cloud Run's Autoscaler Actually Works
Cloud Run uses a request-based autoscaling model built on top of Knative's autoscaler. The key metric isn't CPU or memory — it's concurrent requests per container instance.
The formula is straightforward:
Target instances = Current requests / (Target utilization × Max concurrency)
With default settings (target utilization of 60%, max concurrency of 80):
Target instances = 1000 concurrent requests / (0.6 × 80) = 21 instances
But the autoscaler doesn't react instantaneously. It operates on a tick-based system with a 2-second observation window and a 60-second stable window before scaling down.
Production Scaling Data
Here's what we measured across our service fleet over 30 days:
| Metric | Default Config | Tuned Config |
|---|---|---|
| Scale-up latency (0→1) | 2.1s avg | 0s (min instances) |
| Scale-up latency (1→10) | 4.8s avg | 3.2s avg |
| Scale-up latency (10→50) | 8.3s avg | 6.1s avg |
| Cold start duration | 1.8s avg | 0.9s avg |
| Request timeout rate | 0.3% | 0.02% |
| Monthly cost (per service) | $847 | $1,120 |
The tuned configuration costs 32% more but eliminates virtually all timeout errors.
Concurrency Tuning Strategy
The default max concurrency of 80 is rarely optimal. For CPU-bound workloads, we found the sweet spot through systematic testing:
# cloud-run-service.yaml
apiVersion: serving.knative.dev/v1
kind: Service
metadata:
name: payment-processor
spec:
template:
metadata:
annotations:
autoscaling.knative.dev/minScale: "3"
autoscaling.knative.dev/maxScale: "100"
spec:
containerConcurrency: 25
timeoutSeconds: 30
containers:
- image: gcr.io/project/payment-processor:v2.4.1
resources:
limits:
cpu: "2"
memory: "1Gi"
For I/O-bound services (database queries, external API calls), higher concurrency works:
spec:
template:
spec:
containerConcurrency: 200
containers:
- image: gcr.io/project/api-gateway:v3.1.0
resources:
limits:
cpu: "1"
memory: "512Mi"
Cold Start Mitigation
Cold starts are the silent killer of serverless latency. We implemented a three-layer strategy:
Layer 1: Minimum Instances
Setting minScale to a baseline that handles your steady-state traffic eliminates cold starts for normal load. We use our p25 traffic level as the minimum.
Layer 2: CPU Boost
Cloud Run's startup CPU boost feature allocates additional CPU during container startup:
gcloud run services update payment-processor \
--cpu-boost \
--min-instances 3 \
--region us-central1
This reduced our cold start from 1.8s to 0.9s — a 50% improvement.
Layer 3: Lazy Initialization
We restructured our application to defer non-critical initialization:
// Before: Everything initialized at startup
class PaymentService {
private db: Database;
private cache: Redis;
private validator: SchemaValidator;
constructor() {
this.db = new Database(config.dbUrl); // 400ms
this.cache = new Redis(config.redisUrl); // 200ms
this.validator = new SchemaValidator(schemas); // 300ms
}
}
// After: Critical path only at startup, rest lazy-loaded
class PaymentService {
private _db?: Database;
private _cache?: Redis;
private validator: SchemaValidator;
constructor() {
// Only schema validation is needed for every request
this.validator = new SchemaValidator(schemas);
}
private get db(): Database {
if (!this._db) {
this._db = new Database(config.dbUrl);
}
return this._db;
}
private get cache(): Redis {
if (!this._cache) {
this._cache = new Redis(config.redisUrl);
}
return this._cache;
}
}
Request-Based Metrics Dashboard
We built a custom monitoring stack to understand scaling decisions in real-time:
# Query to track autoscaler behavior
gcloud logging read '
resource.type="cloud_run_revision"
AND resource.labels.service_name="payment-processor"
AND jsonPayload.message=~"Autoscaler"
' --limit 100 --format json | \
jq '[.[] | {
timestamp: .timestamp,
instances: .jsonPayload.desiredInstances,
concurrent_requests: .jsonPayload.observedConcurrency
}]'
The 60-Second Stability Window
One critical behavior that isn't obvious from documentation: Cloud Run won't scale down until metrics have been stable for 60 seconds. This means after a traffic spike, you're paying for peak capacity for at least a minute after traffic normalizes.
For services with frequent short bursts, this creates a "ratcheting" effect where instance count stays elevated. Our solution: accept the cost during business hours (traffic is bursty) and set aggressive max-instances limits during off-peak via scheduled updates.
Key Takeaways
- Set concurrency based on workload type, not defaults. CPU-bound services need 10-30 concurrency; I/O-bound can handle 100-250.
- Minimum instances aren't waste — they're insurance. Calculate the cost of cold-start-induced timeouts vs. idle instance cost.
- Monitor the autoscaler itself, not just your application metrics. The decision-making lag is where latency hides.
- Cold start optimization is cumulative. CPU boost + lazy init + minimum instances together cut our cold starts by 85%.
- Budget for the 60-second stability window. Your cost model needs to account for post-spike tail costs.
Cloud Run's autoscaler is remarkably good once you understand its decision model. The gap between "it works" and "it works well in production" is entirely in the tuning.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.