CloudFront Edge Caching Strategies That Cut Our P95 Latency by 73%
Advanced edge caching patterns with CloudFront Functions, origin shield, and cache key optimization that reduced P95 latency from 820ms to 220ms.

When your API serves users across 40 countries, every millisecond of latency maps to measurable revenue impact. Our e-commerce platform was delivering P95 latency of 820ms for product catalog requests to users in Southeast Asia and the Middle East. After implementing a multi-layer edge caching strategy with CloudFront, we dropped that to 220ms globally — a 73% reduction.
This is the architecture, the cache key engineering, and the operational patterns that made it work.
The Starting Architecture
Before optimization, our architecture was straightforward:
- Origin: ALB in us-east-1 fronting ECS services
- CloudFront: Default distribution with minimal caching
- Cache hit ratio: 23% (pathologically low)
- P95 latency: 820ms (global), 340ms (us-east-1)
The low cache hit ratio told us CloudFront was mostly acting as a TLS terminator, not a cache. The root cause: over-broad cache keys that included irrelevant headers and query parameters.
Strategy 1: Cache Key Engineering
The default CloudFront behavior includes ALL query string parameters and select headers in the cache key. For our product API, this meant:
/api/products?id=123&utm_source=google&utm_medium=cpc&fbclid=abc123&_ga=xyz
Every marketing parameter created a unique cache entry for the same content. We implemented a cache policy that strips irrelevant parameters:
{
"CachePolicyConfig": {
"Name": "OptimizedAPIPolicy",
"DefaultTTL": 300,
"MaxTTL": 86400,
"MinTTL": 60,
"ParametersInCacheKeyAndForwardedToOrigin": {
"EnableAcceptEncodingGzip": true,
"EnableAcceptEncodingBrotli": true,
"HeadersConfig": {
"HeaderBehavior": "whitelist",
"Headers": {
"Items": ["Accept-Language", "X-API-Version"],
"Quantity": 2
}
},
"QueryStringsConfig": {
"QueryStringBehavior": "whitelist",
"QueryStrings": {
"Items": ["id", "category", "page", "limit", "sort"],
"Quantity": 5
}
},
"CookiesConfig": {
"CookieBehavior": "none"
}
}
}
}
Impact of cache key optimization alone:
| Metric | Before | After | Improvement |
|---|---|---|---|
| Cache hit ratio | 23% | 61% | +165% |
| Unique cache keys/hour | 2.4M | 340K | -86% |
| Origin requests/sec | 12,400 | 4,800 | -61% |
Strategy 2: Origin Shield
Origin Shield adds a centralized caching layer between edge locations and your origin. Without it, cache misses from 400+ edge locations all hit your origin independently. With it, only one request per cache miss reaches the origin.
{
"OriginShield": {
"Enabled": true,
"OriginShieldRegion": "us-east-1"
}
}
We chose us-east-1 as the shield region because our origin lives there. The results during a cache invalidation event (product price update):
| Scenario | Without Shield | With Shield | Reduction |
|---|---|---|---|
| Origin requests (invalidation) | 12,400/sec | 340/sec | 97.3% |
| Origin CPU during invalidation | 89% | 12% | 86.5% |
| Time to global consistency | 45sec | 8sec | 82.2% |
Strategy 3: CloudFront Functions for Edge Logic
CloudFront Functions execute in ~1ms at edge locations. We use them for URL normalization, A/B test routing, and geo-based cache key augmentation:
// URL Normalization Function
function handler(event) {
var request = event.request;
var uri = request.uri;
var params = request.querystring;
// Normalize trailing slashes
if (uri.endsWith('/') && uri !== '/') {
uri = uri.slice(0, -1);
}
// Sort query parameters for cache consistency
var sortedParams = {};
var allowedParams = ['id', 'category', 'page', 'limit', 'sort'];
Object.keys(params)
.filter(key => allowedParams.includes(key))
.sort()
.forEach(key => {
sortedParams[key] = params[key];
});
request.uri = uri.toLowerCase();
request.querystring = sortedParams;
// Add geo-based cache key header
var country = event.viewer.country || 'US';
var priceRegion = getPriceRegion(country);
request.headers['x-price-region'] = { value: priceRegion };
return request;
}
function getPriceRegion(country) {
var regions = {
'US': 'na', 'CA': 'na', 'MX': 'na',
'GB': 'eu', 'DE': 'eu', 'FR': 'eu',
'AE': 'mena', 'SA': 'mena', 'QA': 'mena',
'SG': 'apac', 'JP': 'apac', 'AU': 'apac'
};
return regions[country] || 'default';
}
This function runs at all 400+ edge locations with sub-millisecond latency. The URL normalization alone improved our cache hit ratio by 12% by eliminating duplicate entries.
Strategy 4: Tiered TTL Architecture
Not all content has the same freshness requirements. We implemented a tiered TTL strategy:
# CloudFront behaviors configuration
behaviors:
# Static assets - aggressive caching
- path_pattern: "/static/*"
cache_policy:
default_ttl: 31536000 # 1 year
min_ttl: 86400
compress: true
# Product catalog - moderate caching with stale-while-revalidate
- path_pattern: "/api/products/*"
cache_policy:
default_ttl: 300 # 5 minutes
max_ttl: 3600
origin_request_policy: "ProductOriginPolicy"
response_headers_policy:
custom_headers:
- name: "Cache-Control"
value: "public, max-age=300, stale-while-revalidate=60"
# User-specific content - no caching
- path_pattern: "/api/user/*"
cache_policy: "CachingDisabled"
origin_request_policy: "AllViewerHeaders"
# Search results - short TTL with Origin Shield
- path_pattern: "/api/search*"
cache_policy:
default_ttl: 60 # 1 minute
min_ttl: 30
origin_shield:
enabled: true
region: "us-east-1"
Strategy 5: Real-Time Invalidation Pipeline
Caching is only useful if stale data does not cause business problems. We built a real-time invalidation pipeline triggered by DynamoDB Streams:
// invalidation-handler.ts
import { CloudFrontClient, CreateInvalidationCommand } from '@aws-sdk/client-cloudfront';
import { DynamoDBStreamEvent } from 'aws-lambda';
const cf = new CloudFrontClient({});
const DISTRIBUTION_ID = process.env.CF_DISTRIBUTION_ID!;
export async function handler(event: DynamoDBStreamEvent): Promise<void> {
const paths = new Set<string>();
for (const record of event.Records) {
if (record.eventName === 'MODIFY' || record.eventName === 'REMOVE') {
const productId = record.dynamodb?.Keys?.PK?.S?.replace('PRODUCT#', '');
if (productId) {
paths.add(`/api/products/${productId}`);
paths.add(`/api/products/${productId}/*`);
}
}
}
if (paths.size === 0) return;
// Batch invalidations (max 3000 paths per request)
const pathArray = Array.from(paths).slice(0, 3000);
await cf.send(new CreateInvalidationCommand({
DistributionId: DISTRIBUTION_ID,
InvalidationBatch: {
CallerReference: `auto-${Date.now()}`,
Paths: {
Quantity: pathArray.length,
Items: pathArray,
},
},
}));
}
The Final Results
After implementing all five strategies:
| Metric | Before | After | Improvement |
|---|---|---|---|
| Cache hit ratio | 23% | 94.2% | +309% |
| P95 latency (global) | 820ms | 220ms | -73.2% |
| P95 latency (MENA) | 1,240ms | 195ms | -84.3% |
| Origin requests/sec | 12,400 | 720 | -94.2% |
| Monthly CloudFront cost | $3,200 | $4,100 | +28% |
| Monthly origin compute | $18,400 | $4,200 | -77.2% |
| Net monthly savings | - | $13,300 | - |
The 28% increase in CloudFront costs was offset by a 77% reduction in origin compute — the origin fleet shrank from 12 ECS tasks to 3.
Key Takeaways
- Cache key engineering is the highest-leverage optimization: Stripping irrelevant query parameters and normalizing URLs can double your hit ratio overnight.
- Origin Shield pays for itself during invalidations: Without it, a single cache invalidation creates a thundering herd to your origin.
- CloudFront Functions for normalization: Sub-millisecond URL and header normalization at the edge eliminates cache fragmentation.
- Tiered TTLs match business requirements: Not all content needs the same freshness. Classify and configure accordingly.
- Measure net cost, not CDN cost: Higher CDN spend that reduces origin compute is almost always a net positive.
The best CDN configuration is one where your origin barely notices traffic exists. At 94% hit ratio, our origin handles the long tail of uncacheable requests while CloudFront serves the world.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.