AWS Global Accelerator: Reducing Global API Latency by 60% with Anycast Routing
How we used AWS Global Accelerator to cut global API latency by 60%, with real benchmarks, architecture decisions, and cost analysis across 12 regions.

When your API serves customers across 30+ countries, every millisecond of latency compounds into user frustration, abandoned requests, and lost revenue. We were seeing P95 latencies of 800ms for users in Southeast Asia hitting our US-East-1 origin. After deploying AWS Global Accelerator with a multi-region backend strategy, we brought that down to 320ms — a 60% reduction that directly improved our conversion metrics.
This is the story of how we got there, the mistakes we made along the way, and the exact architecture that delivered results.
The Problem: Geographic Distance Kills Performance
Our platform served a real-time collaboration API from a primary deployment in us-east-1. As we expanded internationally, the complaints started rolling in:
- Southeast Asian users: 750-900ms P95 latency
- European users: 280-350ms P95 latency
- South American users: 400-500ms P95 latency
The physics of light traveling through fiber optics meant our single-region architecture had a hard floor on performance. CDNs helped for static content, but our API calls — authenticated, dynamic, and stateful — needed a different solution.
Why Not Just Multi-Region?
The obvious answer is "deploy everywhere." But multi-region comes with massive complexity:
| Approach | Latency Improvement | Complexity | Monthly Cost Increase |
|---|---|---|---|
| Single region + CDN | 10-20% for APIs | Low | $500-2K |
| CloudFront + Lambda@Edge | 30-40% | Medium | $3-8K |
| Full multi-region active-active | 50-70% | Very High | $15-40K |
| Global Accelerator + regional backends | 50-65% | Medium | $5-12K |
We chose Global Accelerator because it offered the latency benefits of multi-region without requiring us to solve distributed consistency for our entire data layer on day one.
Architecture: Anycast Routing with Regional Processing
The architecture works in three layers:
- Anycast entry points: Global Accelerator provides static IP addresses that route traffic to the nearest AWS edge location (over 90 globally)
- AWS backbone transit: Traffic travels over the AWS private backbone instead of the public internet
- Regional endpoint groups: Requests arrive at the optimal regional backend based on health, proximity, and traffic dial settings
// CDK infrastructure for Global Accelerator with multi-region endpoints
import * as cdk from 'aws-cdk-lib';
import * as globalaccelerator from 'aws-cdk-lib/aws-globalaccelerator';
import * as ga_endpoints from 'aws-cdk-lib/aws-globalaccelerator-endpoints';
const accelerator = new globalaccelerator.Accelerator(this, 'ApiAccelerator', {
acceleratorName: 'api-global-accelerator',
ipAddressType: globalaccelerator.IpAddressType.DUAL_STACK,
});
const listener = accelerator.addListener('ApiListener', {
portRanges: [{ fromPort: 443, toPort: 443 }],
protocol: globalaccelerator.ConnectionProtocol.TCP,
clientAffinity: globalaccelerator.ClientAffinity.SOURCE_IP,
});
// Primary region - US East
const usEastGroup = listener.addEndpointGroup('UsEastGroup', {
region: 'us-east-1',
trafficDialPercentage: 100,
healthCheckPath: '/health',
healthCheckIntervalSeconds: 10,
thresholdCount: 3,
});
usEastGroup.addEndpoint(
new ga_endpoints.ApplicationLoadBalancerEndpoint(usEastAlb, {
weight: 128,
preserveClientIp: true,
})
);
// Asia Pacific region
const apSoutheastGroup = listener.addEndpointGroup('ApSoutheastGroup', {
region: 'ap-southeast-1',
trafficDialPercentage: 100,
healthCheckPath: '/health',
healthCheckIntervalSeconds: 10,
thresholdCount: 3,
});
The AWS Backbone Advantage
The single biggest improvement came not from being "closer" to users, but from avoiding the public internet entirely. When a request from Singapore hits the nearest AWS edge location, it travels over AWS's private fiber network to the destination region.
We measured the difference:
| Route | Singapore to US-East-1 | Singapore to AP-Southeast-1 |
|---|---|---|
| Public internet | 280-350ms | 45-80ms |
| AWS backbone (via GA) | 180-220ms | 8-15ms |
| Improvement | 35-40% | 75-85% |
The backbone routing alone — without any multi-region deployment — reduced latency by 35-40% for long-distance routes. Adding regional backends brought the total improvement to 60%.
Client Affinity and Session Consistency
One challenge with anycast routing is maintaining session consistency. A user's requests might route to different regions if network conditions change. We solved this with source IP client affinity:
// Global Accelerator handles affinity at the network layer
// But we also implemented application-level session routing
const sessionRouter = {
getTargetRegion(sessionId: string, sourceIp: string): string {
// Check if session has an established region
const cachedRegion = await redis.get(`session:${sessionId}:region`);
if (cachedRegion) return cachedRegion;
// New session - let Global Accelerator handle routing
// Record the region for future cross-region consistency
const currentRegion = process.env.AWS_REGION;
await redis.setex(`session:${sessionId}:region`, 3600, currentRegion);
return currentRegion;
}
};
Traffic Dial: Controlled Rollout Across Regions
We didn't flip everything on at once. Global Accelerator's traffic dial let us gradually shift load to new regional backends:
Week 1: Deploy AP-Southeast-1 backend, traffic dial at 10% Week 2: Monitor error rates and latency, increase to 50% Week 3: Full traffic (100%) with automatic failover configured Week 4: Deploy EU-West-1, repeat the process
This controlled approach caught a database replication lag issue in week 2 that would have impacted all Asian users if we'd gone to 100% immediately.
Real-World Benchmarks: Before and After
We ran synthetic monitoring from 12 global locations for 30 days before and after the migration:
| Region | Before (P50/P95) | After (P50/P95) | Improvement |
|---|---|---|---|
| Singapore | 420ms / 850ms | 85ms / 180ms | 79% |
| Tokyo | 380ms / 720ms | 95ms / 210ms | 71% |
| Mumbai | 350ms / 680ms | 120ms / 250ms | 63% |
| London | 140ms / 280ms | 55ms / 120ms | 57% |
| Frankfurt | 150ms / 300ms | 60ms / 130ms | 57% |
| Sydney | 450ms / 900ms | 110ms / 240ms | 73% |
| Sao Paulo | 220ms / 450ms | 95ms / 200ms | 56% |
Cost Analysis: Is It Worth It?
Global Accelerator pricing has two components:
- Fixed fee: $0.025/hour per accelerator (~$18/month)
- Data transfer premium: $0.015-0.035/GB depending on region (on top of standard transfer)
For our traffic profile (50TB/month egress, 200M requests/day):
| Cost Component | Monthly Cost |
|---|---|
| Accelerator fixed fee | $18 |
| Data transfer premium (50TB) | $1,250 |
| Additional regional ALBs | $450 |
| Regional compute (Fargate) | $3,200 |
| Cross-region data sync | $800 |
| Total additional cost | $5,718 |
Against a measurable improvement in conversion rate (+2.3% for Asian users) worth approximately $45K/month in additional revenue, the ROI was clear within the first billing cycle.
Failover and Health Checks
Global Accelerator's health checking is aggressive and fast. We configured it to detect failures within 30 seconds:
// Health check configuration for fast failover
const endpointGroup = listener.addEndpointGroup('PrimaryGroup', {
healthCheckPath: '/health/deep',
healthCheckIntervalSeconds: 10,
thresholdCount: 3, // 3 consecutive failures = unhealthy
healthCheckProtocol: globalaccelerator.HealthCheckProtocol.HTTPS,
});
During a planned maintenance window, Global Accelerator shifted all traffic away from us-east-1 within 27 seconds of the first failed health check. Users experienced a single retry at most.
Mistakes We Made
1. Not enabling flow logs initially. Global Accelerator flow logs are essential for debugging routing decisions. Enable them from day one.
2. Forgetting about WebSocket connections. Our real-time features used WebSockets, which require TCP listener configuration and proper idle timeout settings.
3. Underestimating DNS propagation. We kept our old DNS records active for 2 weeks after migration. Some corporate DNS resolvers cache aggressively.
When NOT to Use Global Accelerator
Global Accelerator is not always the right choice:
- Static content: Use CloudFront instead — it's cheaper and caches at the edge
- Single-region, single-country apps: The overhead isn't worth it
- UDP-heavy workloads with small payloads: Consider CloudFront or direct regional endpoints
- Budget under $1K/month for networking: The premium may not justify the improvement
Key Takeaways
-
AWS backbone routing alone provides 35-40% latency reduction for cross-continental traffic without any application changes.
-
Anycast + regional backends is the sweet spot between full multi-region complexity and single-region limitations.
-
Traffic dials enable safe rollouts — always start at 10% when adding new regional backends.
-
Measure from real user locations — synthetic monitoring from 10+ geographic points gives you ground truth.
-
The ROI calculation matters — for us, $5.7K/month in infrastructure delivered $45K/month in conversion improvements. But this only works if you have significant international traffic.
-
Plan for session consistency — anycast routing can shift users between regions. Build application-level affinity as a safety net.
Global Accelerator sits in a sweet spot for teams that need global performance without the operational burden of full active-active multi-region. Start with the backbone routing benefit, then add regional backends as your data layer supports it.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.