Token Lifecycle Management for 50M+ Daily API Calls: OAuth at Scale
Building a token management system handling 50 million daily API calls with sub-millisecond validation, automated rotation, and zero-downtime revocation

When your API handles 50 million requests per day, every millisecond of token validation latency multiplies into hours of aggregate computation. Our initial OAuth implementation added 23ms per request for token introspection. At scale, that translated to 13.4 CPU-days wasted daily on authentication overhead. We rebuilt our token lifecycle management to achieve sub-millisecond validation while maintaining security guarantees: short-lived tokens, instant revocation, and automated rotation. Here is the architecture that serves 580 requests per second per node with 0.4ms P99 auth latency.
The Problem: Token Validation as a Bottleneck
Our original implementation followed the textbook OAuth2 pattern: client presents a token, API validates by calling the authorization server's introspection endpoint. This works at low scale. At 50M requests/day (580 RPS average, 2,400 RPS peak), the introspection endpoint became a single point of failure and a latency tax on every request.
Before optimization:
- Token validation latency: 23ms average, 89ms P99
- Introspection endpoint: 3 instances handling 580 RPS
- Single point of failure: introspection downtime = API downtime
- Token lifetime: 1 hour (long-lived, difficult to revoke)
- Revocation propagation time: up to 60 minutes
Architecture: Distributed Token Validation
We redesigned the system around three principles: validate locally, revoke globally, and rotate frequently.
The key architectural decision: use short-lived JWTs for stateless validation (no network call needed) combined with a distributed revocation list for immediate invalidation when security events require it.
Token Issuance: Short-Lived JWTs with Rotation
Tokens are now issued with a 5-minute lifetime. This dramatically reduces the window of exposure for leaked tokens while requiring a robust refresh mechanism:
import { SignJWT, jwtVerify, createLocalJWKSet } from 'jose';
import { randomUUID } from 'crypto';
interface TokenPayload {
sub: string; // User/service ID
scope: string[]; // Authorized scopes
aud: string; // Target API
client_id: string; // OAuth client
jti: string; // Unique token ID (for revocation)
tier: string; // Rate limit tier
}
class TokenIssuer {
private signingKey: CryptoKey;
private keyId: string;
private readonly TOKEN_TTL = '5m';
private readonly REFRESH_TTL = '24h';
async issueAccessToken(payload: TokenPayload): Promise<string> {
const now = Math.floor(Date.now() / 1000);
const token = await new SignJWT({
sub: payload.sub,
scope: payload.scope,
client_id: payload.client_id,
tier: payload.tier,
})
.setProtectedHeader({ alg: 'ES256', kid: this.keyId })
.setIssuedAt(now)
.setExpirationTime(this.TOKEN_TTL)
.setAudience(payload.aud)
.setJti(randomUUID())
.setIssuer('https://auth.company.com')
.sign(this.signingKey);
// Track issued token for audit and revocation
await this.trackIssuedToken({
jti: payload.jti,
sub: payload.sub,
issued_at: now,
expires_at: now + 300, // 5 minutes
client_id: payload.client_id,
});
return token;
}
async issueRefreshToken(userId: string, clientId: string): Promise<string> {
const refreshToken = randomUUID();
// Store refresh token with metadata
await this.redis.set(
`refresh:${refreshToken}`,
JSON.stringify({
user_id: userId,
client_id: clientId,
issued_at: Date.now(),
rotation_count: 0,
}),
'EX', 86400 // 24-hour expiry
);
return refreshToken;
}
}
We use ES256 (ECDSA with P-256) for signing because it produces smaller signatures than RS256 and offers faster verification, which matters at 580 RPS.
Local Token Validation: Sub-Millisecond
The performance breakthrough comes from validating tokens locally without any network call. Each API node caches the JWKS (JSON Web Key Set) and validates tokens using only CPU:
import { jwtVerify, createLocalJWKSet, JWTPayload } from 'jose';
import { LRUCache } from 'lru-cache';
class TokenValidator {
private jwks: ReturnType<typeof createLocalJWKSet>;
private revocationCache: LRUCache<string, boolean>;
private readonly CLOCK_TOLERANCE = 30; // seconds
constructor(jwksData: object) {
this.jwks = createLocalJWKSet(jwksData);
// Local revocation cache - synced from Redis pub/sub
this.revocationCache = new LRUCache({
max: 100_000, // Track up to 100K revoked token IDs
ttl: 1000 * 60 * 6, // 6 minutes (slightly longer than max token TTL)
});
}
async validateToken(token: string, expectedAudience: string): Promise<TokenPayload> {
// Step 1: Cryptographic validation (no network call)
const { payload } = await jwtVerify(token, this.jwks, {
issuer: 'https://auth.company.com',
audience: expectedAudience,
clockTolerance: this.CLOCK_TOLERANCE,
});
// Step 2: Check local revocation cache (no network call)
if (this.revocationCache.has(payload.jti as string)) {
throw new TokenRevokedError(payload.jti as string);
}
// Step 3: Validate scopes against request
return payload as unknown as TokenPayload;
}
// Called by Redis pub/sub subscriber
onTokenRevoked(jti: string): void {
this.revocationCache.set(jti, true);
}
}
Total validation time: cryptographic signature verification (~0.3ms) plus revocation cache lookup (~0.01ms). No network calls. No database queries.
Distributed Revocation: Real-Time Token Invalidation
Short-lived tokens (5 minutes) limit exposure, but sometimes you need instant revocation: compromised credentials, user logout, or suspicious activity detection. We use Redis pub/sub to propagate revocations to all API nodes within 50ms:
import redis
import json
from datetime import datetime, timedelta
class RevocationService:
"""Manages real-time token revocation across all API nodes."""
def __init__(self, redis_url: str):
self.redis = redis.Redis.from_url(redis_url)
self.pubsub = self.redis.pubsub()
self.CHANNEL = "token:revocations"
def revoke_token(self, jti: str, reason: str, actor: str) -> dict:
"""Immediately revoke a specific token by JTI."""
revocation = {
"jti": jti,
"revoked_at": datetime.utcnow().isoformat(),
"reason": reason,
"actor": actor,
"ttl_seconds": 300, # Match max token lifetime
}
# Persist revocation (for nodes that may be starting up)
self.redis.setex(
f"revoked:{jti}",
300, # Expire after max token TTL
json.dumps(revocation),
)
# Broadcast to all API nodes immediately
self.redis.publish(self.CHANNEL, json.dumps(revocation))
return revocation
def revoke_all_for_user(self, user_id: str, reason: str) -> int:
"""Revoke all active tokens for a user (e.g., password change)."""
# Get all active token JTIs for this user
active_tokens = self.redis.smembers(f"user_tokens:{user_id}")
count = 0
for jti in active_tokens:
self.revoke_token(
jti.decode(),
reason=f"user_revocation: {reason}",
actor="system",
)
count += 1
return count
def revoke_by_client(self, client_id: str, reason: str) -> int:
"""Revoke all tokens issued to a specific OAuth client."""
active_tokens = self.redis.smembers(f"client_tokens:{client_id}")
count = 0
for jti in active_tokens:
self.revoke_token(jti.decode(), reason=reason, actor="security_system")
count += 1
return count
Revocation propagation time: Redis pub/sub delivers to all nodes within 50ms. Combined with the 5-minute token TTL, even if a revocation message is missed, the token expires naturally within minutes.
Key Rotation Without Downtime
Signing keys rotate every 30 days. The rotation process ensures zero-downtime validation:
- Day 0: New key generated and added to JWKS as secondary
- Day 1: New key promoted to primary (new tokens signed with it)
- Day 1-30: Old key remains in JWKS for validation of existing tokens
- Day 30: Old key removed from JWKS (all tokens signed with it have expired)
API nodes refresh JWKS every 60 seconds, ensuring they always have both keys available during rotation windows.
Rate Limiting by Token Tier
Tokens carry a tier claim that determines rate limits. This moves rate limit decisions to the token validation layer:
const TIER_LIMITS: Record<string, { rpm: number; burst: number }> = {
free: { rpm: 60, burst: 10 },
starter: { rpm: 600, burst: 50 },
business: { rpm: 6000, burst: 200 },
enterprise: { rpm: 60000, burst: 1000 },
internal: { rpm: 0, burst: 0 }, // Unlimited
};
function getRateLimit(token: TokenPayload): { rpm: number; burst: number } {
return TIER_LIMITS[token.tier] || TIER_LIMITS.free;
}
Embedding tier in the token avoids a database lookup on every request to determine the caller's plan.
Monitoring and Anomaly Detection
We monitor token operations for security anomalies:
| Signal | Threshold | Action |
|---|---|---|
| Failed validations per client | >100/min | Alert + auto-revoke client |
| Token refresh rate per user | >12/hour | Alert security team |
| Tokens issued from new IP | Any | Log for review |
| Revocations per user | >5/day | Account lockout + investigation |
| JWKS fetch failures | >3 consecutive | Critical alert |
Performance Results
| Metric | Before | After | Improvement |
|---|---|---|---|
| Token validation P50 | 23ms | 0.3ms | -98.7% |
| Token validation P99 | 89ms | 0.4ms | -99.6% |
| Auth CPU overhead (daily) | 13.4 CPU-days | 0.18 CPU-days | -98.7% |
| Revocation propagation time | Up to 60 min | < 50ms | -99.99% |
| Token exposure window | 60 minutes | 5 minutes | -92% |
| Auth-related outages (annual) | 4 | 0 | -100% |
| Single point of failure | Yes (introspection) | No (local validation) | Eliminated |
Lessons Learned
Short tokens eliminate most revocation needs. With 5-minute tokens, 94% of "revocation" scenarios are handled by simply not refreshing. Explicit revocation is only needed for active security incidents.
ES256 over RS256 at scale. The smaller signature size (64 bytes vs 256 bytes) saves bandwidth, and verification is faster on modern hardware. At 50M requests/day, this adds up.
Redis pub/sub is fast enough for revocation. We considered Kafka for revocation propagation but Redis pub/sub's 50ms delivery time is well within acceptable bounds for a 5-minute token TTL.
Embed authorization data in the token. Tier, scopes, and client ID in the token payload eliminate three database lookups per request. At 580 RPS, that is 1,740 fewer database queries per second.
Conclusion
Token management at 50M daily requests requires fundamentally different architecture than textbook OAuth. The combination of short-lived JWTs (local validation, no network call), distributed revocation (Redis pub/sub for instant invalidation), and automated key rotation (zero-downtime every 30 days) gives us sub-millisecond auth latency with stronger security guarantees than our original long-lived token approach. The 98.7% latency reduction freed compute resources equivalent to 13 CPU-days daily, paying for the entire redesign in the first week.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.