Token Lifecycle Management for 50M+ Daily API Calls: OAuth at Scale

Building a token management system handling 50 million daily API calls with sub-millisecond validation, automated rotation, and zero-downtime revocation

#api-security#oauth#tokens#authentication
Cover image for the article: Token Lifecycle Management for 50M+ Daily API Calls: OAuth at Scale

When your API handles 50 million requests per day, every millisecond of token validation latency multiplies into hours of aggregate computation. Our initial OAuth implementation added 23ms per request for token introspection. At scale, that translated to 13.4 CPU-days wasted daily on authentication overhead. We rebuilt our token lifecycle management to achieve sub-millisecond validation while maintaining security guarantees: short-lived tokens, instant revocation, and automated rotation. Here is the architecture that serves 580 requests per second per node with 0.4ms P99 auth latency.

The Problem: Token Validation as a Bottleneck

Our original implementation followed the textbook OAuth2 pattern: client presents a token, API validates by calling the authorization server's introspection endpoint. This works at low scale. At 50M requests/day (580 RPS average, 2,400 RPS peak), the introspection endpoint became a single point of failure and a latency tax on every request.

Before optimization:

  • Token validation latency: 23ms average, 89ms P99
  • Introspection endpoint: 3 instances handling 580 RPS
  • Single point of failure: introspection downtime = API downtime
  • Token lifetime: 1 hour (long-lived, difficult to revoke)
  • Revocation propagation time: up to 60 minutes

Architecture: Distributed Token Validation

We redesigned the system around three principles: validate locally, revoke globally, and rotate frequently.

Token Lifecycle Architecture

The key architectural decision: use short-lived JWTs for stateless validation (no network call needed) combined with a distributed revocation list for immediate invalidation when security events require it.

Token Issuance: Short-Lived JWTs with Rotation

Tokens are now issued with a 5-minute lifetime. This dramatically reduces the window of exposure for leaked tokens while requiring a robust refresh mechanism:

import { SignJWT, jwtVerify, createLocalJWKSet } from 'jose';
import { randomUUID } from 'crypto';

interface TokenPayload {
  sub: string;          // User/service ID
  scope: string[];      // Authorized scopes
  aud: string;          // Target API
  client_id: string;    // OAuth client
  jti: string;          // Unique token ID (for revocation)
  tier: string;         // Rate limit tier
}

class TokenIssuer {
  private signingKey: CryptoKey;
  private keyId: string;
  private readonly TOKEN_TTL = '5m';
  private readonly REFRESH_TTL = '24h';

  async issueAccessToken(payload: TokenPayload): Promise<string> {
    const now = Math.floor(Date.now() / 1000);
    
    const token = await new SignJWT({
      sub: payload.sub,
      scope: payload.scope,
      client_id: payload.client_id,
      tier: payload.tier,
    })
      .setProtectedHeader({ alg: 'ES256', kid: this.keyId })
      .setIssuedAt(now)
      .setExpirationTime(this.TOKEN_TTL)
      .setAudience(payload.aud)
      .setJti(randomUUID())
      .setIssuer('https://auth.company.com')
      .sign(this.signingKey);

    // Track issued token for audit and revocation
    await this.trackIssuedToken({
      jti: payload.jti,
      sub: payload.sub,
      issued_at: now,
      expires_at: now + 300,  // 5 minutes
      client_id: payload.client_id,
    });

    return token;
  }

  async issueRefreshToken(userId: string, clientId: string): Promise<string> {
    const refreshToken = randomUUID();
    
    // Store refresh token with metadata
    await this.redis.set(
      `refresh:${refreshToken}`,
      JSON.stringify({
        user_id: userId,
        client_id: clientId,
        issued_at: Date.now(),
        rotation_count: 0,
      }),
      'EX', 86400  // 24-hour expiry
    );

    return refreshToken;
  }
}

We use ES256 (ECDSA with P-256) for signing because it produces smaller signatures than RS256 and offers faster verification, which matters at 580 RPS.

Local Token Validation: Sub-Millisecond

The performance breakthrough comes from validating tokens locally without any network call. Each API node caches the JWKS (JSON Web Key Set) and validates tokens using only CPU:

import { jwtVerify, createLocalJWKSet, JWTPayload } from 'jose';
import { LRUCache } from 'lru-cache';

class TokenValidator {
  private jwks: ReturnType<typeof createLocalJWKSet>;
  private revocationCache: LRUCache<string, boolean>;
  private readonly CLOCK_TOLERANCE = 30; // seconds

  constructor(jwksData: object) {
    this.jwks = createLocalJWKSet(jwksData);
    
    // Local revocation cache - synced from Redis pub/sub
    this.revocationCache = new LRUCache({
      max: 100_000,  // Track up to 100K revoked token IDs
      ttl: 1000 * 60 * 6,  // 6 minutes (slightly longer than max token TTL)
    });
  }

  async validateToken(token: string, expectedAudience: string): Promise<TokenPayload> {
    // Step 1: Cryptographic validation (no network call)
    const { payload } = await jwtVerify(token, this.jwks, {
      issuer: 'https://auth.company.com',
      audience: expectedAudience,
      clockTolerance: this.CLOCK_TOLERANCE,
    });

    // Step 2: Check local revocation cache (no network call)
    if (this.revocationCache.has(payload.jti as string)) {
      throw new TokenRevokedError(payload.jti as string);
    }

    // Step 3: Validate scopes against request
    return payload as unknown as TokenPayload;
  }

  // Called by Redis pub/sub subscriber
  onTokenRevoked(jti: string): void {
    this.revocationCache.set(jti, true);
  }
}

Total validation time: cryptographic signature verification (~0.3ms) plus revocation cache lookup (~0.01ms). No network calls. No database queries.

Distributed Revocation: Real-Time Token Invalidation

Short-lived tokens (5 minutes) limit exposure, but sometimes you need instant revocation: compromised credentials, user logout, or suspicious activity detection. We use Redis pub/sub to propagate revocations to all API nodes within 50ms:

import redis
import json
from datetime import datetime, timedelta

class RevocationService:
    """Manages real-time token revocation across all API nodes."""
    
    def __init__(self, redis_url: str):
        self.redis = redis.Redis.from_url(redis_url)
        self.pubsub = self.redis.pubsub()
        self.CHANNEL = "token:revocations"
    
    def revoke_token(self, jti: str, reason: str, actor: str) -> dict:
        """Immediately revoke a specific token by JTI."""
        revocation = {
            "jti": jti,
            "revoked_at": datetime.utcnow().isoformat(),
            "reason": reason,
            "actor": actor,
            "ttl_seconds": 300,  # Match max token lifetime
        }
        
        # Persist revocation (for nodes that may be starting up)
        self.redis.setex(
            f"revoked:{jti}",
            300,  # Expire after max token TTL
            json.dumps(revocation),
        )
        
        # Broadcast to all API nodes immediately
        self.redis.publish(self.CHANNEL, json.dumps(revocation))
        
        return revocation
    
    def revoke_all_for_user(self, user_id: str, reason: str) -> int:
        """Revoke all active tokens for a user (e.g., password change)."""
        # Get all active token JTIs for this user
        active_tokens = self.redis.smembers(f"user_tokens:{user_id}")
        
        count = 0
        for jti in active_tokens:
            self.revoke_token(
                jti.decode(),
                reason=f"user_revocation: {reason}",
                actor="system",
            )
            count += 1
        
        return count
    
    def revoke_by_client(self, client_id: str, reason: str) -> int:
        """Revoke all tokens issued to a specific OAuth client."""
        active_tokens = self.redis.smembers(f"client_tokens:{client_id}")
        
        count = 0
        for jti in active_tokens:
            self.revoke_token(jti.decode(), reason=reason, actor="security_system")
            count += 1
        
        return count

Revocation propagation time: Redis pub/sub delivers to all nodes within 50ms. Combined with the 5-minute token TTL, even if a revocation message is missed, the token expires naturally within minutes.

Key Rotation Without Downtime

Signing keys rotate every 30 days. The rotation process ensures zero-downtime validation:

Key Rotation Timeline

  1. Day 0: New key generated and added to JWKS as secondary
  2. Day 1: New key promoted to primary (new tokens signed with it)
  3. Day 1-30: Old key remains in JWKS for validation of existing tokens
  4. Day 30: Old key removed from JWKS (all tokens signed with it have expired)

API nodes refresh JWKS every 60 seconds, ensuring they always have both keys available during rotation windows.

Rate Limiting by Token Tier

Tokens carry a tier claim that determines rate limits. This moves rate limit decisions to the token validation layer:

const TIER_LIMITS: Record<string, { rpm: number; burst: number }> = {
  free:       { rpm: 60,    burst: 10 },
  starter:    { rpm: 600,   burst: 50 },
  business:   { rpm: 6000,  burst: 200 },
  enterprise: { rpm: 60000, burst: 1000 },
  internal:   { rpm: 0,     burst: 0 },  // Unlimited
};

function getRateLimit(token: TokenPayload): { rpm: number; burst: number } {
  return TIER_LIMITS[token.tier] || TIER_LIMITS.free;
}

Embedding tier in the token avoids a database lookup on every request to determine the caller's plan.

Monitoring and Anomaly Detection

We monitor token operations for security anomalies:

SignalThresholdAction
Failed validations per client>100/minAlert + auto-revoke client
Token refresh rate per user>12/hourAlert security team
Tokens issued from new IPAnyLog for review
Revocations per user>5/dayAccount lockout + investigation
JWKS fetch failures>3 consecutiveCritical alert

Performance Results

MetricBeforeAfterImprovement
Token validation P5023ms0.3ms-98.7%
Token validation P9989ms0.4ms-99.6%
Auth CPU overhead (daily)13.4 CPU-days0.18 CPU-days-98.7%
Revocation propagation timeUp to 60 min< 50ms-99.99%
Token exposure window60 minutes5 minutes-92%
Auth-related outages (annual)40-100%
Single point of failureYes (introspection)No (local validation)Eliminated

Lessons Learned

Short tokens eliminate most revocation needs. With 5-minute tokens, 94% of "revocation" scenarios are handled by simply not refreshing. Explicit revocation is only needed for active security incidents.

ES256 over RS256 at scale. The smaller signature size (64 bytes vs 256 bytes) saves bandwidth, and verification is faster on modern hardware. At 50M requests/day, this adds up.

Redis pub/sub is fast enough for revocation. We considered Kafka for revocation propagation but Redis pub/sub's 50ms delivery time is well within acceptable bounds for a 5-minute token TTL.

Embed authorization data in the token. Tier, scopes, and client ID in the token payload eliminate three database lookups per request. At 580 RPS, that is 1,740 fewer database queries per second.

Conclusion

Token management at 50M daily requests requires fundamentally different architecture than textbook OAuth. The combination of short-lived JWTs (local validation, no network call), distributed revocation (Redis pub/sub for instant invalidation), and automated key rotation (zero-downtime every 30 days) gives us sub-millisecond auth latency with stronger security guarantees than our original long-lived token approach. The 98.7% latency reduction freed compute resources equivalent to 13 CPU-days daily, paying for the entire redesign in the first week.

Comments

    No comments yet. Be the first to share your thoughts.