Building Resilient AI Gateways with Multi-Provider Fallback

Architecture and implementation of production AI gateways — rate limiting, circuit breakers, multi-provider failover, and cost-aware routing for LLM APIs.

#ai#api-gateway#rate-limiting#failover
Cover image for the article: Building Resilient AI Gateways with Multi-Provider Fallback

Every production AI system is one provider outage away from a total service failure. In March 2026, a major LLM provider had a 4-hour outage. Teams with multi-provider gateways served traffic without interruption. Teams without them had 4 hours of downtime. Here's how to build the gateway.

Why AI APIs need a gateway

LLM APIs differ from traditional APIs in ways that break standard gateway patterns:

  • Variable latency: 200ms to 30s depending on output length
  • Token-based rate limits: Not just requests/second, but tokens/minute
  • Non-idempotent: Same prompt can produce different outputs
  • Cost asymmetry: One runaway request can cost $5+ (long context + long output)
  • Provider-specific limits: Each provider has different rate limit structures

A gateway purpose-built for AI APIs handles all of these.

Architecture

AI Gateway Architecture

┌─────────────┐     ┌──────────────────┐     ┌──────────────────┐
│ Application │────▶│   AI Gateway     │────▶│  Provider Pool   │
└─────────────┘     │                  │     │  ┌─────────────┐ │
                    │  • Rate Limiter  │     │  │  Anthropic  │ │
                    │  • Circuit Break │     │  │  OpenAI     │ │
                    │  • Cost Guard    │     │  │  Google     │ │
                    │  • Retry Logic   │     │  │  Azure OAI  │ │
                    │  • Audit Log     │     │  └─────────────┘ │
                    └──────────────────┘     └──────────────────┘

Multi-provider failover

The core pattern: maintain a priority-ordered list of providers, route to the primary, and failover to alternates on failure.

from dataclasses import dataclass, field
from enum import Enum
import asyncio
import time

class ProviderStatus(Enum):
    HEALTHY = "healthy"
    DEGRADED = "degraded"
    DOWN = "down"

@dataclass
class ProviderConfig:
    name: str
    model: str
    priority: int
    rate_limit_rpm: int
    rate_limit_tpm: int
    cost_per_1k_input: float
    cost_per_1k_output: float
    status: ProviderStatus = ProviderStatus.HEALTHY
    failure_count: int = 0
    last_failure: float = 0.0
    circuit_open_until: float = 0.0

class AIGateway:
    def __init__(self, providers: list[ProviderConfig]):
        self.providers = sorted(providers, key=lambda p: p.priority)
        self._request_counts: dict[str, list[float]] = {}

    async def generate(self, messages: list[dict], **kwargs) -> dict:
        errors = []

        for provider in self._get_available_providers():
            if not self._check_rate_limit(provider):
                continue

            try:
                result = await self._call_provider(provider, messages, **kwargs)
                self._record_success(provider)
                return {
                    "content": result["content"],
                    "provider": provider.name,
                    "model": provider.model,
                    "usage": result["usage"],
                    "cost": self._calculate_cost(provider, result["usage"]),
                }
            except RateLimitError:
                self._record_rate_limit(provider)
                errors.append(f"{provider.name}: rate limited")
            except ProviderError as e:
                self._record_failure(provider)
                errors.append(f"{provider.name}: {e}")

        raise AllProvidersFailedError(errors)

    def _get_available_providers(self) -> list[ProviderConfig]:
        now = time.time()
        return [
            p for p in self.providers
            if p.status != ProviderStatus.DOWN
            and p.circuit_open_until < now
        ]

    def _check_rate_limit(self, provider: ProviderConfig) -> bool:
        now = time.time()
        window = self._request_counts.get(provider.name, [])
        # Remove requests outside the 60-second window
        window = [t for t in window if now - t < 60]
        self._request_counts[provider.name] = window
        return len(window) < provider.rate_limit_rpm

    def _record_failure(self, provider: ProviderConfig):
        provider.failure_count += 1
        provider.last_failure = time.time()
        # Circuit breaker: open after 3 consecutive failures
        if provider.failure_count >= 3:
            provider.circuit_open_until = time.time() + 60  # 60s cooldown
            provider.status = ProviderStatus.DOWN

    def _record_success(self, provider: ProviderConfig):
        provider.failure_count = 0
        provider.status = ProviderStatus.HEALTHY

Token-aware rate limiting

Traditional rate limiters count requests. AI APIs need token-aware limiting that accounts for both input and estimated output tokens.

interface TokenBucket {
  tokens: number;
  maxTokens: number;
  refillRate: number; // tokens per second
  lastRefill: number;
}

class TokenAwareRateLimiter {
  private buckets: Map<string, TokenBucket> = new Map();

  constructor(
    private providers: Map<string, { tpm: number; rpm: number }>,
  ) {
    for (const [name, limits] of providers) {
      this.buckets.set(name, {
        tokens: limits.tpm,
        maxTokens: limits.tpm,
        refillRate: limits.tpm / 60, // per second
        lastRefill: Date.now(),
      });
    }
  }

  canProceed(provider: string, estimatedTokens: number): boolean {
    const bucket = this.buckets.get(provider);
    if (!bucket) return false;

    this.refill(bucket);
    return bucket.tokens >= estimatedTokens;
  }

  consume(provider: string, actualTokens: number): void {
    const bucket = this.buckets.get(provider);
    if (bucket) {
      bucket.tokens = Math.max(0, bucket.tokens - actualTokens);
    }
  }

  private refill(bucket: TokenBucket): void {
    const now = Date.now();
    const elapsed = (now - bucket.lastRefill) / 1000;
    bucket.tokens = Math.min(bucket.maxTokens, bucket.tokens + elapsed * bucket.refillRate);
    bucket.lastRefill = now;
  }

  estimateTokens(messages: Array<{ content: string }>): number {
    // Rough estimation: 1 token ≈ 4 characters
    const inputTokens = messages.reduce((sum, m) => sum + Math.ceil(m.content.length / 4), 0);
    // Estimate output as 50% of input (conservative)
    return inputTokens + Math.ceil(inputTokens * 0.5);
  }
}

Cost guardrails

Prevent runaway costs with per-request and per-minute spending limits:

GuardThresholdAction
Per-request cost>$0.50Block + alert
Per-minute spend>$10Throttle to 50%
Per-hour spend>$200Switch to cheaper model
Daily budget>$2,000Emergency shutdown
Single-user hourly>$5Rate limit user

Health checking and circuit breaking

The circuit breaker pattern adapted for AI providers:

Closed (normal): Requests flow to the provider. Track failure rate in sliding window.

Open (tripped): No requests sent. After cooldown period, transition to half-open.

Half-open (testing): Send single probe request. If it succeeds, close circuit. If it fails, reopen.

AI-specific adaptation: Don't count timeouts on long-generation requests as failures. A 25-second response is normal for complex prompts. Only count hard errors (5xx, connection refused, malformed response) toward the circuit breaker threshold.

Model equivalence mapping

When failing over between providers, you need model equivalence:

Capability TierAnthropicOpenAIGoogle
High reasoningclaude-opusgpt-4ogemini-pro
Balancedclaude-sonnetgpt-4o-minigemini-flash
Fast/cheapclaude-haikugpt-4o-minigemini-flash

The gateway maps your application's capability request to the appropriate model per provider. Your application requests "tier: balanced" rather than a specific model name.

Monitoring the gateway

Key metrics to export:

  • Provider availability: Uptime per provider, MTTR after outages
  • Failover events: How often each backup provider is activated
  • Token budget utilization: % of rate limit consumed per provider
  • Cost distribution: Spend split across providers
  • Latency by provider: Detect degradation before full outage

Key takeaways

  • Every production AI system needs multi-provider failover — single-provider dependency is unacceptable risk
  • Token-aware rate limiting prevents both quota exhaustion and cost overruns
  • Circuit breakers need AI-specific tuning — long latency isn't failure, 5xx is
  • Cost guardrails at multiple granularities (request, minute, hour, day) prevent budget disasters
  • Abstract model selection to capability tiers, not provider-specific model names
  • Monitor failover events — frequent failovers signal an unreliable primary that should be reconsidered

Build the gateway before you need it. The 4-hour outage will happen at the worst possible time.

Comments

    No comments yet. Be the first to share your thoughts.