Building Resilient AI Gateways with Multi-Provider Fallback
Architecture and implementation of production AI gateways — rate limiting, circuit breakers, multi-provider failover, and cost-aware routing for LLM APIs.

Every production AI system is one provider outage away from a total service failure. In March 2026, a major LLM provider had a 4-hour outage. Teams with multi-provider gateways served traffic without interruption. Teams without them had 4 hours of downtime. Here's how to build the gateway.
Why AI APIs need a gateway
LLM APIs differ from traditional APIs in ways that break standard gateway patterns:
- Variable latency: 200ms to 30s depending on output length
- Token-based rate limits: Not just requests/second, but tokens/minute
- Non-idempotent: Same prompt can produce different outputs
- Cost asymmetry: One runaway request can cost $5+ (long context + long output)
- Provider-specific limits: Each provider has different rate limit structures
A gateway purpose-built for AI APIs handles all of these.
Architecture
┌─────────────┐ ┌──────────────────┐ ┌──────────────────┐
│ Application │────▶│ AI Gateway │────▶│ Provider Pool │
└─────────────┘ │ │ │ ┌─────────────┐ │
│ • Rate Limiter │ │ │ Anthropic │ │
│ • Circuit Break │ │ │ OpenAI │ │
│ • Cost Guard │ │ │ Google │ │
│ • Retry Logic │ │ │ Azure OAI │ │
│ • Audit Log │ │ └─────────────┘ │
└──────────────────┘ └──────────────────┘
Multi-provider failover
The core pattern: maintain a priority-ordered list of providers, route to the primary, and failover to alternates on failure.
from dataclasses import dataclass, field
from enum import Enum
import asyncio
import time
class ProviderStatus(Enum):
HEALTHY = "healthy"
DEGRADED = "degraded"
DOWN = "down"
@dataclass
class ProviderConfig:
name: str
model: str
priority: int
rate_limit_rpm: int
rate_limit_tpm: int
cost_per_1k_input: float
cost_per_1k_output: float
status: ProviderStatus = ProviderStatus.HEALTHY
failure_count: int = 0
last_failure: float = 0.0
circuit_open_until: float = 0.0
class AIGateway:
def __init__(self, providers: list[ProviderConfig]):
self.providers = sorted(providers, key=lambda p: p.priority)
self._request_counts: dict[str, list[float]] = {}
async def generate(self, messages: list[dict], **kwargs) -> dict:
errors = []
for provider in self._get_available_providers():
if not self._check_rate_limit(provider):
continue
try:
result = await self._call_provider(provider, messages, **kwargs)
self._record_success(provider)
return {
"content": result["content"],
"provider": provider.name,
"model": provider.model,
"usage": result["usage"],
"cost": self._calculate_cost(provider, result["usage"]),
}
except RateLimitError:
self._record_rate_limit(provider)
errors.append(f"{provider.name}: rate limited")
except ProviderError as e:
self._record_failure(provider)
errors.append(f"{provider.name}: {e}")
raise AllProvidersFailedError(errors)
def _get_available_providers(self) -> list[ProviderConfig]:
now = time.time()
return [
p for p in self.providers
if p.status != ProviderStatus.DOWN
and p.circuit_open_until < now
]
def _check_rate_limit(self, provider: ProviderConfig) -> bool:
now = time.time()
window = self._request_counts.get(provider.name, [])
# Remove requests outside the 60-second window
window = [t for t in window if now - t < 60]
self._request_counts[provider.name] = window
return len(window) < provider.rate_limit_rpm
def _record_failure(self, provider: ProviderConfig):
provider.failure_count += 1
provider.last_failure = time.time()
# Circuit breaker: open after 3 consecutive failures
if provider.failure_count >= 3:
provider.circuit_open_until = time.time() + 60 # 60s cooldown
provider.status = ProviderStatus.DOWN
def _record_success(self, provider: ProviderConfig):
provider.failure_count = 0
provider.status = ProviderStatus.HEALTHY
Token-aware rate limiting
Traditional rate limiters count requests. AI APIs need token-aware limiting that accounts for both input and estimated output tokens.
interface TokenBucket {
tokens: number;
maxTokens: number;
refillRate: number; // tokens per second
lastRefill: number;
}
class TokenAwareRateLimiter {
private buckets: Map<string, TokenBucket> = new Map();
constructor(
private providers: Map<string, { tpm: number; rpm: number }>,
) {
for (const [name, limits] of providers) {
this.buckets.set(name, {
tokens: limits.tpm,
maxTokens: limits.tpm,
refillRate: limits.tpm / 60, // per second
lastRefill: Date.now(),
});
}
}
canProceed(provider: string, estimatedTokens: number): boolean {
const bucket = this.buckets.get(provider);
if (!bucket) return false;
this.refill(bucket);
return bucket.tokens >= estimatedTokens;
}
consume(provider: string, actualTokens: number): void {
const bucket = this.buckets.get(provider);
if (bucket) {
bucket.tokens = Math.max(0, bucket.tokens - actualTokens);
}
}
private refill(bucket: TokenBucket): void {
const now = Date.now();
const elapsed = (now - bucket.lastRefill) / 1000;
bucket.tokens = Math.min(bucket.maxTokens, bucket.tokens + elapsed * bucket.refillRate);
bucket.lastRefill = now;
}
estimateTokens(messages: Array<{ content: string }>): number {
// Rough estimation: 1 token ≈ 4 characters
const inputTokens = messages.reduce((sum, m) => sum + Math.ceil(m.content.length / 4), 0);
// Estimate output as 50% of input (conservative)
return inputTokens + Math.ceil(inputTokens * 0.5);
}
}
Cost guardrails
Prevent runaway costs with per-request and per-minute spending limits:
| Guard | Threshold | Action |
|---|---|---|
| Per-request cost | >$0.50 | Block + alert |
| Per-minute spend | >$10 | Throttle to 50% |
| Per-hour spend | >$200 | Switch to cheaper model |
| Daily budget | >$2,000 | Emergency shutdown |
| Single-user hourly | >$5 | Rate limit user |
Health checking and circuit breaking
The circuit breaker pattern adapted for AI providers:
Closed (normal): Requests flow to the provider. Track failure rate in sliding window.
Open (tripped): No requests sent. After cooldown period, transition to half-open.
Half-open (testing): Send single probe request. If it succeeds, close circuit. If it fails, reopen.
AI-specific adaptation: Don't count timeouts on long-generation requests as failures. A 25-second response is normal for complex prompts. Only count hard errors (5xx, connection refused, malformed response) toward the circuit breaker threshold.
Model equivalence mapping
When failing over between providers, you need model equivalence:
| Capability Tier | Anthropic | OpenAI | |
|---|---|---|---|
| High reasoning | claude-opus | gpt-4o | gemini-pro |
| Balanced | claude-sonnet | gpt-4o-mini | gemini-flash |
| Fast/cheap | claude-haiku | gpt-4o-mini | gemini-flash |
The gateway maps your application's capability request to the appropriate model per provider. Your application requests "tier: balanced" rather than a specific model name.
Monitoring the gateway
Key metrics to export:
- Provider availability: Uptime per provider, MTTR after outages
- Failover events: How often each backup provider is activated
- Token budget utilization: % of rate limit consumed per provider
- Cost distribution: Spend split across providers
- Latency by provider: Detect degradation before full outage
Key takeaways
- Every production AI system needs multi-provider failover — single-provider dependency is unacceptable risk
- Token-aware rate limiting prevents both quota exhaustion and cost overruns
- Circuit breakers need AI-specific tuning — long latency isn't failure, 5xx is
- Cost guardrails at multiple granularities (request, minute, hour, day) prevent budget disasters
- Abstract model selection to capability tiers, not provider-specific model names
- Monitor failover events — frequent failovers signal an unreliable primary that should be reconsidered
Build the gateway before you need it. The 4-hour outage will happen at the worst possible time.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.