Monitoring and Governing Claude API Costs at Scale
How we built a cost governance platform that reduced Claude API spend by 38% while maintaining output quality through intelligent routing, caching, and budget enforcement.

When you're spending $47,000/month on Claude API calls across 12 teams and 23 applications, "just monitor the dashboard" doesn't cut it. We needed granular cost attribution, budget enforcement, intelligent model routing, and automated optimization — a full governance platform that prevents runaway costs while preserving the quality teams depend on.
We built this platform and reduced our monthly Claude spend by 38% ($47K to $29K) without any team reporting degraded output quality. Here's the architecture, the optimization strategies that worked, and the governance policies that keep costs predictable.
The Problem
Our Claude usage had grown organically across the organization:
- 12 teams using Claude through different applications
- 23 distinct applications making API calls
- No cost attribution — everyone shared one API key
- No quality baselines — nobody knew if cheaper models would suffice
- Unpredictable spikes — one bad prompt loop could cost $3,000 in an hour
- No caching — identical queries were processed (and billed) repeatedly
The $47K/month bill was growing 15% month-over-month with no visibility into what was driving costs.
Architecture
The governance platform sits as a proxy layer between applications and the Anthropic API.
Proxy Layer and Instrumentation
Every Claude API call routes through our governance proxy, which adds attribution, enforces policies, and collects metrics.
import Anthropic from '@anthropic-ai/sdk';
import { Redis } from 'ioredis';
interface RequestMetadata {
team: string;
application: string;
feature: string;
requestId: string;
userId: string;
priority: 'critical' | 'high' | 'normal' | 'low' | 'batch';
}
interface CostRecord {
timestamp: Date;
team: string;
application: string;
model: string;
inputTokens: number;
outputTokens: number;
cost: number;
latencyMs: number;
cached: boolean;
routedModel: string;
originalModel: string;
}
class GovernanceProxy {
private anthropic: Anthropic;
private redis: Redis;
private budgets: BudgetManager;
private router: ModelRouter;
private cache: SemanticCache;
constructor() {
this.anthropic = new Anthropic();
this.redis = new Redis(process.env.REDIS_URL!);
this.budgets = new BudgetManager(this.redis);
this.router = new ModelRouter();
this.cache = new SemanticCache(this.redis);
}
async createMessage(
params: Anthropic.MessageCreateParams,
metadata: RequestMetadata
): Promise<Anthropic.Message> {
// Step 1: Check budget
const budgetCheck = await this.budgets.checkBudget(metadata);
if (!budgetCheck.allowed) {
throw new BudgetExceededError(
`Team ${metadata.team} has exceeded its ${budgetCheck.period} budget. ` +
`Used: $${budgetCheck.used}, Limit: $${budgetCheck.limit}`
);
}
// Step 2: Check semantic cache
const cachedResponse = await this.cache.lookup(params);
if (cachedResponse) {
await this.recordUsage(metadata, params, cachedResponse, true);
return cachedResponse;
}
// Step 3: Intelligent model routing
const routedParams = await this.router.route(params, metadata);
// Step 4: Execute request
const startTime = Date.now();
const response = await this.anthropic.messages.create(routedParams);
const latencyMs = Date.now() - startTime;
// Step 5: Record usage and costs
await this.recordUsage(metadata, routedParams, response, false, latencyMs);
// Step 6: Cache response for future reuse
await this.cache.store(params, response);
return response;
}
private async recordUsage(
metadata: RequestMetadata,
params: Anthropic.MessageCreateParams,
response: Anthropic.Message,
cached: boolean,
latencyMs?: number
): Promise<void> {
const cost = this.calculateCost(
params.model,
response.usage.input_tokens,
response.usage.output_tokens
);
const record: CostRecord = {
timestamp: new Date(),
team: metadata.team,
application: metadata.application,
model: params.model,
inputTokens: response.usage.input_tokens,
outputTokens: response.usage.output_tokens,
cost,
latencyMs: latencyMs || 0,
cached,
routedModel: params.model,
originalModel: params.model
};
// Emit to metrics pipeline
await this.redis.xadd('cost:stream', '*', ...Object.entries(record).flat());
// Update budget tracking
await this.budgets.recordSpend(metadata.team, metadata.application, cost);
}
private calculateCost(model: string, inputTokens: number, outputTokens: number): number {
const pricing: Record<string, { input: number; output: number }> = {
'claude-sonnet-4-20250514': { input: 3.0 / 1_000_000, output: 15.0 / 1_000_000 },
'claude-haiku-4-20250514': { input: 0.25 / 1_000_000, output: 1.25 / 1_000_000 }
};
const rates = pricing[model] || pricing['claude-sonnet-4-20250514'];
return (inputTokens * rates.input) + (outputTokens * rates.output);
}
}
Intelligent Model Routing
Not every request needs the most expensive model. The router analyzes request characteristics to select the optimal model.
import anthropic
import hashlib
from dataclasses import dataclass
@dataclass
class RoutingDecision:
target_model: str
reason: str
estimated_savings: float
quality_risk: str # "none", "low", "medium"
class ModelRouter:
"""Routes requests to the optimal model based on task complexity."""
def __init__(self):
self.client = anthropic.Anthropic()
self.routing_rules = self._load_rules()
self.quality_baselines = self._load_baselines()
def route(self, params: dict, metadata: dict) -> dict:
"""Determine optimal model for this request."""
decision = self._make_routing_decision(params, metadata)
# Override model if routing suggests a different one
if decision.target_model != params.get("model"):
params = {**params, "model": decision.target_model}
return params
def _make_routing_decision(self, params: dict, metadata: dict) -> RoutingDecision:
# Rule 1: Batch/low-priority requests use Haiku
if metadata.get("priority") in ("low", "batch"):
return RoutingDecision(
target_model="claude-haiku-4-20250514",
reason="low priority request",
estimated_savings=0.85,
quality_risk="low"
)
# Rule 2: Simple classification/extraction tasks use Haiku
if self._is_simple_task(params):
return RoutingDecision(
target_model="claude-haiku-4-20250514",
reason="simple classification/extraction task",
estimated_savings=0.85,
quality_risk="none"
)
# Rule 3: Short context + structured output = Haiku candidate
input_tokens = self._estimate_tokens(params)
if input_tokens < 1000 and self._requests_structured_output(params):
return RoutingDecision(
target_model="claude-haiku-4-20250514",
reason="short context with structured output",
estimated_savings=0.85,
quality_risk="low"
)
# Default: use requested model
return RoutingDecision(
target_model=params.get("model", "claude-sonnet-4-20250514"),
reason="complex task requiring full capability",
estimated_savings=0.0,
quality_risk="none"
)
def _is_simple_task(self, params: dict) -> bool:
"""Detect tasks that don't need Sonnet."""
system = params.get("system", "")
messages = params.get("messages", [])
last_message = messages[-1]["content"] if messages else ""
simple_indicators = [
"classify", "categorize", "extract", "summarize in one sentence",
"return JSON", "yes or no", "true or false", "pick one"
]
content = f"{system} {last_message}".lower()
return any(indicator in content for indicator in simple_indicators)
Semantic Caching
Many applications send semantically identical requests — same intent, slightly different wording. We cache at the semantic level, not just exact match.
class SemanticCache {
private redis: Redis;
private ttlSeconds = 3600; // 1 hour default
constructor(redis: Redis) {
this.redis = redis;
}
async lookup(params: Anthropic.MessageCreateParams): Promise<Anthropic.Message | null> {
// Generate cache key from semantic content
const key = this.generateCacheKey(params);
const cached = await this.redis.get(`cache:${key}`);
if (cached) {
return JSON.parse(cached);
}
// Check for semantic near-matches
const nearKey = await this.findSemanticMatch(params);
if (nearKey) {
const nearCached = await this.redis.get(`cache:${nearKey}`);
if (nearCached) return JSON.parse(nearCached);
}
return null;
}
async store(params: Anthropic.MessageCreateParams, response: Anthropic.Message): Promise<void> {
const key = this.generateCacheKey(params);
// Only cache deterministic requests (low temperature)
if ((params.temperature || 1.0) > 0.3) return;
// Don't cache tool-use responses (they depend on external state)
if (params.tools && params.tools.length > 0) return;
await this.redis.setex(
`cache:${key}`,
this.ttlSeconds,
JSON.stringify(response)
);
}
private generateCacheKey(params: Anthropic.MessageCreateParams): string {
// Normalize and hash the semantic content
const normalized = {
model: params.model,
system: params.system || '',
messages: params.messages.map(m => ({
role: m.role,
content: typeof m.content === 'string' ? m.content.trim() : m.content
}))
};
return hashlib.createHash('sha256')
.update(JSON.stringify(normalized))
.digest('hex');
}
}
Budget Enforcement
Teams get monthly budgets with configurable alert thresholds and hard limits.
class BudgetManager:
def __init__(self, redis):
self.redis = redis
async def check_budget(self, metadata: dict) -> dict:
team = metadata["team"]
app = metadata["application"]
# Check team-level budget
team_budget = await self._get_budget(f"team:{team}")
team_spent = await self._get_spent(f"team:{team}")
if team_spent >= team_budget["hard_limit"]:
return {"allowed": False, "period": "monthly",
"used": team_spent, "limit": team_budget["hard_limit"]}
# Check application-level budget
app_budget = await self._get_budget(f"app:{team}:{app}")
app_spent = await self._get_spent(f"app:{team}:{app}")
if app_spent >= app_budget["hard_limit"]:
return {"allowed": False, "period": "monthly",
"used": app_spent, "limit": app_budget["hard_limit"]}
# Check hourly rate limit (prevents runaway loops)
hourly_spent = await self._get_hourly_spent(f"{team}:{app}")
hourly_limit = app_budget.get("hourly_limit", 500)
if hourly_spent >= hourly_limit:
await self._alert_team(team, "hourly_limit_hit", {
"application": app, "spent": hourly_spent, "limit": hourly_limit
})
return {"allowed": False, "period": "hourly",
"used": hourly_spent, "limit": hourly_limit}
# Send warning alerts at thresholds
utilization = team_spent / team_budget["hard_limit"]
if utilization > 0.8 and not await self._alert_sent(team, "80_pct"):
await self._alert_team(team, "budget_80_pct", {
"used": team_spent, "limit": team_budget["hard_limit"]
})
return {"allowed": True, "remaining": team_budget["hard_limit"] - team_spent}
Optimization Results
After implementing the full governance platform:
| Optimization | Monthly Savings | Quality Impact |
|---|---|---|
| Model routing (Sonnet → Haiku) | $11,200 (24%) | None measured |
| Semantic caching | $4,700 (10%) | None (identical responses) |
| Prompt optimization (shorter prompts) | $2,800 (6%) | None |
| Budget enforcement (prevented waste) | $1,400 (3%) | N/A |
| Batch API for non-urgent | $1,900 (4%) | Latency increase (acceptable) |
| Total | $18,000 (38%) | No degradation |
Cost Attribution Dashboard
Every team now sees their usage broken down by:
- Application and feature
- Model used (and whether it was auto-routed)
- Cache hit rate
- Cost per request average
- Month-over-month trend
- Budget utilization with forecast
Governance Policies
We implemented these organizational policies:
- Default model = Haiku — Teams must justify Sonnet usage for each use case
- Hourly rate limits — No application can spend more than $500/hour without override
- Mandatory caching — Applications with <20% cache hit rate must explain why
- Prompt review — New applications go through prompt efficiency review before launch
- Quarterly audits — Each team reviews their usage patterns with recommendations
Lessons Learned
Routing saves more than caching. Model routing saved 24% vs. caching's 10%. Most organizations over-index on caching and under-invest in model selection.
Budget alerts aren't enough — you need hard limits. Before hard limits, teams would acknowledge alerts and keep spending. Hard limits force optimization conversations.
Batch API is free money. Any request that doesn't need real-time response should use the Batch API. It's a 50% cost reduction for zero quality trade-off.
Measure quality before optimizing. We established quality baselines for each application before routing to cheaper models. Without baselines, you're optimizing blind.
Conclusion
AI cost governance at scale requires the same rigor as cloud cost management: attribution, budgets, optimization, and enforcement. The proxy architecture gives you full visibility and control without requiring application changes. Start with attribution (know who's spending what), then add routing (use the cheapest model that works), then add caching (don't pay twice for the same answer), and finally add enforcement (prevent waste). The 38% reduction was achievable because most organizations over-provision AI capabilities the same way they over-provision infrastructure — by defaulting to the most expensive option without measuring whether it's necessary.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.