Monitoring and Governing Claude API Costs at Scale

How we built a cost governance platform that reduced Claude API spend by 38% while maintaining output quality through intelligent routing, caching, and budget enforcement.

#claude#cost-optimization#governance#ai
Cover image for the article: Monitoring and Governing Claude API Costs at Scale

When you're spending $47,000/month on Claude API calls across 12 teams and 23 applications, "just monitor the dashboard" doesn't cut it. We needed granular cost attribution, budget enforcement, intelligent model routing, and automated optimization — a full governance platform that prevents runaway costs while preserving the quality teams depend on.

We built this platform and reduced our monthly Claude spend by 38% ($47K to $29K) without any team reporting degraded output quality. Here's the architecture, the optimization strategies that worked, and the governance policies that keep costs predictable.

The Problem

Our Claude usage had grown organically across the organization:

  • 12 teams using Claude through different applications
  • 23 distinct applications making API calls
  • No cost attribution — everyone shared one API key
  • No quality baselines — nobody knew if cheaper models would suffice
  • Unpredictable spikes — one bad prompt loop could cost $3,000 in an hour
  • No caching — identical queries were processed (and billed) repeatedly

The $47K/month bill was growing 15% month-over-month with no visibility into what was driving costs.

Architecture

The governance platform sits as a proxy layer between applications and the Anthropic API.

Cost Governance Architecture

Proxy Layer and Instrumentation

Every Claude API call routes through our governance proxy, which adds attribution, enforces policies, and collects metrics.

import Anthropic from '@anthropic-ai/sdk';
import { Redis } from 'ioredis';

interface RequestMetadata {
  team: string;
  application: string;
  feature: string;
  requestId: string;
  userId: string;
  priority: 'critical' | 'high' | 'normal' | 'low' | 'batch';
}

interface CostRecord {
  timestamp: Date;
  team: string;
  application: string;
  model: string;
  inputTokens: number;
  outputTokens: number;
  cost: number;
  latencyMs: number;
  cached: boolean;
  routedModel: string;
  originalModel: string;
}

class GovernanceProxy {
  private anthropic: Anthropic;
  private redis: Redis;
  private budgets: BudgetManager;
  private router: ModelRouter;
  private cache: SemanticCache;

  constructor() {
    this.anthropic = new Anthropic();
    this.redis = new Redis(process.env.REDIS_URL!);
    this.budgets = new BudgetManager(this.redis);
    this.router = new ModelRouter();
    this.cache = new SemanticCache(this.redis);
  }

  async createMessage(
    params: Anthropic.MessageCreateParams,
    metadata: RequestMetadata
  ): Promise<Anthropic.Message> {
    // Step 1: Check budget
    const budgetCheck = await this.budgets.checkBudget(metadata);
    if (!budgetCheck.allowed) {
      throw new BudgetExceededError(
        `Team ${metadata.team} has exceeded its ${budgetCheck.period} budget. ` +
        `Used: $${budgetCheck.used}, Limit: $${budgetCheck.limit}`
      );
    }

    // Step 2: Check semantic cache
    const cachedResponse = await this.cache.lookup(params);
    if (cachedResponse) {
      await this.recordUsage(metadata, params, cachedResponse, true);
      return cachedResponse;
    }

    // Step 3: Intelligent model routing
    const routedParams = await this.router.route(params, metadata);

    // Step 4: Execute request
    const startTime = Date.now();
    const response = await this.anthropic.messages.create(routedParams);
    const latencyMs = Date.now() - startTime;

    // Step 5: Record usage and costs
    await this.recordUsage(metadata, routedParams, response, false, latencyMs);

    // Step 6: Cache response for future reuse
    await this.cache.store(params, response);

    return response;
  }

  private async recordUsage(
    metadata: RequestMetadata,
    params: Anthropic.MessageCreateParams,
    response: Anthropic.Message,
    cached: boolean,
    latencyMs?: number
  ): Promise<void> {
    const cost = this.calculateCost(
      params.model,
      response.usage.input_tokens,
      response.usage.output_tokens
    );

    const record: CostRecord = {
      timestamp: new Date(),
      team: metadata.team,
      application: metadata.application,
      model: params.model,
      inputTokens: response.usage.input_tokens,
      outputTokens: response.usage.output_tokens,
      cost,
      latencyMs: latencyMs || 0,
      cached,
      routedModel: params.model,
      originalModel: params.model
    };

    // Emit to metrics pipeline
    await this.redis.xadd('cost:stream', '*', ...Object.entries(record).flat());
    
    // Update budget tracking
    await this.budgets.recordSpend(metadata.team, metadata.application, cost);
  }

  private calculateCost(model: string, inputTokens: number, outputTokens: number): number {
    const pricing: Record<string, { input: number; output: number }> = {
      'claude-sonnet-4-20250514': { input: 3.0 / 1_000_000, output: 15.0 / 1_000_000 },
      'claude-haiku-4-20250514': { input: 0.25 / 1_000_000, output: 1.25 / 1_000_000 }
    };
    const rates = pricing[model] || pricing['claude-sonnet-4-20250514'];
    return (inputTokens * rates.input) + (outputTokens * rates.output);
  }
}

Intelligent Model Routing

Not every request needs the most expensive model. The router analyzes request characteristics to select the optimal model.

import anthropic
import hashlib
from dataclasses import dataclass

@dataclass
class RoutingDecision:
    target_model: str
    reason: str
    estimated_savings: float
    quality_risk: str  # "none", "low", "medium"

class ModelRouter:
    """Routes requests to the optimal model based on task complexity."""

    def __init__(self):
        self.client = anthropic.Anthropic()
        self.routing_rules = self._load_rules()
        self.quality_baselines = self._load_baselines()

    def route(self, params: dict, metadata: dict) -> dict:
        """Determine optimal model for this request."""
        
        decision = self._make_routing_decision(params, metadata)
        
        # Override model if routing suggests a different one
        if decision.target_model != params.get("model"):
            params = {**params, "model": decision.target_model}

        return params

    def _make_routing_decision(self, params: dict, metadata: dict) -> RoutingDecision:
        # Rule 1: Batch/low-priority requests use Haiku
        if metadata.get("priority") in ("low", "batch"):
            return RoutingDecision(
                target_model="claude-haiku-4-20250514",
                reason="low priority request",
                estimated_savings=0.85,
                quality_risk="low"
            )

        # Rule 2: Simple classification/extraction tasks use Haiku
        if self._is_simple_task(params):
            return RoutingDecision(
                target_model="claude-haiku-4-20250514",
                reason="simple classification/extraction task",
                estimated_savings=0.85,
                quality_risk="none"
            )

        # Rule 3: Short context + structured output = Haiku candidate
        input_tokens = self._estimate_tokens(params)
        if input_tokens < 1000 and self._requests_structured_output(params):
            return RoutingDecision(
                target_model="claude-haiku-4-20250514",
                reason="short context with structured output",
                estimated_savings=0.85,
                quality_risk="low"
            )

        # Default: use requested model
        return RoutingDecision(
            target_model=params.get("model", "claude-sonnet-4-20250514"),
            reason="complex task requiring full capability",
            estimated_savings=0.0,
            quality_risk="none"
        )

    def _is_simple_task(self, params: dict) -> bool:
        """Detect tasks that don't need Sonnet."""
        system = params.get("system", "")
        messages = params.get("messages", [])
        last_message = messages[-1]["content"] if messages else ""

        simple_indicators = [
            "classify", "categorize", "extract", "summarize in one sentence",
            "return JSON", "yes or no", "true or false", "pick one"
        ]
        
        content = f"{system} {last_message}".lower()
        return any(indicator in content for indicator in simple_indicators)

Semantic Caching

Many applications send semantically identical requests — same intent, slightly different wording. We cache at the semantic level, not just exact match.

class SemanticCache {
  private redis: Redis;
  private ttlSeconds = 3600; // 1 hour default

  constructor(redis: Redis) {
    this.redis = redis;
  }

  async lookup(params: Anthropic.MessageCreateParams): Promise<Anthropic.Message | null> {
    // Generate cache key from semantic content
    const key = this.generateCacheKey(params);

    const cached = await this.redis.get(`cache:${key}`);
    if (cached) {
      return JSON.parse(cached);
    }

    // Check for semantic near-matches
    const nearKey = await this.findSemanticMatch(params);
    if (nearKey) {
      const nearCached = await this.redis.get(`cache:${nearKey}`);
      if (nearCached) return JSON.parse(nearCached);
    }

    return null;
  }

  async store(params: Anthropic.MessageCreateParams, response: Anthropic.Message): Promise<void> {
    const key = this.generateCacheKey(params);

    // Only cache deterministic requests (low temperature)
    if ((params.temperature || 1.0) > 0.3) return;

    // Don't cache tool-use responses (they depend on external state)
    if (params.tools && params.tools.length > 0) return;

    await this.redis.setex(
      `cache:${key}`,
      this.ttlSeconds,
      JSON.stringify(response)
    );
  }

  private generateCacheKey(params: Anthropic.MessageCreateParams): string {
    // Normalize and hash the semantic content
    const normalized = {
      model: params.model,
      system: params.system || '',
      messages: params.messages.map(m => ({
        role: m.role,
        content: typeof m.content === 'string' ? m.content.trim() : m.content
      }))
    };

    return hashlib.createHash('sha256')
      .update(JSON.stringify(normalized))
      .digest('hex');
  }
}

Budget Enforcement

Teams get monthly budgets with configurable alert thresholds and hard limits.

class BudgetManager:
    def __init__(self, redis):
        self.redis = redis

    async def check_budget(self, metadata: dict) -> dict:
        team = metadata["team"]
        app = metadata["application"]

        # Check team-level budget
        team_budget = await self._get_budget(f"team:{team}")
        team_spent = await self._get_spent(f"team:{team}")

        if team_spent >= team_budget["hard_limit"]:
            return {"allowed": False, "period": "monthly", 
                    "used": team_spent, "limit": team_budget["hard_limit"]}

        # Check application-level budget
        app_budget = await self._get_budget(f"app:{team}:{app}")
        app_spent = await self._get_spent(f"app:{team}:{app}")

        if app_spent >= app_budget["hard_limit"]:
            return {"allowed": False, "period": "monthly",
                    "used": app_spent, "limit": app_budget["hard_limit"]}

        # Check hourly rate limit (prevents runaway loops)
        hourly_spent = await self._get_hourly_spent(f"{team}:{app}")
        hourly_limit = app_budget.get("hourly_limit", 500)

        if hourly_spent >= hourly_limit:
            await self._alert_team(team, "hourly_limit_hit", {
                "application": app, "spent": hourly_spent, "limit": hourly_limit
            })
            return {"allowed": False, "period": "hourly",
                    "used": hourly_spent, "limit": hourly_limit}

        # Send warning alerts at thresholds
        utilization = team_spent / team_budget["hard_limit"]
        if utilization > 0.8 and not await self._alert_sent(team, "80_pct"):
            await self._alert_team(team, "budget_80_pct", {
                "used": team_spent, "limit": team_budget["hard_limit"]
            })

        return {"allowed": True, "remaining": team_budget["hard_limit"] - team_spent}

Optimization Results

After implementing the full governance platform:

OptimizationMonthly SavingsQuality Impact
Model routing (Sonnet → Haiku)$11,200 (24%)None measured
Semantic caching$4,700 (10%)None (identical responses)
Prompt optimization (shorter prompts)$2,800 (6%)None
Budget enforcement (prevented waste)$1,400 (3%)N/A
Batch API for non-urgent$1,900 (4%)Latency increase (acceptable)
Total$18,000 (38%)No degradation

Cost Attribution Dashboard

Every team now sees their usage broken down by:

  • Application and feature
  • Model used (and whether it was auto-routed)
  • Cache hit rate
  • Cost per request average
  • Month-over-month trend
  • Budget utilization with forecast

Governance Policies

We implemented these organizational policies:

  1. Default model = Haiku — Teams must justify Sonnet usage for each use case
  2. Hourly rate limits — No application can spend more than $500/hour without override
  3. Mandatory caching — Applications with <20% cache hit rate must explain why
  4. Prompt review — New applications go through prompt efficiency review before launch
  5. Quarterly audits — Each team reviews their usage patterns with recommendations

Lessons Learned

Routing saves more than caching. Model routing saved 24% vs. caching's 10%. Most organizations over-index on caching and under-invest in model selection.

Budget alerts aren't enough — you need hard limits. Before hard limits, teams would acknowledge alerts and keep spending. Hard limits force optimization conversations.

Batch API is free money. Any request that doesn't need real-time response should use the Batch API. It's a 50% cost reduction for zero quality trade-off.

Measure quality before optimizing. We established quality baselines for each application before routing to cheaper models. Without baselines, you're optimizing blind.

Conclusion

AI cost governance at scale requires the same rigor as cloud cost management: attribution, budgets, optimization, and enforcement. The proxy architecture gives you full visibility and control without requiring application changes. Start with attribution (know who's spending what), then add routing (use the cheapest model that works), then add caching (don't pay twice for the same answer), and finally add enforcement (prevent waste). The 38% reduction was achievable because most organizations over-provision AI capabilities the same way they over-provision infrastructure — by defaulting to the most expensive option without measuring whether it's necessary.

Comments

    No comments yet. Be the first to share your thoughts.