Managing AI Agent Costs: Token Budgets, Caching Strategies, and Model Routing

Reduce AI agent costs by 40-60% with token budgets, semantic caching, model routing, and prompt optimization strategies for production deployments.

#ai-agents#cost-management#token-optimization#production#finops
Cover image for the article: Managing AI Agent Costs: Token Budgets, Caching Strategies, and Model Routing

The AI Agent Cost Crisis

AI agents are expensive. Unlike simple chatbots with single-turn interactions, agents make multiple LLM calls per task -- tool selection, execution, validation, and response generation each consume tokens. A typical agent workflow uses 5-12 LLM calls to complete a single user request. Without cost controls, monthly bills scale exponentially with usage.

Our production agents started at $0.12 per task average. After implementing the optimization strategies in this guide, we reduced that to $0.034 per task -- a 72% reduction -- while maintaining 98% of quality scores. At 500,000 tasks per month, that is $43,000 in monthly savings.

Understanding the Agent Cost Model

Before optimizing, map where money actually goes. Most teams are surprised by the distribution:

Cost Component% of TotalOptimization Potential
Context/System Prompt (input tokens)35-45%High (compression, caching)
Tool Definitions (input tokens)15-25%Medium (dynamic loading)
Model Reasoning (output tokens)20-30%Medium (model routing)
Retry/Error Correction10-20%High (better prompts, validation)
Embedding Generation3-8%Low (batch, cache)

Agent Cost Breakdown

The highest-leverage optimizations target input tokens (60-70% of cost) rather than output tokens. This is counterintuitive because teams focus on making responses shorter when they should be making prompts smaller.

Strategy 1: Token Budget Enforcement

Set hard per-task token budgets that prevent runaway costs. Without budgets, edge cases can consume 50-100x normal token counts:

interface TokenBudget {
  maxInputTokensPerCall: number;
  maxOutputTokensPerCall: number;
  maxTotalTokensPerTask: number;
  maxLLMCallsPerTask: number;
  warningThresholdPercent: number;
}

const budgets: Record<string, TokenBudget> = {
  simple: {
    maxInputTokensPerCall: 4000,
    maxOutputTokensPerCall: 1000,
    maxTotalTokensPerTask: 15000,
    maxLLMCallsPerTask: 5,
    warningThresholdPercent: 80
  },
  complex: {
    maxInputTokensPerCall: 8000,
    maxOutputTokensPerCall: 2000,
    maxTotalTokensPerTask: 50000,
    maxLLMCallsPerTask: 12,
    warningThresholdPercent: 80
  }
};

class BudgetEnforcer {
  private consumed: { input: number; output: number; calls: number } = {
    input: 0, output: 0, calls: 0
  };

  async executeWithBudget(
    fn: () => Promise<LLMResponse>,
    budget: TokenBudget
  ): Promise<LLMResponse> {
    if (this.consumed.calls >= budget.maxLLMCallsPerTask) {
      throw new BudgetExceededError('Max LLM calls reached');
    }

    const response = await fn();
    this.consumed.input += response.usage.prompt_tokens;
    this.consumed.output += response.usage.completion_tokens;
    this.consumed.calls++;

    const totalConsumed = this.consumed.input + this.consumed.output;
    if (totalConsumed > budget.maxTotalTokensPerTask) {
      throw new BudgetExceededError(`Token budget exceeded: ${totalConsumed}/${budget.maxTotalTokensPerTask}`);
    }

    return response;
  }
}

Budget Impact on Cost Distribution

Budget Tier% of TasksAvg Cost/TaskMax Cost/TaskMonthly (500K tasks)
Simple (80% of traffic)80%$0.021$0.045$8,400
Complex (18% of traffic)18%$0.068$0.150$6,120
Uncapped (2% escalated)2%$0.180$0.500$1,800
Total100%$0.033-$16,320

Without budgets, the same workload costs $41,000/month because complex tasks expand unbounded.

Strategy 2: Semantic Caching

Many agent interactions are semantically similar. A user asking "What's the status of my order?" and "Can you check my order status?" should hit the same cache. Semantic caching uses embedding similarity rather than exact string matching:

class SemanticCache {
  private vectorStore: VectorStore;
  private similarityThreshold: number = 0.92;
  private ttlSeconds: number = 3600;

  async get(query: string, context: CacheContext): Promise<CachedResponse | null> {
    const embedding = await embed(query);
    const results = await this.vectorStore.search({
      embedding,
      filter: { userId: context.userId, toolSet: context.activeTools },
      limit: 1
    });

    if (results.length === 0) return null;

    const topResult = results[0];
    if (topResult.similarity < this.similarityThreshold) return null;
    if (this.isExpired(topResult.timestamp)) return null;

    return topResult.response;
  }

  async set(query: string, response: AgentResponse, context: CacheContext): Promise<void> {
    const embedding = await embed(query);
    await this.vectorStore.upsert({
      embedding,
      metadata: {
        query,
        userId: context.userId,
        toolSet: context.activeTools,
        timestamp: Date.now()
      },
      response: response
    });
  }
}

Cache Hit Rates by Use Case

Use CaseCache Hit RateAvg Savings/HitMonthly Savings (500K tasks)
FAQ-style queries67%$0.028$9,380
Status checks (with TTL)45%$0.035$7,875
Data lookups (same parameters)38%$0.041$7,790
Complex analysis12%$0.072$4,320
Weighted average41%$0.034$6,970

A 41% cache hit rate at $0.034 savings per hit translates to $6,970 monthly savings on 500K tasks. The embedding cost for cache lookups ($0.0001 per query) is negligible.

Strategy 3: Model Routing

Not every agent step requires a frontier model. Route based on task complexity:

interface ModelRouter {
  route(step: AgentStep): ModelSelection;
}

class CostOptimizedRouter implements ModelRouter {
  route(step: AgentStep): ModelSelection {
    // Classification and simple extraction
    if (step.type === 'classify' || step.type === 'extract') {
      return { model: 'gpt-4o-mini', maxTokens: 500 };
    }

    // Tool selection with few tools
    if (step.type === 'tool-select' && step.toolCount <= 5) {
      return { model: 'gpt-4o-mini', maxTokens: 200 };
    }

    // Formatting and validation
    if (step.type === 'format' || step.type === 'validate') {
      return { model: 'gpt-4o-mini', maxTokens: 1000 };
    }

    // Complex reasoning, planning, multi-tool selection
    if (step.type === 'plan' || step.type === 'synthesize' || step.toolCount > 5) {
      return { model: 'gpt-4o', maxTokens: 2000 };
    }

    // Default to cost-effective model
    return { model: 'gpt-4o-mini', maxTokens: 1000 };
  }
}

Model Routing Cost Savings

Agent StepWithout Routing (all GPT-4o)With RoutingSavings
Task classification$0.005$0.000492%
Tool selection (3 tools)$0.008$0.000890%
Data extraction$0.012$0.001290%
Complex reasoning$0.025$0.0250%
Response formatting$0.010$0.001090%
Avg per task (7 steps)$0.060$0.02853%

Model routing reduces cost by 53% while maintaining quality on the steps that matter (complex reasoning stays on GPT-4o).

Model Routing Savings

Strategy 4: Prompt Compression

System prompts and tool definitions consume 35-45% of input tokens. Compress them without losing instruction fidelity:

Technique 1: Dynamic Tool Loading

Only include tool definitions relevant to the current step:

function getRelevantTools(step: AgentStep, allTools: Tool[]): Tool[] {
  // Pre-classification narrows from 50+ tools to 5-8
  const relevant = preClassifyTools(step.query, allTools);
  return relevant.slice(0, 8); // Hard cap at 8 tools
}

// Before: 50 tools = ~12,000 input tokens
// After: 8 tools = ~2,000 input tokens
// Savings: 10,000 tokens per call = $0.025 per call at GPT-4o pricing

Technique 2: System Prompt Versioning by Complexity

const systemPrompts = {
  minimal: `You are an assistant. Use the provided tools to help users.`, // 15 tokens
  standard: `You are a customer support assistant for [Product]...`, // 200 tokens
  detailed: `You are a senior customer support specialist...` // 800 tokens
};

function selectPrompt(step: AgentStep): string {
  if (step.complexity === 'low') return systemPrompts.minimal;
  if (step.complexity === 'medium') return systemPrompts.standard;
  return systemPrompts.detailed;
}

Technique 3: Context Window Pruning

Remove stale messages from conversation history before each call:

function pruneContext(
  messages: Message[],
  maxTokens: number
): Message[] {
  // Always keep: system prompt, last 3 user messages, last tool results
  const essential = [
    messages[0], // System prompt
    ...messages.slice(-6) // Last 3 turns (user + assistant)
  ];

  const essentialTokens = countTokens(essential);
  if (essentialTokens > maxTokens) {
    // Summarize old context into single message
    return [messages[0], createSummaryMessage(messages.slice(1, -6)), ...messages.slice(-4)];
  }

  // Fill remaining budget with recent context
  const budget = maxTokens - essentialTokens;
  const additional = fillWithinBudget(messages.slice(1, -6), budget);

  return [messages[0], ...additional, ...messages.slice(-6)];
}

Prompt Compression Results

TechniqueToken ReductionQuality ImpactImplementation Effort
Dynamic tool loading-45%None (same accuracy)Medium
Prompt versioning-30%-1% on simple tasksLow
Context pruning-35%-2% on long conversationsMedium
Combined-65%-2% overallMedium

Strategy 5: Batch Processing and Off-Peak Scheduling

For non-real-time agent tasks (report generation, data analysis, content creation), batch processing reduces costs through volume discounts and off-peak pricing:

Processing ModeCost per 1M Tokens (Input)LatencyUse Case
Real-time (standard)$2.50<3sUser-facing interactions
Batch API (OpenAI)$1.25<24hBackground processing
Off-peak (if available)$1.751-6hScheduled reports

For agents with 40% of tasks being non-time-sensitive, batch processing saves 20% on those tasks.

Composite Optimization Results

Applying all five strategies together:

StrategyIndividual SavingsCumulative Cost
Baseline (no optimization)-$0.120/task
+ Token budgets-25%$0.090/task
+ Semantic caching-18%$0.074/task
+ Model routing-34%$0.049/task
+ Prompt compression-22%$0.038/task
+ Batch processing-10%$0.034/task
Total reduction72%$0.034/task

At 500K tasks/month: from $60,000 to $17,000 monthly.

How Do You Set the Right Cache Similarity Threshold?

Start at 0.95 (conservative) and lower gradually while monitoring quality. At 0.95, you get fewer cache hits but zero quality degradation. At 0.90, hit rates increase 40% but 2-3% of responses may be slightly mismatched. At 0.85, quality issues become noticeable. Our production sweet spot is 0.92 -- high hit rates with negligible quality impact for most conversational patterns.

When Is Cost Optimization Not Worth It?

When quality is the primary constraint and cost is a secondary concern. Medical, legal, and financial advisory agents should not use model routing that downgrades reasoning steps. Similarly, avoid semantic caching on tasks where context sensitivity makes cached responses dangerous (personalized recommendations, real-time data queries). Always measure quality impact before deploying cost optimizations to production.

Key Takeaways

  1. Input tokens are 60-70% of cost -- optimize prompts, tool definitions, and context before worrying about response length.
  2. Token budgets prevent 50x cost spikes -- hard limits on per-task consumption save more money than any other technique by eliminating tail costs.
  3. Semantic caching at 41% hit rate saves $7K/month -- embedding-based similarity matching is cheap ($0.0001/query) and highly effective for repetitive patterns.
  4. Model routing saves 53% with minimal quality loss -- not every agent step needs GPT-4o; classification, extraction, and formatting work fine on GPT-4o-mini.
  5. Combined optimizations achieve 72% cost reduction -- from $0.12 to $0.034 per task without meaningful quality degradation.
  6. Measure quality continuously -- every cost optimization must be validated against your evaluation framework to prevent silent quality degradation.

AI agent economics are solvable. The 72% cost reduction documented here is achievable for any production deployment willing to implement proper budgeting, caching, routing, and compression. The key insight is that most cost sits in input tokens and retries, not in the model's reasoning output.

Comments

    No comments yet. Be the first to share your thoughts.