Reducing LLM Inference Costs by 73% with Smart Batching

Practical strategies for cutting LLM inference costs through intelligent request batching, model routing, and prompt optimization without sacrificing quality.

#llm#cost-optimization#inference#ai
Cover image for the article: Reducing LLM Inference Costs by 73% with Smart Batching

LLM inference costs scale linearly with traffic unless you architect against it. After optimizing inference pipelines serving 2M+ daily requests, I reduced our monthly spend from $47K to $12.7K — a 73% reduction — without degrading response quality. Here's the playbook.

The cost problem nobody talks about

Most teams focus on model selection when optimizing costs. That matters, but it's only one lever. The real savings come from three areas most engineers overlook:

  1. Request batching — grouping similar requests to maximize throughput per dollar
  2. Semantic routing — sending simple queries to smaller, cheaper models
  3. Prompt compression — reducing token count without losing intent

The compounding effect of all three is what gets you to 70%+ savings.

Architecture overview

LLM Cost Optimization Architecture

The system sits between your application layer and the LLM providers. It intercepts requests, classifies them by complexity, batches compatible ones, and routes to the optimal model.

┌─────────────┐     ┌──────────────┐     ┌─────────────────┐
│ Application │────▶│  AI Gateway  │────▶│  Model Router   │
└─────────────┘     └──────────────┘     └─────────────────┘
                           │                       │
                    ┌──────┴──────┐         ┌──────┴──────┐
                    │ Batch Queue │         │ Cost Tracker │
                    └─────────────┘         └─────────────┘

Strategy 1: Intelligent request batching

Not all requests need immediate responses. Background tasks like summarization, classification, and data extraction can tolerate 50-200ms of additional latency for batching.

import asyncio
from dataclasses import dataclass, field
from typing import Any
import time

@dataclass
class BatchedInferenceEngine:
    max_batch_size: int = 32
    max_wait_ms: float = 100.0
    _queue: list = field(default_factory=list)
    _lock: asyncio.Lock = field(default_factory=asyncio.Lock)

    async def infer(self, prompt: str, priority: str = "normal") -> str:
        if priority == "realtime":
            return await self._single_infer(prompt)

        future = asyncio.get_event_loop().create_future()
        async with self._lock:
            self._queue.append({"prompt": prompt, "future": future, "ts": time.time()})
            if len(self._queue) >= self.max_batch_size:
                await self._flush_batch()

        return await future

    async def _flush_batch(self):
        batch = self._queue[:self.max_batch_size]
        self._queue = self._queue[self.max_batch_size:]

        prompts = [item["prompt"] for item in batch]
        results = await self._batch_api_call(prompts)

        for item, result in zip(batch, results):
            item["future"].set_result(result)

    async def _batch_api_call(self, prompts: list[str]) -> list[str]:
        # Batch API call reduces per-request overhead by ~60%
        # Using provider batch endpoints (e.g., Anthropic Batch API)
        response = await anthropic_client.messages.batch_create(
            requests=[
                {"custom_id": f"req_{i}", "params": {"model": "claude-sonnet-4-20250514", "messages": [{"role": "user", "content": p}]}}
                for i, p in enumerate(prompts)
            ]
        )
        return [r.result.content[0].text for r in response.results]

Batching alone saved us 31% — the batch API pricing discount combined with reduced overhead per request compounds quickly.

Strategy 2: Semantic complexity routing

Not every request needs your most powerful model. A classification query doesn't need the same model as a complex reasoning task.

interface ModelRoute {
  model: string;
  costPer1kTokens: number;
  maxComplexity: number;
}

const ROUTES: ModelRoute[] = [
  { model: "claude-haiku", costPer1kTokens: 0.00025, maxComplexity: 3 },
  { model: "claude-sonnet", costPer1kTokens: 0.003, maxComplexity: 7 },
  { model: "claude-opus", costPer1kTokens: 0.015, maxComplexity: 10 },
];

async function routeRequest(prompt: string): Promise<ModelRoute> {
  // Use a lightweight classifier to estimate complexity (1-10)
  const complexity = await classifyComplexity(prompt);

  // Route to cheapest model that can handle the complexity
  const route = ROUTES.find((r) => r.maxComplexity >= complexity);
  return route ?? ROUTES[ROUTES.length - 1];
}

async function classifyComplexity(prompt: string): Promise<number> {
  // Lightweight heuristics first, ML classifier for edge cases
  const tokenCount = prompt.split(/\s+/).length;
  const hasCodeBlock = /```/.test(prompt);
  const hasMultiStep = /(first|then|finally|step)/i.test(prompt);

  let score = 3; // baseline
  if (tokenCount > 500) score += 2;
  if (hasCodeBlock) score += 2;
  if (hasMultiStep) score += 2;

  return Math.min(score, 10);
}

In production, 62% of our requests route to the cheapest model tier. The quality difference is imperceptible for simple tasks — classification accuracy drops by only 0.3% on Haiku vs Opus for binary decisions.

Strategy 3: Prompt compression

Every token costs money. Most prompts carry redundant context, verbose instructions, and unnecessary formatting.

Our prompt compression pipeline:

TechniqueToken ReductionQuality Impact
Remove examples from few-shot (use 2 instead of 5)40-60%-1.2% accuracy
Compress system prompts with abbreviations15-25%None measurable
Cache and reference previous context30-50%None measurable
Use structured input format (JSON vs prose)20-35%+0.8% accuracy

Benchmarks: Before and after

After implementing all three strategies across our production pipeline:

MetricBeforeAfterChange
Monthly cost$47,200$12,700-73%
Avg latency (p50)1,200ms890ms-26%
Avg latency (p99)4,800ms3,100ms-35%
Quality score (human eval)4.2/54.1/5-2.4%
Daily request volume2.1M2.1M—

The latency improvement is a bonus — batching introduces small delays for background tasks but eliminates queue contention for real-time requests.

The cost tracking dashboard

You can't optimize what you don't measure. We tag every request with:

  • Model used
  • Token count (input + output)
  • Route decision reason
  • Batch vs single
  • Estimated cost

This feeds a real-time dashboard that alerts when cost-per-request drifts above thresholds by endpoint.

Implementation order

If you're starting from scratch, implement in this order:

  1. Prompt compression (week 1) — immediate savings, zero infrastructure changes
  2. Semantic routing (weeks 2-3) — requires a complexity classifier, moderate effort
  3. Request batching (weeks 3-4) — requires queue infrastructure, highest savings ceiling

Each strategy works independently. You don't need all three to see meaningful savings.

What I'd do differently

Looking back, I'd start with better observability. We spent two weeks optimizing a route that handled 3% of traffic because we assumed it was expensive. Measure first, then optimize the expensive paths.

I'd also invest earlier in prompt versioning. When you're A/B testing prompt compression, you need to track which version of which prompt generated which result. That audit trail pays for itself when debugging quality regressions.

Key takeaways

  • Batch APIs are underutilized — if your use case tolerates 100ms delay, you're leaving money on the table
  • Most requests don't need your most expensive model — a lightweight classifier saves 40%+ on model costs alone
  • Prompt engineering isn't just about quality — it's about cost per token
  • Measure cost per request, not just total monthly spend

The 73% reduction isn't a one-time win. As traffic grows, these optimizations compound. Our cost now scales sub-linearly with request volume, which changes the unit economics of every AI feature we ship.

Comments

    No comments yet. Be the first to share your thoughts.