Reducing LLM Inference Costs by 73% with Smart Batching
Practical strategies for cutting LLM inference costs through intelligent request batching, model routing, and prompt optimization without sacrificing quality.

LLM inference costs scale linearly with traffic unless you architect against it. After optimizing inference pipelines serving 2M+ daily requests, I reduced our monthly spend from $47K to $12.7K — a 73% reduction — without degrading response quality. Here's the playbook.
The cost problem nobody talks about
Most teams focus on model selection when optimizing costs. That matters, but it's only one lever. The real savings come from three areas most engineers overlook:
- Request batching — grouping similar requests to maximize throughput per dollar
- Semantic routing — sending simple queries to smaller, cheaper models
- Prompt compression — reducing token count without losing intent
The compounding effect of all three is what gets you to 70%+ savings.
Architecture overview
The system sits between your application layer and the LLM providers. It intercepts requests, classifies them by complexity, batches compatible ones, and routes to the optimal model.
┌─────────────┐ ┌──────────────┐ ┌─────────────────┐
│ Application │────▶│ AI Gateway │────▶│ Model Router │
└─────────────┘ └──────────────┘ └─────────────────┘
│ │
┌──────┴──────┐ ┌──────┴──────┐
│ Batch Queue │ │ Cost Tracker │
└─────────────┘ └─────────────┘
Strategy 1: Intelligent request batching
Not all requests need immediate responses. Background tasks like summarization, classification, and data extraction can tolerate 50-200ms of additional latency for batching.
import asyncio
from dataclasses import dataclass, field
from typing import Any
import time
@dataclass
class BatchedInferenceEngine:
max_batch_size: int = 32
max_wait_ms: float = 100.0
_queue: list = field(default_factory=list)
_lock: asyncio.Lock = field(default_factory=asyncio.Lock)
async def infer(self, prompt: str, priority: str = "normal") -> str:
if priority == "realtime":
return await self._single_infer(prompt)
future = asyncio.get_event_loop().create_future()
async with self._lock:
self._queue.append({"prompt": prompt, "future": future, "ts": time.time()})
if len(self._queue) >= self.max_batch_size:
await self._flush_batch()
return await future
async def _flush_batch(self):
batch = self._queue[:self.max_batch_size]
self._queue = self._queue[self.max_batch_size:]
prompts = [item["prompt"] for item in batch]
results = await self._batch_api_call(prompts)
for item, result in zip(batch, results):
item["future"].set_result(result)
async def _batch_api_call(self, prompts: list[str]) -> list[str]:
# Batch API call reduces per-request overhead by ~60%
# Using provider batch endpoints (e.g., Anthropic Batch API)
response = await anthropic_client.messages.batch_create(
requests=[
{"custom_id": f"req_{i}", "params": {"model": "claude-sonnet-4-20250514", "messages": [{"role": "user", "content": p}]}}
for i, p in enumerate(prompts)
]
)
return [r.result.content[0].text for r in response.results]
Batching alone saved us 31% — the batch API pricing discount combined with reduced overhead per request compounds quickly.
Strategy 2: Semantic complexity routing
Not every request needs your most powerful model. A classification query doesn't need the same model as a complex reasoning task.
interface ModelRoute {
model: string;
costPer1kTokens: number;
maxComplexity: number;
}
const ROUTES: ModelRoute[] = [
{ model: "claude-haiku", costPer1kTokens: 0.00025, maxComplexity: 3 },
{ model: "claude-sonnet", costPer1kTokens: 0.003, maxComplexity: 7 },
{ model: "claude-opus", costPer1kTokens: 0.015, maxComplexity: 10 },
];
async function routeRequest(prompt: string): Promise<ModelRoute> {
// Use a lightweight classifier to estimate complexity (1-10)
const complexity = await classifyComplexity(prompt);
// Route to cheapest model that can handle the complexity
const route = ROUTES.find((r) => r.maxComplexity >= complexity);
return route ?? ROUTES[ROUTES.length - 1];
}
async function classifyComplexity(prompt: string): Promise<number> {
// Lightweight heuristics first, ML classifier for edge cases
const tokenCount = prompt.split(/\s+/).length;
const hasCodeBlock = /```/.test(prompt);
const hasMultiStep = /(first|then|finally|step)/i.test(prompt);
let score = 3; // baseline
if (tokenCount > 500) score += 2;
if (hasCodeBlock) score += 2;
if (hasMultiStep) score += 2;
return Math.min(score, 10);
}
In production, 62% of our requests route to the cheapest model tier. The quality difference is imperceptible for simple tasks — classification accuracy drops by only 0.3% on Haiku vs Opus for binary decisions.
Strategy 3: Prompt compression
Every token costs money. Most prompts carry redundant context, verbose instructions, and unnecessary formatting.
Our prompt compression pipeline:
| Technique | Token Reduction | Quality Impact |
|---|---|---|
| Remove examples from few-shot (use 2 instead of 5) | 40-60% | -1.2% accuracy |
| Compress system prompts with abbreviations | 15-25% | None measurable |
| Cache and reference previous context | 30-50% | None measurable |
| Use structured input format (JSON vs prose) | 20-35% | +0.8% accuracy |
Benchmarks: Before and after
After implementing all three strategies across our production pipeline:
| Metric | Before | After | Change |
|---|---|---|---|
| Monthly cost | $47,200 | $12,700 | -73% |
| Avg latency (p50) | 1,200ms | 890ms | -26% |
| Avg latency (p99) | 4,800ms | 3,100ms | -35% |
| Quality score (human eval) | 4.2/5 | 4.1/5 | -2.4% |
| Daily request volume | 2.1M | 2.1M | — |
The latency improvement is a bonus — batching introduces small delays for background tasks but eliminates queue contention for real-time requests.
The cost tracking dashboard
You can't optimize what you don't measure. We tag every request with:
- Model used
- Token count (input + output)
- Route decision reason
- Batch vs single
- Estimated cost
This feeds a real-time dashboard that alerts when cost-per-request drifts above thresholds by endpoint.
Implementation order
If you're starting from scratch, implement in this order:
- Prompt compression (week 1) — immediate savings, zero infrastructure changes
- Semantic routing (weeks 2-3) — requires a complexity classifier, moderate effort
- Request batching (weeks 3-4) — requires queue infrastructure, highest savings ceiling
Each strategy works independently. You don't need all three to see meaningful savings.
What I'd do differently
Looking back, I'd start with better observability. We spent two weeks optimizing a route that handled 3% of traffic because we assumed it was expensive. Measure first, then optimize the expensive paths.
I'd also invest earlier in prompt versioning. When you're A/B testing prompt compression, you need to track which version of which prompt generated which result. That audit trail pays for itself when debugging quality regressions.
Key takeaways
- Batch APIs are underutilized — if your use case tolerates 100ms delay, you're leaving money on the table
- Most requests don't need your most expensive model — a lightweight classifier saves 40%+ on model costs alone
- Prompt engineering isn't just about quality — it's about cost per token
- Measure cost per request, not just total monthly spend
The 73% reduction isn't a one-time win. As traffic grows, these optimizations compound. Our cost now scales sub-linearly with request volume, which changes the unit economics of every AI feature we ship.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.