Reducing LLM Response Latency from 3.2s to 800ms in Production

Practical techniques for cutting LLM response latency by 75% through streaming, caching, prompt optimization, and model routing without sacrificing output quality.

#llm#latency#optimization#production
Cover image for the article: Reducing LLM Response Latency from 3.2s to 800ms in Production

A 3.2-second response time kills user engagement. Users start typing their next message, switch tabs, or abandon the interaction entirely. When we measured, every 500ms of added latency reduced user satisfaction scores by 12% and task completion rates by 8%. Reducing our LLM pipeline from 3.2s to 800ms time-to-first-token required optimizing every layer of the stack.

The Problem: Where 3.2 Seconds Goes

Breaking down our original latency:

  • Network to API provider: 120ms
  • Prompt construction + template rendering: 45ms
  • Context retrieval (RAG): 380ms
  • Queue wait time: 280ms
  • Model inference (time to first token): 1,850ms
  • Token generation (full response): 520ms additional
  • Post-processing + safety checks: 85ms
  • Total: ~3,280ms

Each component needs different optimization strategies. The biggest wins came from model routing, semantic caching, and prompt compression — not from faster hardware.

Architecture: Latency-Optimized LLM Pipeline

The optimized pipeline runs retrieval and prompt preparation in parallel, routes to the fastest capable model, and streams responses through safety checks without buffering.

LLM Latency Optimization Architecture

Semantic Cache with Similarity Matching

The highest-impact optimization: avoid calling the LLM at all for semantically similar queries. Our cache hit rate reaches 35% for customer support and FAQ-style queries.

import hashlib
import time
import numpy as np
from dataclasses import dataclass
from typing import Optional
import asyncio

@dataclass
class CacheEntry:
    query_embedding: np.ndarray
    response: str
    created_at: float
    hit_count: int
    model_used: str
    token_count: int

@dataclass
class CacheResult:
    hit: bool
    response: Optional[str] = None
    similarity: float = 0.0
    latency_ms: float = 0.0
    entry_age_seconds: float = 0.0

class SemanticCache:
    def __init__(self, embedding_model, similarity_threshold: float = 0.95,
                 max_entries: int = 100000, ttl_seconds: int = 3600):
        self.embedding_model = embedding_model
        self.threshold = similarity_threshold
        self.max_entries = max_entries
        self.ttl = ttl_seconds
        self.entries: dict[str, CacheEntry] = {}
        self._embeddings_matrix: Optional[np.ndarray] = None
        self._entry_keys: list[str] = []

    async def get(self, query: str, context_hash: Optional[str] = None) -> CacheResult:
        start = time.time()

        if not self.entries:
            return CacheResult(hit=False, latency_ms=(time.time() - start) * 1000)

        query_embedding = await self.embedding_model.encode(query)
        
        # Find most similar cached query
        if self._embeddings_matrix is None:
            self._rebuild_matrix()

        similarities = np.dot(self._embeddings_matrix, query_embedding)
        best_idx = np.argmax(similarities)
        best_similarity = float(similarities[best_idx])

        if best_similarity >= self.threshold:
            key = self._entry_keys[best_idx]
            entry = self.entries[key]
            
            # Check TTL
            age = time.time() - entry.created_at
            if age > self.ttl:
                self._evict(key)
                return CacheResult(hit=False, latency_ms=(time.time() - start) * 1000)

            entry.hit_count += 1
            return CacheResult(
                hit=True,
                response=entry.response,
                similarity=best_similarity,
                latency_ms=(time.time() - start) * 1000,
                entry_age_seconds=age,
            )

        return CacheResult(hit=False, similarity=best_similarity,
                          latency_ms=(time.time() - start) * 1000)

    async def put(self, query: str, response: str, model_used: str, token_count: int):
        query_embedding = await self.embedding_model.encode(query)
        key = hashlib.md5(query.encode()).hexdigest()

        self.entries[key] = CacheEntry(
            query_embedding=query_embedding,
            response=response,
            created_at=time.time(),
            hit_count=0,
            model_used=model_used,
            token_count=token_count,
        )

        self._embeddings_matrix = None  # Invalidate matrix cache
        self._entry_keys.append(key)

        if len(self.entries) > self.max_entries:
            self._evict_lru()

    def _rebuild_matrix(self):
        self._entry_keys = list(self.entries.keys())
        embeddings = [self.entries[k].query_embedding for k in self._entry_keys]
        self._embeddings_matrix = np.array(embeddings)

    def _evict(self, key: str):
        del self.entries[key]
        self._embeddings_matrix = None

    def _evict_lru(self):
        oldest_key = min(self.entries, key=lambda k: self.entries[k].created_at)
        self._evict(oldest_key)

    @property
    def stats(self) -> dict:
        total_hits = sum(e.hit_count for e in self.entries.values())
        return {
            "entries": len(self.entries),
            "total_hits": total_hits,
            "avg_entry_age": np.mean([time.time() - e.created_at for e in self.entries.values()]) if self.entries else 0,
        }

Intelligent Model Router

Route queries to the fastest model that can handle the complexity. Simple queries go to smaller, faster models; complex queries go to larger models.

interface ModelConfig {
  modelId: string;
  provider: string;
  maxTokens: number;
  avgLatencyMs: number;
  costPer1kTokens: number;
  capabilities: string[];
  qualityScore: number;  // 0-1, from eval benchmarks
}

interface RoutingDecision {
  modelId: string;
  reason: string;
  estimatedLatencyMs: number;
  confidence: number;
}

interface QueryClassification {
  complexity: 'simple' | 'moderate' | 'complex';
  requiresReasoning: boolean;
  requiresCreativity: boolean;
  requiresFactualAccuracy: boolean;
  estimatedOutputTokens: number;
  domain: string;
}

class LLMRouter {
  private models: ModelConfig[];
  private classifier: { classify: (query: string) => Promise<QueryClassification> };
  private latencyTracker: Map<string, number[]>;

  constructor(models: ModelConfig[], classifier: { classify: (query: string) => Promise<QueryClassification> }) {
    this.models = models.sort((a, b) => a.avgLatencyMs - b.avgLatencyMs);
    this.classifier = classifier;
    this.latencyTracker = new Map();
  }

  async route(query: string, maxLatencyMs: number = 2000): Promise<RoutingDecision> {
    const classification = await this.classifier.classify(query);

    // Filter models by capability requirements
    const capable = this.models.filter(m => this.meetsRequirements(m, classification));

    // Filter by latency budget
    const withinBudget = capable.filter(m => {
      const recentLatency = this.getRecentP95(m.modelId);
      return recentLatency <= maxLatencyMs;
    });

    if (withinBudget.length === 0) {
      // Fallback: fastest capable model regardless of budget
      const fastest = capable[0];
      return {
        modelId: fastest.modelId,
        reason: `No model within ${maxLatencyMs}ms budget, using fastest capable: ${fastest.modelId}`,
        estimatedLatencyMs: this.getRecentP95(fastest.modelId),
        confidence: 0.6,
      };
    }

    // Select optimal model: minimize latency while meeting quality threshold
    const qualityThreshold = classification.complexity === 'complex' ? 0.8 : 0.6;
    const qualityFiltered = withinBudget.filter(m => m.qualityScore >= qualityThreshold);
    const selected = qualityFiltered.length > 0
      ? qualityFiltered[0]  // Fastest that meets quality bar
      : withinBudget[0];     // Fastest within budget

    return {
      modelId: selected.modelId,
      reason: `${classification.complexity} query, routed to ${selected.modelId} (quality: ${selected.qualityScore}, est: ${selected.avgLatencyMs}ms)`,
      estimatedLatencyMs: this.getRecentP95(selected.modelId),
      confidence: 0.85,
    };
  }

  private meetsRequirements(model: ModelConfig, classification: QueryClassification): boolean {
    if (classification.requiresReasoning && !model.capabilities.includes('reasoning')) {
      return false;
    }
    if (classification.estimatedOutputTokens > model.maxTokens) {
      return false;
    }
    if (classification.complexity === 'complex' && model.qualityScore < 0.7) {
      return false;
    }
    return true;
  }

  private getRecentP95(modelId: string): number {
    const latencies = this.latencyTracker.get(modelId);
    if (!latencies || latencies.length === 0) {
      const model = this.models.find(m => m.modelId === modelId);
      return model?.avgLatencyMs ?? 2000;
    }
    const sorted = [...latencies].sort((a, b) => a - b);
    return sorted[Math.floor(sorted.length * 0.95)];
  }

  recordLatency(modelId: string, latencyMs: number): void {
    const existing = this.latencyTracker.get(modelId) || [];
    existing.push(latencyMs);
    if (existing.length > 1000) existing.shift();
    this.latencyTracker.set(modelId, existing);
  }
}

Additional Optimization Techniques

Prompt Compression

Long prompts increase time-to-first-token linearly. Compress prompts by:

  • Removing redundant few-shot examples based on query similarity
  • Using shorter system prompts with abbreviations models understand
  • Compressing retrieved context through extractive summarization
  • Eliminating boilerplate that doesn't affect output quality

We reduced average prompt length from 3,200 tokens to 1,400 tokens with no measurable quality loss.

Speculative Decoding

For self-hosted models, speculative decoding uses a small draft model to predict multiple tokens ahead, then validates with the large model in a single forward pass. This achieves 2-3x speedup for generation-heavy responses.

Streaming with Progressive Enhancement

Stream the first response chunk immediately while running expensive post-processing (citation verification, fact-checking) in parallel. Users perceive the response as faster even if total completion time is similar.

Connection Pooling and Keep-Alive

Maintain persistent HTTP/2 connections to LLM providers. Cold connection setup adds 50-150ms per request. Connection pooling eliminates this for all but the first request.

Benchmarks: Latency Reduction Results

OptimizationLatency ReductionQuality Impact
Semantic cache (35% hit rate)-65% (on hits)None (cached responses)
Model routing (simple→small model)-55%-2% on simple queries
Prompt compression-28%-0.5% (negligible)
Parallel context retrieval-15%None
Streaming TTFT-40% perceivedNone
Connection pooling-8%None
Combined pipeline-75% (3.2s → 800ms TTFT)-1.2% average

Before vs After (P95 Metrics)

MetricBeforeAfter
Time to first token3,200ms800ms
Full response time4,800ms2,100ms
Throughput (req/sec)45180
Cost per request (avg)$0.0082$0.0034
User satisfaction score3.6/54.4/5
Task completion rate72%88%

Implementation Priority

For teams starting LLM latency optimization:

  1. Streaming (1 day) — Immediate perceived improvement, zero quality cost
  2. Connection pooling (hours) — Free latency reduction
  3. Semantic caching (1 week) — Highest ROI for repetitive query patterns
  4. Model routing (2 weeks) — Requires query classification training
  5. Prompt compression (1 week) — Requires quality validation per use case
  6. Speculative decoding (2 weeks) — Only for self-hosted models

Lessons Learned

Time-to-first-token matters more than total time. Users are patient with streaming responses as long as they start quickly. Optimize TTFT aggressively even at the cost of total generation time.

Cache invalidation is the hard part. Semantic caches can serve stale responses when the underlying data changes. Tie cache invalidation to data update events, not just TTL.

Small models are better than you think. For 60% of our queries, a 7B parameter model produces responses indistinguishable from a 70B model. Proper routing saves both latency and cost.

Measure end-to-end, not just model time. Our biggest initial win came from fixing a queue backup issue, not optimizing inference. Profile the entire request lifecycle.

Conclusion

LLM latency optimization is a systems problem, not just a model problem. The 75% reduction from 3.2s to 800ms came from attacking every layer: caching eliminates redundant inference, routing sends queries to the fastest capable model, prompt compression reduces input processing time, and streaming makes the remaining latency feel shorter. Start with streaming and caching for immediate wins, then layer on routing and compression as you build the infrastructure.

Comments

    No comments yet. Be the first to share your thoughts.