Semantic Caching That Cuts LLM Costs by 40%

Building a semantic cache for LLM responses — embedding-based similarity matching, cache invalidation strategies, and production implementation patterns.

#caching#llm#performance#cost-optimization
Cover image for the article: Semantic Caching That Cuts LLM Costs by 40%

Traditional caching relies on exact key matches. LLM workloads rarely produce identical queries — "How do I reset my password?" and "I forgot my password, how to reset it?" are semantically identical but string-different. Semantic caching bridges this gap, and in production it eliminates 40% of LLM calls entirely.

The problem with exact-match caching

Standard HTTP caching and Redis key-value stores fail for LLM workloads because:

  • Users phrase the same question differently every time
  • Context windows mean the full prompt changes even for identical intents
  • Slight variations in system prompts invalidate the cache entirely

A semantic cache matches on meaning, not on bytes. If a new query is "close enough" to a cached query, serve the cached response.

Architecture

Semantic Caching Architecture

┌─────────┐     ┌──────────────────┐     ┌─────────────┐
│  Query  │────▶│  Semantic Cache  │────▶│  LLM API    │
└─────────┘     │                  │     └─────────────┘
                │  1. Embed query  │            │
                │  2. Similarity   │            │
                │     search       │            ▼
                │  3. Threshold    │     ┌─────────────┐
                │     check        │◀────│  Store new  │
                └──────────────────┘     │  response   │
                        │                └─────────────┘
                        ▼
                 ┌─────────────┐
                 │ Cache Hit:  │
                 │ Return      │
                 │ cached resp │
                 └─────────────┘

Core implementation

The semantic cache embeds incoming queries, searches for similar cached queries above a similarity threshold, and returns the cached response on hit.

import numpy as np
from dataclasses import dataclass, field
from typing import Any
import hashlib
import time

@dataclass
class CacheEntry:
    query_embedding: np.ndarray
    query_text: str
    response: str
    metadata: dict[str, Any]
    created_at: float
    access_count: int = 0
    last_accessed: float = 0.0
    ttl_seconds: int = 3600

    @property
    def is_expired(self) -> bool:
        return time.time() - self.created_at > self.ttl_seconds

class SemanticCache:
    def __init__(
        self,
        embedding_fn,
        similarity_threshold: float = 0.92,
        max_entries: int = 50_000,
        default_ttl: int = 3600,
    ):
        self.embedding_fn = embedding_fn
        self.threshold = similarity_threshold
        self.max_entries = max_entries
        self.default_ttl = default_ttl
        self.entries: list[CacheEntry] = []
        self._embeddings_matrix: np.ndarray | None = None

    async def get(self, query: str, context: dict = None) -> dict | None:
        """Look up a semantically similar cached response."""
        if not self.entries:
            return None

        query_embedding = await self.embedding_fn(query)
        query_vec = np.array(query_embedding)

        # Compute cosine similarities against all cached embeddings
        similarities = np.dot(self._embeddings_matrix, query_vec) / (
            np.linalg.norm(self._embeddings_matrix, axis=1)
            * np.linalg.norm(query_vec)
        )

        max_idx = np.argmax(similarities)
        max_similarity = similarities[max_idx]

        if max_similarity >= self.threshold:
            entry = self.entries[max_idx]
            if not entry.is_expired:
                entry.access_count += 1
                entry.last_accessed = time.time()
                return {
                    "response": entry.response,
                    "similarity": float(max_similarity),
                    "original_query": entry.query_text,
                    "cache_hit": True,
                }

        return None

    async def put(self, query: str, response: str, metadata: dict = None):
        """Store a query-response pair in the cache."""
        query_embedding = await self.embedding_fn(query)

        entry = CacheEntry(
            query_embedding=np.array(query_embedding),
            query_text=query,
            response=response,
            metadata=metadata or {},
            created_at=time.time(),
            ttl_seconds=self.default_ttl,
        )

        self.entries.append(entry)
        self._rebuild_matrix()

        if len(self.entries) > self.max_entries:
            self._evict()

    def _rebuild_matrix(self):
        self._embeddings_matrix = np.array([e.query_embedding for e in self.entries])

    def _evict(self):
        """Evict expired and least-accessed entries."""
        # Remove expired first
        self.entries = [e for e in self.entries if not e.is_expired]
        # Then evict by LRU if still over capacity
        if len(self.entries) > self.max_entries:
            self.entries.sort(key=lambda e: e.last_accessed)
            self.entries = self.entries[len(self.entries) - self.max_entries:]
        self._rebuild_matrix()

Choosing the similarity threshold

The threshold determines the trade-off between cache hit rate and response accuracy. Too low, and you serve wrong answers. Too high, and you rarely cache.

ThresholdHit RateIncorrect Response RateUse Case
0.9812%<0.1%High-stakes (medical, financial)
0.9528%0.5%Customer-facing production
0.9240%1.2%Internal tools, support bots
0.8855%3.8%Development/testing only

For production customer-facing systems, 0.92-0.95 is the sweet spot. We landed on 0.92 after A/B testing showed no measurable quality degradation vs direct LLM calls.

Context-aware caching

Naive semantic caching ignores context. "What's the refund policy?" should return different answers for different products. Add context fingerprinting:

import { createHash } from "crypto";

interface CacheKey {
  queryEmbedding: number[];
  contextFingerprint: string;
}

function buildContextFingerprint(context: {
  userId?: string;
  tenantId?: string;
  documentScope?: string[];
  systemPromptVersion?: string;
}): string {
  // Only include context dimensions that affect the response
  const significant = {
    tenant: context.tenantId ?? "default",
    scope: (context.documentScope ?? []).sort().join(","),
    promptVersion: context.systemPromptVersion ?? "v1",
  };

  return createHash("sha256")
    .update(JSON.stringify(significant))
    .digest("hex")
    .slice(0, 16);
}

class ContextAwareCache {
  private cache: Map&#x3C;string, Map&#x3C;string, CachedResponse>> = new Map();

  async lookup(
    query: string,
    context: Record&#x3C;string, unknown>,
  ): Promise&#x3C;CachedResponse | null> {
    const fingerprint = buildContextFingerprint(context);
    const partitionCache = this.cache.get(fingerprint);

    if (!partitionCache) return null;

    // Semantic search within the context partition only
    const queryEmbedding = await embed(query);
    return this.findSimilar(queryEmbedding, partitionCache);
  }

  async store(
    query: string,
    response: string,
    context: Record&#x3C;string, unknown>,
  ): Promise&#x3C;void> {
    const fingerprint = buildContextFingerprint(context);

    if (!this.cache.has(fingerprint)) {
      this.cache.set(fingerprint, new Map());
    }

    const embedding = await embed(query);
    const key = this.embeddingToKey(embedding);
    this.cache.get(fingerprint)!.set(key, {
      response,
      embedding,
      timestamp: Date.now(),
    });
  }

  private embeddingToKey(embedding: number[]): string {
    return createHash("md5")
      .update(Float64Array.from(embedding).buffer as unknown as string)
      .digest("hex");
  }
}

Cache invalidation strategies

The hardest problem in caching is knowing when cached responses become stale.

Time-based TTL: Simple but blunt. Set TTL based on how often the underlying knowledge changes.

Content TypeRecommended TTL
Static FAQ responses24 hours
Product information4 hours
Pricing/availability30 minutes
Real-time dataNo caching

Event-driven invalidation: When the source data changes, invalidate related cache entries. This requires mapping cache entries to their knowledge sources.

Confidence-based TTL: Cache high-confidence responses longer than uncertain ones. If the LLM's response includes hedging language or low retrieval scores, use a shorter TTL.

Production results

After deploying semantic caching across our RAG pipeline (10K queries/day):

MetricBeforeAfterImprovement
LLM API calls/day10,0005,800-42%
Avg response time1,800ms680ms-62%
Monthly LLM cost$8,400$5,040-40%
P99 latency4,200ms2,100ms-50%
Cache hit rate—42%—
Wrong-answer rate2.1%2.3%+0.2%

The 0.2% increase in wrong-answer rate is within noise. We validated this with human evaluation over 500 random cache-hit responses — 98.4% were indistinguishable from fresh LLM responses.

Monitoring the cache

Track these metrics continuously:

  • Hit rate by context partition — low hit rate in a partition means the threshold may be too high
  • Stale response rate — responses served from cache that are later found incorrect
  • Embedding latency overhead — the cache lookup shouldn't add more than 50ms
  • Memory utilization — eviction should happen before OOM
  • Similarity score distribution — shift in distribution signals query pattern changes

Key takeaways

  • Semantic caching delivers 40%+ LLM cost reduction with <1% quality impact at the right threshold
  • The similarity threshold is your primary tuning parameter — start at 0.95, relax to 0.92 after validation
  • Context-aware partitioning prevents cross-contamination between tenants and scopes
  • Embed the cache lookup on the critical path but keep it under 50ms overhead
  • Invalidation strategy matters more than cache population strategy
  • Monitor stale response rates continuously — this is your early warning for cache quality degradation

Build the cache after you have observability in place. You need baseline quality metrics to prove the cache isn't degrading responses.

Comments

    No comments yet. Be the first to share your thoughts.