Building AI-Powered Search Autocomplete at Scale
Architecture and implementation of intelligent search autocomplete using embedding models, behavioral signals, and real-time ranking for sub-50ms suggestions

Search autocomplete is the most latency-sensitive AI feature in most products. Users expect suggestions to appear within 50ms of each keystroke. At the same time, those suggestions need to be semantically relevant, personalized, and contextually appropriate. This is a deceptively hard problem that combines real-time systems engineering with AI ranking.
This article covers the architecture for a production autocomplete system serving 10M+ queries per day with sub-50ms P95 latency.
System Requirements
| Requirement | Target | Rationale |
|---|---|---|
| Latency (P50) | < 20ms | User-perceived instantaneous |
| Latency (P95) | < 50ms | Maintain responsiveness |
| Relevance (NDCG@5) | > 0.75 | High-quality suggestions |
| Freshness | < 5 min | New content appears quickly |
| Personalization | Per-user | Adapts to individual behavior |
| Availability | 99.99% | Search is a core path |
Architecture Overview
The system uses a three-stage pipeline: prefix retrieval, semantic expansion, and personalized ranking.
| Stage | Purpose | Latency Budget | Technology |
|---|---|---|---|
| Prefix Trie | Fast prefix matching | < 5ms | Redis / Custom trie |
| Semantic Expansion | Related completions | < 15ms | Embedding similarity |
| Personalized Ranking | User-specific ordering | < 10ms | Lightweight model |
| Business Rules | Filters, boosts, diversity | < 5ms | Rule engine |
Stage 1: Prefix Retrieval
The trie index handles the initial prefix matching with sub-5ms latency:
from typing import List, Tuple
import redis
import json
class PrefixIndex:
def __init__(self, redis_client: redis.Redis):
self.redis = redis_client
def index_suggestion(self, text: str, score: float, metadata: dict = None):
"""Add a suggestion to the prefix index."""
normalized = text.lower().strip()
# Index all prefixes from 2 characters up
for i in range(2, len(normalized) + 1):
prefix = normalized[:i]
self.redis.zadd(
f"autocomplete:{prefix}",
{json.dumps({"text": text, "meta": metadata}): score},
)
# Trim to top 50 per prefix for memory efficiency
self.redis.zremrangebyrank(f"autocomplete:{prefix}", 0, -51)
def query(self, prefix: str, limit: int = 10) -> List[dict]:
"""Retrieve top suggestions for a prefix."""
normalized = prefix.lower().strip()
if len(normalized) < 2:
return []
results = self.redis.zrevrange(
f"autocomplete:{normalized}", 0, limit - 1, withscores=True
)
suggestions = []
for item, score in results:
data = json.loads(item)
data["score"] = score
suggestions.append(data)
return suggestions
def update_score(self, text: str, boost: float):
"""Boost suggestion score based on user interactions."""
normalized = text.lower().strip()
for i in range(2, len(normalized) + 1):
prefix = normalized[:i]
self.redis.zincrby(
f"autocomplete:{prefix}",
boost,
json.dumps({"text": text}),
)
Stage 2: Semantic Expansion
Expand beyond exact prefix matches using embedding similarity:
import numpy as np
from sentence_transformers import SentenceTransformer
import faiss
class SemanticExpander:
def __init__(self, model_name: str = "BAAI/bge-small-en-v1.5"):
self.model = SentenceTransformer(model_name)
self.index = None
self.suggestions = []
def build_index(self, suggestions: List[str]):
"""Build FAISS index over all suggestion embeddings."""
self.suggestions = suggestions
embeddings = self.model.encode(
suggestions, batch_size=256, show_progress_bar=False,
normalize_embeddings=True,
)
# Use IndexFlatIP for cosine similarity (normalized vectors)
self.index = faiss.IndexFlatIP(embeddings.shape[1])
self.index.add(embeddings.astype(np.float32))
def expand(self, partial_query: str, prefix_results: List[dict],
k: int = 10) -> List[dict]:
"""Find semantically related suggestions."""
if not self.index or len(partial_query) < 3:
return []
# Encode the partial query
query_embedding = self.model.encode(
[partial_query], normalize_embeddings=True
).astype(np.float32)
# Search for similar suggestions
scores, indices = self.index.search(query_embedding, k * 2)
# Filter out prefix results (already included)
prefix_texts = {r["text"].lower() for r in prefix_results}
expanded = []
for idx, score in zip(indices[0], scores[0]):
if idx < 0:
continue
suggestion = self.suggestions[idx]
if suggestion.lower() not in prefix_texts:
expanded.append({
"text": suggestion,
"score": float(score),
"source": "semantic",
})
if len(expanded) >= k:
break
return expanded
Stage 3: Personalized Ranking
Re-rank combined results using user history and context:
import torch
import torch.nn as nn
class PersonalizedRanker(nn.Module):
"""Lightweight ranking model for personalized autocomplete."""
def __init__(self, suggestion_dim: int = 384, user_dim: int = 64,
context_dim: int = 32):
super().__init__()
total_dim = suggestion_dim + user_dim + context_dim
self.scorer = nn.Sequential(
nn.Linear(total_dim, 128),
nn.ReLU(),
nn.Dropout(0.1),
nn.Linear(128, 64),
nn.ReLU(),
nn.Linear(64, 1),
)
def forward(self, suggestion_emb, user_emb, context_emb):
combined = torch.cat([suggestion_emb, user_emb, context_emb], dim=-1)
return self.scorer(combined).squeeze(-1)
class RankingService:
def __init__(self, model: PersonalizedRanker, feature_store):
self.model = model
self.features = feature_store
self.model.eval()
@torch.inference_mode()
def rank(self, candidates: List[dict], user_id: str,
context: dict) -> List[dict]:
"""Rank candidates using personalization model."""
user_features = self._get_user_features(user_id)
context_features = self._encode_context(context)
suggestion_embeddings = torch.stack([
torch.tensor(c.get("embedding", [0] * 384))
for c in candidates
])
user_emb = user_features.unsqueeze(0).expand(len(candidates), -1)
ctx_emb = context_features.unsqueeze(0).expand(len(candidates), -1)
scores = self.model(suggestion_embeddings, user_emb, ctx_emb)
# Combine model scores with base scores
for i, candidate in enumerate(candidates):
candidate["final_score"] = (
0.6 * scores[i].item() + 0.4 * candidate.get("score", 0)
)
return sorted(candidates, key=lambda x: x["final_score"], reverse=True)
def _get_user_features(self, user_id: str) -> torch.Tensor:
"""Get precomputed user embedding from feature store."""
features = self.features.get(f"user:{user_id}:search_profile")
if features:
return torch.tensor(features, dtype=torch.float32)
return torch.zeros(64) # Cold start default
Behavioral Signal Collection
Track user interactions to improve suggestion quality over time:
| Signal | Weight | Decay | Update Frequency |
|---|---|---|---|
| Click on suggestion | +1.0 | 7-day half-life | Real-time |
| Search after suggestion shown | +0.3 | 3-day half-life | Real-time |
| Suggestion ignored | -0.1 | 1-day half-life | Batch (hourly) |
| Conversion after suggestion | +3.0 | 14-day half-life | Near real-time |
| Session search count | Context | N/A | Per session |
Performance Benchmarks
Production metrics from a system serving 10M queries/day:
| Metric | Value | Target |
|---|---|---|
| P50 latency | 14ms | < 20ms |
| P95 latency | 38ms | < 50ms |
| P99 latency | 62ms | < 100ms |
| NDCG@5 | 0.78 | > 0.75 |
| Suggestion click-through rate | 34.2% | > 30% |
| Zero-result rate | 2.1% | < 5% |
| Queries per second (peak) | 8,400 | Scale to 15K |
| Infrastructure cost | $4,200/month | Budget: $5K |
Caching Strategy
Caching is critical for meeting latency requirements:
class AutocompleteCache:
"""Multi-level caching for autocomplete responses."""
def __init__(self, redis_client, local_cache_size: int = 10000):
self.redis = redis_client
self.local_cache = {} # LRU cache for hottest prefixes
self.local_cache_size = local_cache_size
def get(self, prefix: str, user_id: str = None) -> list:
"""Check caches in order: local -> Redis -> None."""
# Level 1: Local in-memory (< 0.1ms)
cache_key = f"{prefix}:{user_id or 'global'}"
if cache_key in self.local_cache:
return self.local_cache[cache_key]
# Level 2: Redis (< 2ms)
cached = self.redis.get(f"ac_cache:{cache_key}")
if cached:
result = json.loads(cached)
self._update_local(cache_key, result)
return result
return None
def set(self, prefix: str, results: list, user_id: str = None,
ttl: int = 300):
"""Cache results at both levels."""
cache_key = f"{prefix}:{user_id or 'global'}"
self.redis.setex(f"ac_cache:{cache_key}", ttl, json.dumps(results))
self._update_local(cache_key, results)
Key Takeaways
- Sub-50ms latency requires a multi-stage architecture. No single approach achieves both relevance and speed. Use fast prefix retrieval, then enrich with semantics.
- Semantic expansion catches 20-30% more relevant suggestions that pure prefix matching misses (typo-tolerance, synonyms, related concepts).
- Personalization improves CTR by 15-25%. Even a simple user-history-based boost significantly outperforms one-size-fits-all ranking.
- Cache aggressively at multiple levels. The top 1000 prefixes cover 60-70% of queries. Keep them in memory for sub-1ms responses.
- Behavioral signals compound over time. Click-through data makes the system smarter with every interaction. Invest in signal collection infrastructure early.
Search autocomplete is where AI meets real-time systems engineering. The best systems feel magical to users precisely because the complexity is invisible - they just get relevant suggestions instantly, every time.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.