Building AI-Powered Search Autocomplete at Scale

Architecture and implementation of intelligent search autocomplete using embedding models, behavioral signals, and real-time ranking for sub-50ms suggestions

#search#autocomplete#embeddings#real-time
Cover image for the article: Building AI-Powered Search Autocomplete at Scale

Search autocomplete is the most latency-sensitive AI feature in most products. Users expect suggestions to appear within 50ms of each keystroke. At the same time, those suggestions need to be semantically relevant, personalized, and contextually appropriate. This is a deceptively hard problem that combines real-time systems engineering with AI ranking.

This article covers the architecture for a production autocomplete system serving 10M+ queries per day with sub-50ms P95 latency.

System Requirements

RequirementTargetRationale
Latency (P50)< 20msUser-perceived instantaneous
Latency (P95)< 50msMaintain responsiveness
Relevance (NDCG@5)> 0.75High-quality suggestions
Freshness< 5 minNew content appears quickly
PersonalizationPer-userAdapts to individual behavior
Availability99.99%Search is a core path

Chart

Architecture Overview

The system uses a three-stage pipeline: prefix retrieval, semantic expansion, and personalized ranking.

StagePurposeLatency BudgetTechnology
Prefix TrieFast prefix matching< 5msRedis / Custom trie
Semantic ExpansionRelated completions< 15msEmbedding similarity
Personalized RankingUser-specific ordering< 10msLightweight model
Business RulesFilters, boosts, diversity< 5msRule engine

Stage 1: Prefix Retrieval

The trie index handles the initial prefix matching with sub-5ms latency:

from typing import List, Tuple
import redis
import json

class PrefixIndex:
    def __init__(self, redis_client: redis.Redis):
        self.redis = redis_client

    def index_suggestion(self, text: str, score: float, metadata: dict = None):
        """Add a suggestion to the prefix index."""
        normalized = text.lower().strip()
        # Index all prefixes from 2 characters up
        for i in range(2, len(normalized) + 1):
            prefix = normalized[:i]
            self.redis.zadd(
                f"autocomplete:{prefix}",
                {json.dumps({"text": text, "meta": metadata}): score},
            )
            # Trim to top 50 per prefix for memory efficiency
            self.redis.zremrangebyrank(f"autocomplete:{prefix}", 0, -51)

    def query(self, prefix: str, limit: int = 10) -> List[dict]:
        """Retrieve top suggestions for a prefix."""
        normalized = prefix.lower().strip()
        if len(normalized) &#x3C; 2:
            return []

        results = self.redis.zrevrange(
            f"autocomplete:{normalized}", 0, limit - 1, withscores=True
        )

        suggestions = []
        for item, score in results:
            data = json.loads(item)
            data["score"] = score
            suggestions.append(data)

        return suggestions

    def update_score(self, text: str, boost: float):
        """Boost suggestion score based on user interactions."""
        normalized = text.lower().strip()
        for i in range(2, len(normalized) + 1):
            prefix = normalized[:i]
            self.redis.zincrby(
                f"autocomplete:{prefix}",
                boost,
                json.dumps({"text": text}),
            )

Stage 2: Semantic Expansion

Expand beyond exact prefix matches using embedding similarity:

import numpy as np
from sentence_transformers import SentenceTransformer
import faiss

class SemanticExpander:
    def __init__(self, model_name: str = "BAAI/bge-small-en-v1.5"):
        self.model = SentenceTransformer(model_name)
        self.index = None
        self.suggestions = []

    def build_index(self, suggestions: List[str]):
        """Build FAISS index over all suggestion embeddings."""
        self.suggestions = suggestions
        embeddings = self.model.encode(
            suggestions, batch_size=256, show_progress_bar=False,
            normalize_embeddings=True,
        )
        # Use IndexFlatIP for cosine similarity (normalized vectors)
        self.index = faiss.IndexFlatIP(embeddings.shape[1])
        self.index.add(embeddings.astype(np.float32))

    def expand(self, partial_query: str, prefix_results: List[dict],
               k: int = 10) -> List[dict]:
        """Find semantically related suggestions."""
        if not self.index or len(partial_query) &#x3C; 3:
            return []

        # Encode the partial query
        query_embedding = self.model.encode(
            [partial_query], normalize_embeddings=True
        ).astype(np.float32)

        # Search for similar suggestions
        scores, indices = self.index.search(query_embedding, k * 2)

        # Filter out prefix results (already included)
        prefix_texts = {r["text"].lower() for r in prefix_results}
        expanded = []

        for idx, score in zip(indices[0], scores[0]):
            if idx &#x3C; 0:
                continue
            suggestion = self.suggestions[idx]
            if suggestion.lower() not in prefix_texts:
                expanded.append({
                    "text": suggestion,
                    "score": float(score),
                    "source": "semantic",
                })

            if len(expanded) >= k:
                break

        return expanded

Stage 3: Personalized Ranking

Re-rank combined results using user history and context:

import torch
import torch.nn as nn

class PersonalizedRanker(nn.Module):
    """Lightweight ranking model for personalized autocomplete."""

    def __init__(self, suggestion_dim: int = 384, user_dim: int = 64,
                 context_dim: int = 32):
        super().__init__()
        total_dim = suggestion_dim + user_dim + context_dim

        self.scorer = nn.Sequential(
            nn.Linear(total_dim, 128),
            nn.ReLU(),
            nn.Dropout(0.1),
            nn.Linear(128, 64),
            nn.ReLU(),
            nn.Linear(64, 1),
        )

    def forward(self, suggestion_emb, user_emb, context_emb):
        combined = torch.cat([suggestion_emb, user_emb, context_emb], dim=-1)
        return self.scorer(combined).squeeze(-1)


class RankingService:
    def __init__(self, model: PersonalizedRanker, feature_store):
        self.model = model
        self.features = feature_store
        self.model.eval()

    @torch.inference_mode()
    def rank(self, candidates: List[dict], user_id: str,
             context: dict) -> List[dict]:
        """Rank candidates using personalization model."""
        user_features = self._get_user_features(user_id)
        context_features = self._encode_context(context)

        suggestion_embeddings = torch.stack([
            torch.tensor(c.get("embedding", [0] * 384))
            for c in candidates
        ])

        user_emb = user_features.unsqueeze(0).expand(len(candidates), -1)
        ctx_emb = context_features.unsqueeze(0).expand(len(candidates), -1)

        scores = self.model(suggestion_embeddings, user_emb, ctx_emb)

        # Combine model scores with base scores
        for i, candidate in enumerate(candidates):
            candidate["final_score"] = (
                0.6 * scores[i].item() + 0.4 * candidate.get("score", 0)
            )

        return sorted(candidates, key=lambda x: x["final_score"], reverse=True)

    def _get_user_features(self, user_id: str) -> torch.Tensor:
        """Get precomputed user embedding from feature store."""
        features = self.features.get(f"user:{user_id}:search_profile")
        if features:
            return torch.tensor(features, dtype=torch.float32)
        return torch.zeros(64)  # Cold start default

Behavioral Signal Collection

Track user interactions to improve suggestion quality over time:

SignalWeightDecayUpdate Frequency
Click on suggestion+1.07-day half-lifeReal-time
Search after suggestion shown+0.33-day half-lifeReal-time
Suggestion ignored-0.11-day half-lifeBatch (hourly)
Conversion after suggestion+3.014-day half-lifeNear real-time
Session search countContextN/APer session

Performance Benchmarks

Production metrics from a system serving 10M queries/day:

MetricValueTarget
P50 latency14ms< 20ms
P95 latency38ms< 50ms
P99 latency62ms< 100ms
NDCG@50.78> 0.75
Suggestion click-through rate34.2%> 30%
Zero-result rate2.1%< 5%
Queries per second (peak)8,400Scale to 15K
Infrastructure cost$4,200/monthBudget: $5K

Caching Strategy

Caching is critical for meeting latency requirements:

class AutocompleteCache:
    """Multi-level caching for autocomplete responses."""

    def __init__(self, redis_client, local_cache_size: int = 10000):
        self.redis = redis_client
        self.local_cache = {}  # LRU cache for hottest prefixes
        self.local_cache_size = local_cache_size

    def get(self, prefix: str, user_id: str = None) -> list:
        """Check caches in order: local -> Redis -> None."""
        # Level 1: Local in-memory (&#x3C; 0.1ms)
        cache_key = f"{prefix}:{user_id or 'global'}"
        if cache_key in self.local_cache:
            return self.local_cache[cache_key]

        # Level 2: Redis (&#x3C; 2ms)
        cached = self.redis.get(f"ac_cache:{cache_key}")
        if cached:
            result = json.loads(cached)
            self._update_local(cache_key, result)
            return result

        return None

    def set(self, prefix: str, results: list, user_id: str = None,
            ttl: int = 300):
        """Cache results at both levels."""
        cache_key = f"{prefix}:{user_id or 'global'}"
        self.redis.setex(f"ac_cache:{cache_key}", ttl, json.dumps(results))
        self._update_local(cache_key, results)

Key Takeaways

  • Sub-50ms latency requires a multi-stage architecture. No single approach achieves both relevance and speed. Use fast prefix retrieval, then enrich with semantics.
  • Semantic expansion catches 20-30% more relevant suggestions that pure prefix matching misses (typo-tolerance, synonyms, related concepts).
  • Personalization improves CTR by 15-25%. Even a simple user-history-based boost significantly outperforms one-size-fits-all ranking.
  • Cache aggressively at multiple levels. The top 1000 prefixes cover 60-70% of queries. Keep them in memory for sub-1ms responses.
  • Behavioral signals compound over time. Click-through data makes the system smarter with every interaction. Invest in signal collection infrastructure early.

Search autocomplete is where AI meets real-time systems engineering. The best systems feel magical to users precisely because the complexity is invisible - they just get relevant suggestions instantly, every time.

Comments

    No comments yet. Be the first to share your thoughts.