Hybrid Search: Combining Keyword and Semantic Retrieval for Production

Building a production hybrid search system that combines BM25 keyword matching with vector semantic search for superior relevance at scale.

#ai-search#hybrid-retrieval#vector-search#production
Cover image for the article: Hybrid Search: Combining Keyword and Semantic Retrieval for Production

Neither pure keyword search nor pure semantic search delivers optimal results alone. Keyword search excels at exact matches and rare terms but misses semantic intent. Vector search captures meaning but struggles with specific identifiers, acronyms, and exact phrases. Hybrid search combines both, and getting the combination right is the difference between "pretty good" and "users stop complaining about search."

After building hybrid search systems serving 20M queries/day with p95 latency under 120ms, here's the architecture that delivers consistently superior relevance.

The Problem: Single-Method Search Limitations

Consider searching a technical knowledge base for "k8s OOM killed pod restart loop." Pure keyword search finds documents containing those exact terms but misses articles about "Kubernetes memory limit exceeded causing container restarts." Pure vector search captures the semantic intent but might rank a general article about memory management above a specific troubleshooting guide that uses the exact error message.

Real users need both:

  • Keyword precision: "error code E-4521" must match exactly
  • Semantic understanding: "how to reduce cloud costs" should match "infrastructure spend optimization"
  • Hybrid intelligence: Combine signals for the best of both worlds

Architecture: Dual-Retrieval with Fusion

The system runs keyword and vector retrieval in parallel, then fuses results using a learned scoring function that adapts to query characteristics.

Hybrid Search Architecture

Dual Retrieval Engine

The retrieval engine queries both keyword and vector indexes in parallel, then combines results using Reciprocal Rank Fusion (RRF) with learned weights.

import numpy as np
from dataclasses import dataclass
from typing import Optional
import asyncio
import time

@dataclass
class SearchResult:
    document_id: str
    title: str
    content_snippet: str
    keyword_score: Optional[float] = None
    vector_score: Optional[float] = None
    fusion_score: float = 0.0
    source: str = "hybrid"

@dataclass
class QueryAnalysis:
    original_query: str
    is_exact_match: bool
    has_identifiers: bool
    semantic_complexity: float
    suggested_keyword_weight: float
    suggested_vector_weight: float

class HybridSearchEngine:
    def __init__(self, keyword_index, vector_index, embedding_model,
                 rrf_k: int = 60, default_keyword_weight: float = 0.4,
                 default_vector_weight: float = 0.6):
        self.keyword_index = keyword_index
        self.vector_index = vector_index
        self.embedding_model = embedding_model
        self.rrf_k = rrf_k
        self.default_kw_weight = default_keyword_weight
        self.default_vec_weight = default_vector_weight

    async def search(self, query: str, top_k: int = 20, 
                     filters: Optional[dict] = None) -> list[SearchResult]:
        # Analyze query to determine optimal weighting
        query_analysis = self._analyze_query(query)

        # Run both retrievers in parallel
        keyword_task = asyncio.create_task(
            self._keyword_search(query, top_k * 2, filters)
        )
        vector_task = asyncio.create_task(
            self._vector_search(query, top_k * 2, filters)
        )

        keyword_results, vector_results = await asyncio.gather(
            keyword_task, vector_task
        )

        # Fuse results with adaptive weighting
        fused = self._reciprocal_rank_fusion(
            keyword_results=keyword_results,
            vector_results=vector_results,
            keyword_weight=query_analysis.suggested_keyword_weight,
            vector_weight=query_analysis.suggested_vector_weight,
        )

        return fused[:top_k]

    def _analyze_query(self, query: str) -> QueryAnalysis:
        import re
        # Detect exact match intent (quotes, specific IDs, error codes)
        has_quotes = '"' in query
        has_identifiers = bool(re.search(r'[A-Z]{2,}-\d+|error\s+\w+\d+|v\d+\.\d+', query))
        
        # Short queries with specific terms favor keyword search
        word_count = len(query.split())
        is_exact = has_quotes or (has_identifiers and word_count <= 3)

        # Longer, natural language queries favor semantic search
        semantic_complexity = min(word_count / 10.0, 1.0)

        if is_exact:
            kw_weight, vec_weight = 0.8, 0.2
        elif has_identifiers:
            kw_weight, vec_weight = 0.6, 0.4
        elif semantic_complexity > 0.5:
            kw_weight, vec_weight = 0.3, 0.7
        else:
            kw_weight, vec_weight = self.default_kw_weight, self.default_vec_weight

        return QueryAnalysis(
            original_query=query,
            is_exact_match=is_exact,
            has_identifiers=has_identifiers,
            semantic_complexity=semantic_complexity,
            suggested_keyword_weight=kw_weight,
            suggested_vector_weight=vec_weight,
        )

    def _reciprocal_rank_fusion(
        self,
        keyword_results: list[SearchResult],
        vector_results: list[SearchResult],
        keyword_weight: float,
        vector_weight: float,
    ) -> list[SearchResult]:
        scores: dict[str, float] = {}
        result_map: dict[str, SearchResult] = {}

        # Score keyword results by reciprocal rank
        for rank, result in enumerate(keyword_results):
            rrf_score = keyword_weight / (self.rrf_k + rank + 1)
            scores[result.document_id] = scores.get(result.document_id, 0) + rrf_score
            if result.document_id not in result_map:
                result_map[result.document_id] = result
            result_map[result.document_id].keyword_score = result.keyword_score

        # Score vector results by reciprocal rank
        for rank, result in enumerate(vector_results):
            rrf_score = vector_weight / (self.rrf_k + rank + 1)
            scores[result.document_id] = scores.get(result.document_id, 0) + rrf_score
            if result.document_id not in result_map:
                result_map[result.document_id] = result
            result_map[result.document_id].vector_score = result.vector_score

        # Sort by fusion score
        sorted_ids = sorted(scores.keys(), key=lambda x: scores[x], reverse=True)
        results = []
        for doc_id in sorted_ids:
            result = result_map[doc_id]
            result.fusion_score = scores[doc_id]
            result.source = "hybrid"
            results.append(result)

        return results

    async def _keyword_search(self, query: str, top_k: int, 
                              filters: Optional[dict]) -> list[SearchResult]:
        # BM25 search against Elasticsearch/OpenSearch
        return []

    async def _vector_search(self, query: str, top_k: int, 
                             filters: Optional[dict]) -> list[SearchResult]:
        # Embed query, then ANN search against vector index
        embedding = await self.embedding_model.encode(query)
        return []

Indexing Pipeline with Chunking Strategy

Documents must be chunked intelligently for both keyword and vector indexes. The chunking strategy balances context preservation with retrieval granularity.

interface DocumentChunk {
  chunkId: string;
  documentId: string;
  content: string;
  embedding?: number[];
  metadata: {
    title: string;
    section: string;
    chunkIndex: number;
    totalChunks: number;
    tokenCount: number;
    overlap: number;
  };
}

interface ChunkingConfig {
  maxTokens: number;
  overlapTokens: number;
  splitStrategy: 'sentence' | 'paragraph' | 'semantic';
  preserveHeaders: boolean;
  minChunkTokens: number;
}

class DocumentIndexer {
  private chunkConfig: ChunkingConfig;
  private embeddingModel: { encode: (texts: string[]) => Promise<number[][]> };
  private keywordIndex: { index: (id: string, doc: Record<string, unknown>) => Promise<void> };
  private vectorIndex: { upsert: (vectors: Array<{ id: string; values: number[]; metadata: Record<string, unknown> }>) => Promise<void> };

  constructor(
    chunkConfig: ChunkingConfig,
    embeddingModel: { encode: (texts: string[]) => Promise<number[][]> },
    keywordIndex: { index: (id: string, doc: Record<string, unknown>) => Promise<void> },
    vectorIndex: { upsert: (vectors: Array<{ id: string; values: number[]; metadata: Record<string, unknown> }>) => Promise<void> }
  ) {
    this.chunkConfig = chunkConfig;
    this.embeddingModel = embeddingModel;
    this.keywordIndex = keywordIndex;
    this.vectorIndex = vectorIndex;
  }

  async indexDocument(documentId: string, title: string, content: string): Promise<DocumentChunk[]> {
    const chunks = this.chunkDocument(documentId, title, content);

    // Batch embed all chunks
    const texts = chunks.map(c => c.content);
    const embeddings = await this.embeddingModel.encode(texts);

    for (let i = 0; i < chunks.length; i++) {
      chunks[i].embedding = embeddings[i];
    }

    // Index in both stores in parallel
    await Promise.all([
      this.indexKeyword(chunks),
      this.indexVector(chunks),
    ]);

    return chunks;
  }

  private chunkDocument(documentId: string, title: string, content: string): DocumentChunk[] {
    const sections = this.splitIntoSections(content);
    const chunks: DocumentChunk[] = [];
    let chunkIndex = 0;

    for (const section of sections) {
      const sectionChunks = this.splitSection(section.content, this.chunkConfig);

      for (const chunkContent of sectionChunks) {
        chunks.push({
          chunkId: `${documentId}_chunk_${chunkIndex}`,
          documentId,
          content: this.chunkConfig.preserveHeaders
            ? `${title} > ${section.header}\n\n${chunkContent}`
            : chunkContent,
          metadata: {
            title,
            section: section.header,
            chunkIndex,
            totalChunks: 0, // Updated after all chunks created
            tokenCount: this.countTokens(chunkContent),
            overlap: this.chunkConfig.overlapTokens,
          },
        });
        chunkIndex++;
      }
    }

    // Update total count
    for (const chunk of chunks) {
      chunk.metadata.totalChunks = chunks.length;
    }

    return chunks;
  }

  private splitIntoSections(content: string): Array<{ header: string; content: string }> {
    const headerPattern = /^#{1,3}\s+(.+)$/gm;
    const sections: Array<{ header: string; content: string }> = [];
    let lastIndex = 0;
    let lastHeader = 'Introduction';
    let match: RegExpExecArray | null;

    while ((match = headerPattern.exec(content)) !== null) {
      if (match.index > lastIndex) {
        sections.push({ header: lastHeader, content: content.slice(lastIndex, match.index).trim() });
      }
      lastHeader = match[1];
      lastIndex = match.index + match[0].length;
    }

    if (lastIndex < content.length) {
      sections.push({ header: lastHeader, content: content.slice(lastIndex).trim() });
    }

    return sections.filter(s => s.content.length > 0);
  }

  private splitSection(content: string, config: ChunkingConfig): string[] {
    const sentences = content.split(/(?<=[.!?])\s+/);
    const chunks: string[] = [];
    let current: string[] = [];
    let currentTokens = 0;

    for (const sentence of sentences) {
      const sentenceTokens = this.countTokens(sentence);

      if (currentTokens + sentenceTokens > config.maxTokens && current.length > 0) {
        chunks.push(current.join(' '));
        // Keep overlap sentences
        const overlapSentences = this.getOverlapSentences(current, config.overlapTokens);
        current = overlapSentences;
        currentTokens = this.countTokens(current.join(' '));
      }

      current.push(sentence);
      currentTokens += sentenceTokens;
    }

    if (current.length > 0 && currentTokens >= config.minChunkTokens) {
      chunks.push(current.join(' '));
    }

    return chunks;
  }

  private getOverlapSentences(sentences: string[], targetTokens: number): string[] {
    const overlap: string[] = [];
    let tokens = 0;
    for (let i = sentences.length - 1; i >= 0 && tokens < targetTokens; i--) {
      overlap.unshift(sentences[i]);
      tokens += this.countTokens(sentences[i]);
    }
    return overlap;
  }

  private countTokens(text: string): number {
    return Math.ceil(text.length / 4); // Approximate
  }

  private async indexKeyword(chunks: DocumentChunk[]): Promise<void> {
    for (const chunk of chunks) {
      await this.keywordIndex.index(chunk.chunkId, {
        content: chunk.content,
        document_id: chunk.documentId,
        title: chunk.metadata.title,
        section: chunk.metadata.section,
      });
    }
  }

  private async indexVector(chunks: DocumentChunk[]): Promise<void> {
    const vectors = chunks.map(chunk => ({
      id: chunk.chunkId,
      values: chunk.embedding!,
      metadata: {
        documentId: chunk.documentId,
        title: chunk.metadata.title,
        section: chunk.metadata.section,
        tokenCount: chunk.metadata.tokenCount,
      },
    }));
    await this.vectorIndex.upsert(vectors);
  }
}

Query Enhancement Techniques

Query Expansion

For keyword search, expand the query with synonyms and related terms. For vector search, generate multiple query embeddings from different phrasings and combine results.

Contextual Re-ranking

After initial retrieval, apply a cross-encoder re-ranker that scores query-document pairs together. This is more expensive but significantly more accurate than bi-encoder similarity.

Negative Feedback Loop

Track queries where users click past the first page or reformulate their query. Use these as training signals for improving the fusion weights.

Benchmarks: Hybrid vs Single-Method

Evaluated on an internal technical knowledge base with 500K documents and 10K annotated queries:

MethodNDCG@10MRRP@5Latency (p95)
BM25 keyword only0.420.510.3835ms
Vector only (e5-large)0.480.560.4485ms
Hybrid (fixed weights)0.550.630.5195ms
Hybrid (adaptive weights)0.610.690.57105ms
Hybrid + re-ranker0.670.740.63118ms

Production Performance at Scale

MetricValue
Daily queries20M
Index size (documents)2.4M
Index size (chunks)18M
P50 latency42ms
P95 latency112ms
P99 latency185ms

Lessons Learned

Adaptive weighting matters more than the fusion algorithm. RRF with query-adaptive weights outperforms sophisticated learned fusion with static weights. Invest in query analysis over fusion complexity.

Chunk size affects keyword and vector differently. Smaller chunks (128-256 tokens) work better for vector search. Larger chunks (512-1024 tokens) work better for keyword search. Index at multiple granularities if you can afford the storage.

Exact match signals are undervalued. When a user query exactly matches a document title or section header, that should dominate all other signals regardless of semantic scores.

Monitor retrieval diversity. Pure vector search tends to return semantically similar but redundant results. Introduce diversity penalties (MMR) to ensure the top results cover different aspects of the query.

Conclusion

Hybrid search is not just keyword + vector — it's intelligent routing that adapts to query intent. The combination of query analysis, parallel retrieval, adaptive fusion, and cross-encoder re-ranking delivers relevance that neither method achieves alone. Start with RRF fusion using fixed weights, add query analysis for adaptive weighting, and introduce re-ranking for your highest-value queries.

Comments

    No comments yet. Be the first to share your thoughts.