HNSW vs IVF: Vector Index Benchmarks for Production Similarity Search
Comprehensive benchmarks comparing HNSW and IVF vector indexes across recall, latency, memory usage, and build time for real-world embedding search workloads

Vector similarity search underpins modern AI applications - from semantic search and recommendations to RAG pipelines. The choice of index algorithm determines your system's recall, latency, and memory profile. HNSW and IVF are the two dominant approaches, each with distinct tradeoffs.
I benchmarked both algorithms across multiple dataset sizes, dimensionalities, and hardware configurations to provide concrete guidance for production deployments.
Algorithm Overview
HNSW (Hierarchical Navigable Small World) builds a multi-layer graph where each node connects to its approximate nearest neighbors. Search traverses from the top layer down, narrowing the candidate set at each level.
IVF (Inverted File Index) partitions the vector space into clusters using k-means, then searches only the closest clusters during query time. Variants include IVF-Flat (exact search within clusters) and IVF-PQ (product quantization for compression).
Benchmark Setup
| Parameter | Configuration |
|---|---|
| Datasets | SIFT-1M, GloVe-1.2M, Custom embeddings (5M, 25M) |
| Dimensions | 128, 384, 768, 1536 |
| Hardware | AWS r6i.4xlarge (128GB RAM), g5.xlarge (A10G GPU) |
| Libraries | FAISS 1.7.4, hnswlib 0.7.0, pgvector 0.5.1 |
| Metrics | Recall@10, QPS, P95 latency, memory, build time |
Core Benchmarks: 1M Vectors, 768 Dimensions
This represents a typical RAG or semantic search workload with OpenAI-style embeddings:
| Index | Recall@10 | QPS (single thread) | P95 Latency | Memory | Build Time |
|---|---|---|---|---|---|
| HNSW (M=32, ef=128) | 0.982 | 1,850 | 2.1ms | 6.2 GB | 14 min |
| HNSW (M=64, ef=256) | 0.995 | 920 | 4.3ms | 9.8 GB | 28 min |
| IVF-Flat (nlist=1024, nprobe=32) | 0.961 | 3,200 | 1.4ms | 3.1 GB | 8 min |
| IVF-Flat (nlist=4096, nprobe=128) | 0.988 | 1,100 | 3.8ms | 3.2 GB | 12 min |
| IVF-PQ (nlist=1024, m=48) | 0.924 | 8,500 | 0.5ms | 0.8 GB | 22 min |
| IVF-PQ (nlist=4096, m=96) | 0.952 | 4,200 | 0.9ms | 1.4 GB | 35 min |
Scaling Analysis: 25M Vectors
At larger scale, the tradeoffs shift significantly:
| Index | Recall@10 | QPS | Memory | Build Time |
|---|---|---|---|---|
| HNSW (M=32, ef=128) | 0.978 | 680 | 156 GB | 6.2 hours |
| HNSW (M=16, ef=64) | 0.951 | 1,450 | 82 GB | 3.1 hours |
| IVF-Flat (nlist=16384, nprobe=64) | 0.972 | 1,200 | 78 GB | 45 min |
| IVF-PQ (nlist=16384, m=64) | 0.941 | 12,000 | 8.2 GB | 2.1 hours |
| IVF-HNSW-PQ (hybrid) | 0.963 | 5,800 | 12 GB | 3.5 hours |
At 25M vectors, HNSW's memory overhead becomes the dominant constraint. IVF-PQ uses 19x less memory with acceptable recall loss.
Implementation: HNSW with FAISS
import faiss
import numpy as np
import time
class HNSWIndex:
def __init__(self, dim: int, M: int = 32, ef_construction: int = 200):
self.dim = dim
self.index = faiss.IndexHNSWFlat(dim, M)
self.index.hnsw.efConstruction = ef_construction
def build(self, vectors: np.ndarray):
"""Build index from numpy array of vectors."""
start = time.time()
self.index.add(vectors)
build_time = time.time() - start
print(f"Built HNSW index: {len(vectors)} vectors in {build_time:.1f}s")
return build_time
def search(self, query: np.ndarray, k: int = 10, ef_search: int = 128):
"""Search with configurable ef parameter."""
self.index.hnsw.efSearch = ef_search
distances, indices = self.index.search(query.reshape(1, -1), k)
return indices[0], distances[0]
def batch_search(self, queries: np.ndarray, k: int = 10, ef_search: int = 128):
"""Batch search for throughput benchmarking."""
self.index.hnsw.efSearch = ef_search
distances, indices = self.index.search(queries, k)
return indices, distances
Implementation: IVF with Product Quantization
class IVFPQIndex:
def __init__(self, dim: int, nlist: int = 4096, m: int = 64, nbits: int = 8):
self.dim = dim
self.nlist = nlist
# Coarse quantizer
quantizer = faiss.IndexFlatL2(dim)
self.index = faiss.IndexIVFPQ(quantizer, dim, nlist, m, nbits)
def build(self, vectors: np.ndarray):
"""Train and build IVF-PQ index."""
start = time.time()
# Train on a subset for large datasets
train_size = min(len(vectors), 500_000)
train_vectors = vectors[np.random.choice(len(vectors), train_size, replace=False)]
self.index.train(train_vectors)
self.index.add(vectors)
build_time = time.time() - start
print(f"Built IVF-PQ index: {len(vectors)} vectors in {build_time:.1f}s")
return build_time
def search(self, query: np.ndarray, k: int = 10, nprobe: int = 64):
"""Search with configurable nprobe."""
self.index.nprobe = nprobe
distances, indices = self.index.search(query.reshape(1, -1), k)
return indices[0], distances[0]
Recall vs Latency Tradeoffs
The key tuning parameters for each algorithm:
HNSW: ef_search parameter
| ef_search | Recall@10 | Latency (P95) | QPS |
|---|---|---|---|
| 32 | 0.921 | 0.8ms | 4,200 |
| 64 | 0.958 | 1.2ms | 3,100 |
| 128 | 0.982 | 2.1ms | 1,850 |
| 256 | 0.995 | 4.3ms | 920 |
| 512 | 0.998 | 8.7ms | 460 |
IVF-Flat: nprobe parameter
| nprobe | Recall@10 | Latency (P95) | QPS |
|---|---|---|---|
| 8 | 0.891 | 0.4ms | 8,500 |
| 16 | 0.932 | 0.7ms | 5,400 |
| 32 | 0.961 | 1.4ms | 3,200 |
| 64 | 0.981 | 2.8ms | 1,600 |
| 128 | 0.988 | 5.2ms | 780 |
Decision Framework
Use this framework to choose your index:
| Criteria | Best Choice | Reason |
|---|---|---|
| Recall > 0.99 required | HNSW (high ef) | Better recall ceiling |
| Memory constrained | IVF-PQ | 10-20x less memory |
| High throughput (>10K QPS) | IVF-PQ | Better parallelization |
| Frequent updates | HNSW | No retraining needed |
| Dataset > 50M vectors | IVF-PQ or hybrid | HNSW memory prohibitive |
| Low latency (< 1ms P95) | IVF-PQ | Fastest absolute latency |
| Cold start / fast builds | IVF-Flat | 3-10x faster index build |
Hybrid Approach: IVF-HNSW
FAISS supports a hybrid where IVF's coarse quantizer uses HNSW for faster cluster assignment:
def build_hybrid_index(vectors: np.ndarray, dim: int):
"""IVF with HNSW coarse quantizer - best of both worlds."""
nlist = int(np.sqrt(len(vectors))) # Rule of thumb
# HNSW as coarse quantizer
quantizer = faiss.IndexHNSWFlat(dim, 32)
quantizer.hnsw.efConstruction = 200
# IVF-PQ with HNSW quantizer
index = faiss.IndexIVFPQ(quantizer, dim, nlist, 64, 8)
# Train
train_size = min(len(vectors), 1_000_000)
train_data = vectors[:train_size]
index.train(train_data)
index.add(vectors)
return index
Production Deployment Recommendations
# Configuration for common workload sizes
CONFIGS = {
"small": { # < 1M vectors
"index": "HNSW",
"params": {"M": 32, "ef_construction": 200, "ef_search": 128},
"expected_memory": "~6GB per 1M vectors (768d)",
},
"medium": { # 1M - 10M vectors
"index": "IVF-HNSW-PQ",
"params": {"nlist": 8192, "m": 64, "nprobe": 64},
"expected_memory": "~1.5GB per 1M vectors (768d)",
},
"large": { # > 10M vectors
"index": "IVF-PQ",
"params": {"nlist": 65536, "m": 96, "nprobe": 128},
"expected_memory": "~0.5GB per 1M vectors (768d)",
},
}
Key Takeaways
- HNSW delivers the highest recall but at significant memory cost (6-10GB per million 768-dim vectors). Choose it when recall matters more than infrastructure cost.
- IVF-PQ is the scaling champion. At 25M+ vectors, its 20x memory advantage makes it the only practical choice for single-node deployments.
- The hybrid IVF-HNSW approach offers an excellent middle ground - near-HNSW recall with IVF-level memory efficiency.
- Tuning matters more than algorithm choice. A well-tuned IVF can outperform a default HNSW configuration. Invest time in parameter sweeps.
- Plan for growth. If you are at 1M vectors today and expect 10M in a year, start with IVF-PQ rather than migrating later under pressure.
The best vector index is the one that matches your specific requirements for recall, latency, memory, and update frequency. There is no universal winner.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.