Pinecone vs Weaviate vs pgvector: Production Benchmarks
Head-to-head comparison of vector databases under production load — latency, throughput, cost, and operational complexity with real benchmark data.

Choosing a vector database for production RAG is a consequential decision. After running all three — Pinecone, Weaviate, and pgvector — in production with 10M+ vectors, I have real numbers to share. The winner depends on your constraints.
Why this comparison exists
Every benchmark published by a vector database vendor shows their product winning. I ran my own benchmarks under conditions that match real production workloads: concurrent queries, mixed read/write traffic, and varying recall requirements.
Test methodology
All benchmarks ran against the same dataset: 12M embeddings (1536 dimensions, OpenAI text-embedding-3-small) from a production knowledge base.
Hardware baseline:
- Pinecone: p2 pod type, 2 replicas
- Weaviate: 3-node cluster on c5.2xlarge instances
- pgvector: RDS r6g.2xlarge with HNSW indexes
Query pattern: 70% similarity search, 20% filtered search, 10% upserts — mirroring our production traffic.
Latency benchmarks
P50 query latency (ms)
| Concurrent Users | Pinecone | Weaviate | pgvector |
|---|---|---|---|
| 10 | 12 | 18 | 8 |
| 50 | 15 | 24 | 22 |
| 100 | 19 | 31 | 45 |
| 500 | 28 | 52 | 180 |
| 1000 | 41 | 78 | 420 |
P99 query latency (ms)
| Concurrent Users | Pinecone | Weaviate | pgvector |
|---|---|---|---|
| 10 | 35 | 45 | 15 |
| 50 | 48 | 62 | 55 |
| 100 | 65 | 95 | 120 |
| 500 | 110 | 180 | 850 |
| 1000 | 185 | 310 | 2100 |
pgvector wins at low concurrency. It lives in the same database as your application data, so there's no network hop. But it degrades badly under load because PostgreSQL's connection model wasn't designed for vector similarity workloads.
Recall accuracy at scale
Recall@10 with HNSW index (ef_search=128):
| Dataset Size | Pinecone | Weaviate | pgvector |
|---|---|---|---|
| 100K vectors | 0.98 | 0.97 | 0.99 |
| 1M vectors | 0.97 | 0.96 | 0.97 |
| 10M vectors | 0.96 | 0.94 | 0.95 |
All three deliver acceptable recall. Pinecone maintains the most consistent recall as dataset size grows, likely due to their proprietary indexing optimizations.
The code: querying each system
Pinecone
from pinecone import Pinecone
pc = Pinecone(api_key="your-key")
index = pc.Index("production-embeddings")
def search_pinecone(query_vector: list[float], top_k: int = 10, filters: dict = None):
results = index.query(
vector=query_vector,
top_k=top_k,
filter=filters,
include_metadata=True
)
return [
{"id": match.id, "score": match.score, "metadata": match.metadata}
for match in results.matches
]
# Filtered search example
results = search_pinecone(
query_vector=embedding,
filters={"category": {"$eq": "engineering"}, "date": {"$gte": "2025-01-01"}}
)
pgvector with connection pooling
import { Pool } from "pg";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Critical for pgvector performance
});
async function searchPgvector(
queryVector: number[],
topK: number = 10,
category?: string,
): Promise<SearchResult[]> {
const vectorStr = `[${queryVector.join(",")}]`;
let query = `
SELECT id, content, metadata,
1 - (embedding <=> $1::vector) as similarity
FROM documents
`;
const params: any[] = [vectorStr];
if (category) {
query += ` WHERE metadata->>'category' = $2`;
params.push(category);
}
query += ` ORDER BY embedding <=> $1::vector LIMIT $${params.length + 1}`;
params.push(topK);
const { rows } = await pool.query(query, params);
return rows;
}
Cost comparison (monthly, 10M vectors)
| Component | Pinecone | Weaviate (self-hosted) | pgvector |
|---|---|---|---|
| Compute/hosting | $280 | $520 | $380 |
| Storage | included | $45 | $25 |
| Network/transfer | $15 | $30 | $0 |
| Operational overhead | Low | High | Medium |
| Total | $295 | $595 | $405 |
Pinecone is cheapest when you factor in operational cost. Self-hosted Weaviate requires cluster management, upgrades, and monitoring that eat engineering time. pgvector piggybacks on existing Postgres infrastructure but needs careful tuning.
Operational complexity
Pinecone: Fully managed. Zero operational burden. You trade control for simplicity. Scaling is a slider. The downside: vendor lock-in and limited query expressiveness.
Weaviate: Powerful hybrid search (vector + BM25), great schema flexibility. But running a Weaviate cluster is real infrastructure work — node failures, rebalancing, version upgrades.
pgvector: If you're already running PostgreSQL, the marginal complexity is low. But you inherit all of Postgres's limitations for this workload: no horizontal scaling, shared resources with OLTP queries, and VACUUM pressure from frequent updates.
When to choose each
Choose Pinecone when:
- You need predictable latency at high concurrency (>200 QPS)
- Your team doesn't have dedicated infrastructure engineers
- Vendor lock-in is acceptable for your use case
- You need metadata filtering with consistent performance
Choose Weaviate when:
- You need hybrid search (vector + keyword)
- Your data model is complex with multiple vector fields
- You want to self-host for compliance/data residency
- You have infrastructure engineers to manage the cluster
Choose pgvector when:
- Your dataset is under 5M vectors
- Query concurrency stays below 100 QPS
- You want to keep vectors alongside relational data
- You're optimizing for simplicity over peak performance
- You need transactional consistency between vectors and metadata
The migration lesson
We started with pgvector because it was familiar. At 2M vectors and 50 QPS, it was fine. At 8M vectors and 200 QPS, p99 latencies crossed 2 seconds. We migrated to Pinecone for the hot path and kept pgvector for batch analytics queries.
The migration took 3 weeks. If I were starting over, I'd begin with the managed option and only self-host when specific requirements demand it.
Key takeaways
- pgvector is excellent for prototyping and low-traffic production, but plan your escape route early
- Pinecone wins on latency predictability and operational simplicity at scale
- Weaviate's hybrid search is genuinely differentiated if your retrieval needs keyword + semantic
- Always benchmark with your actual query patterns — synthetic benchmarks mislead
- The cost difference between options is smaller than the engineering time cost of a mid-growth migration
Measure at your expected scale, not your current scale. The database that works at 100K vectors may not work at 10M.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.