Pinecone vs Weaviate vs pgvector: Production Benchmarks

Head-to-head comparison of vector databases under production load — latency, throughput, cost, and operational complexity with real benchmark data.

#vector-database#ai#rag#comparison
Cover image for the article: Pinecone vs Weaviate vs pgvector: Production Benchmarks

Choosing a vector database for production RAG is a consequential decision. After running all three — Pinecone, Weaviate, and pgvector — in production with 10M+ vectors, I have real numbers to share. The winner depends on your constraints.

Why this comparison exists

Every benchmark published by a vector database vendor shows their product winning. I ran my own benchmarks under conditions that match real production workloads: concurrent queries, mixed read/write traffic, and varying recall requirements.

Test methodology

All benchmarks ran against the same dataset: 12M embeddings (1536 dimensions, OpenAI text-embedding-3-small) from a production knowledge base.

Hardware baseline:

  • Pinecone: p2 pod type, 2 replicas
  • Weaviate: 3-node cluster on c5.2xlarge instances
  • pgvector: RDS r6g.2xlarge with HNSW indexes

Query pattern: 70% similarity search, 20% filtered search, 10% upserts — mirroring our production traffic.

Vector Database Architecture Comparison

Latency benchmarks

P50 query latency (ms)

Concurrent UsersPineconeWeaviatepgvector
1012188
50152422
100193145
5002852180
10004178420

P99 query latency (ms)

Concurrent UsersPineconeWeaviatepgvector
10354515
50486255
1006595120
500110180850
10001853102100

pgvector wins at low concurrency. It lives in the same database as your application data, so there's no network hop. But it degrades badly under load because PostgreSQL's connection model wasn't designed for vector similarity workloads.

Recall accuracy at scale

Recall@10 with HNSW index (ef_search=128):

Dataset SizePineconeWeaviatepgvector
100K vectors0.980.970.99
1M vectors0.970.960.97
10M vectors0.960.940.95

All three deliver acceptable recall. Pinecone maintains the most consistent recall as dataset size grows, likely due to their proprietary indexing optimizations.

The code: querying each system

Pinecone

from pinecone import Pinecone

pc = Pinecone(api_key="your-key")
index = pc.Index("production-embeddings")

def search_pinecone(query_vector: list[float], top_k: int = 10, filters: dict = None):
    results = index.query(
        vector=query_vector,
        top_k=top_k,
        filter=filters,
        include_metadata=True
    )
    return [
        {"id": match.id, "score": match.score, "metadata": match.metadata}
        for match in results.matches
    ]

# Filtered search example
results = search_pinecone(
    query_vector=embedding,
    filters={"category": {"$eq": "engineering"}, "date": {"$gte": "2025-01-01"}}
)

pgvector with connection pooling

import { Pool } from "pg";

const pool = new Pool({
  connectionString: process.env.DATABASE_URL,
  max: 20, // Critical for pgvector performance
});

async function searchPgvector(
  queryVector: number[],
  topK: number = 10,
  category?: string,
): Promise<SearchResult[]> {
  const vectorStr = `[${queryVector.join(",")}]`;

  let query = `
    SELECT id, content, metadata,
           1 - (embedding <=> $1::vector) as similarity
    FROM documents
  `;
  const params: any[] = [vectorStr];

  if (category) {
    query += ` WHERE metadata->>'category' = $2`;
    params.push(category);
  }

  query += ` ORDER BY embedding <=> $1::vector LIMIT $${params.length + 1}`;
  params.push(topK);

  const { rows } = await pool.query(query, params);
  return rows;
}

Cost comparison (monthly, 10M vectors)

ComponentPineconeWeaviate (self-hosted)pgvector
Compute/hosting$280$520$380
Storageincluded$45$25
Network/transfer$15$30$0
Operational overheadLowHighMedium
Total$295$595$405

Pinecone is cheapest when you factor in operational cost. Self-hosted Weaviate requires cluster management, upgrades, and monitoring that eat engineering time. pgvector piggybacks on existing Postgres infrastructure but needs careful tuning.

Operational complexity

Pinecone: Fully managed. Zero operational burden. You trade control for simplicity. Scaling is a slider. The downside: vendor lock-in and limited query expressiveness.

Weaviate: Powerful hybrid search (vector + BM25), great schema flexibility. But running a Weaviate cluster is real infrastructure work — node failures, rebalancing, version upgrades.

pgvector: If you're already running PostgreSQL, the marginal complexity is low. But you inherit all of Postgres's limitations for this workload: no horizontal scaling, shared resources with OLTP queries, and VACUUM pressure from frequent updates.

When to choose each

Choose Pinecone when:

  • You need predictable latency at high concurrency (>200 QPS)
  • Your team doesn't have dedicated infrastructure engineers
  • Vendor lock-in is acceptable for your use case
  • You need metadata filtering with consistent performance

Choose Weaviate when:

  • You need hybrid search (vector + keyword)
  • Your data model is complex with multiple vector fields
  • You want to self-host for compliance/data residency
  • You have infrastructure engineers to manage the cluster

Choose pgvector when:

  • Your dataset is under 5M vectors
  • Query concurrency stays below 100 QPS
  • You want to keep vectors alongside relational data
  • You're optimizing for simplicity over peak performance
  • You need transactional consistency between vectors and metadata

The migration lesson

We started with pgvector because it was familiar. At 2M vectors and 50 QPS, it was fine. At 8M vectors and 200 QPS, p99 latencies crossed 2 seconds. We migrated to Pinecone for the hot path and kept pgvector for batch analytics queries.

The migration took 3 weeks. If I were starting over, I'd begin with the managed option and only self-host when specific requirements demand it.

Key takeaways

  • pgvector is excellent for prototyping and low-traffic production, but plan your escape route early
  • Pinecone wins on latency predictability and operational simplicity at scale
  • Weaviate's hybrid search is genuinely differentiated if your retrieval needs keyword + semantic
  • Always benchmark with your actual query patterns — synthetic benchmarks mislead
  • The cost difference between options is smaller than the engineering time cost of a mid-growth migration

Measure at your expected scale, not your current scale. The database that works at 100K vectors may not work at 10M.

Comments

    No comments yet. Be the first to share your thoughts.