Choosing Embedding Models: Quality vs Latency vs Cost

A practical guide to selecting embedding models for production — benchmarking OpenAI, Cohere, Voyage AI, and open-source options across quality, speed, and cost.

#embeddings#ai#vector-search#comparison
Cover image for the article: Choosing Embedding Models: Quality vs Latency vs Cost

The embedding model you choose determines the ceiling of your retrieval system. A perfect reranker can't fix embeddings that don't capture semantic similarity. After testing 12 embedding models across 4 production use cases, here's what actually matters.

Why embedding selection is hard

Leaderboard rankings (MTEB, BEIR) measure general performance across academic datasets. Production workloads are specific: your domain vocabulary, your document lengths, your query patterns. A model that ranks #1 on MTEB might rank #5 for your legal document retrieval task.

The three dimensions that matter:

  1. Retrieval quality — does it find the right documents?
  2. Latency — can you embed at query time without noticeable delay?
  3. Cost — what's the per-document and per-query expense at scale?

You're always trading between these three.

Models tested

ModelDimensionsMax TokensProvider
text-embedding-3-large30728191OpenAI
text-embedding-3-small15368191OpenAI
embed-v41024512Cohere
voyage-3102416000Voyage AI
e5-large-v21024512Open-source
bge-large-en-v1.51024512Open-source
nomic-embed-text-v1.57688192Nomic (open)

Embedding Model Decision Tree

Quality benchmarks (domain-specific)

I tested against our production dataset: 500K technical documents with 2,000 labeled query-document pairs across four domains.

Recall@10 by domain

ModelEngineering DocsLegalMedicalGeneral
text-embedding-3-large0.940.910.890.96
voyage-30.930.930.910.94
embed-v40.910.900.880.93
text-embedding-3-small0.890.870.850.92
nomic-embed-text-v1.50.880.850.830.91
bge-large-en-v1.50.870.860.840.90
e5-large-v20.860.840.820.89

Voyage-3 excels on domain-specific content (legal, medical) due to its longer context window capturing more document structure. OpenAI's large model wins on engineering docs and general queries.

Latency benchmarks

Measured at query time — single embedding generation latency.

ModelP50 (ms)P99 (ms)Self-hosted?
text-embedding-3-large45120No
text-embedding-3-small2875No
embed-v43595No
voyage-352140No
nomic-embed-text-v1.5815Yes (GPU)
bge-large-en-v1.51222Yes (GPU)
e5-large-v21425Yes (GPU)

Self-hosted models on a single A10G GPU are 3-5x faster at inference because there's no network round-trip. The quality trade-off is real but manageable for many use cases.

Cost analysis (1M documents, 500 tokens avg)

ModelEmbedding CostMonthly Query Cost (100K/day)Storage (1M vectors)
text-embedding-3-large$65$6.50/day12 GB
text-embedding-3-small$10$1.00/day6 GB
embed-v4$5.50$0.55/day4 GB
voyage-3$60$6.00/day4 GB
nomic-embed-text (self-hosted)$0*$0*3 GB
bge-large (self-hosted)$0*$0*4 GB

*Self-hosted: GPU instance cost ($0.50-$1.50/hr) amortized across all workloads.

The evaluation pipeline

Don't trust benchmarks — run your own evaluation. Here's the framework:

import numpy as np
from dataclasses import dataclass
from typing import Protocol

class EmbeddingProvider(Protocol):
    async def embed(self, texts: list[str]) -> list[list[float]]: ...

@dataclass
class EvalResult:
    model: str
    recall_at_10: float
    mrr: float
    latency_p50_ms: float
    latency_p99_ms: float
    cost_per_1k_embeddings: float

async def evaluate_model(
    provider: EmbeddingProvider,
    queries: list[str],
    relevant_docs: dict[str, list[str]],  # query -> relevant doc IDs
    corpus_embeddings: np.ndarray,
    corpus_ids: list[str],
) -> EvalResult:
    """Evaluate an embedding model on your domain-specific data."""
    recalls = []
    mrrs = []

    for query in queries:
        query_embedding = await provider.embed([query])
        query_vec = np.array(query_embedding[0])

        # Cosine similarity
        similarities = np.dot(corpus_embeddings, query_vec) / (
            np.linalg.norm(corpus_embeddings, axis=1) * np.linalg.norm(query_vec)
        )

        top_10_indices = np.argsort(similarities)[-10:][::-1]
        top_10_ids = [corpus_ids[i] for i in top_10_indices]

        # Recall@10
        relevant = set(relevant_docs[query])
        retrieved = set(top_10_ids)
        recalls.append(len(relevant & retrieved) / len(relevant))

        # MRR
        for rank, doc_id in enumerate(top_10_ids, 1):
            if doc_id in relevant:
                mrrs.append(1.0 / rank)
                break
        else:
            mrrs.append(0.0)

    return EvalResult(
        model=provider.name,
        recall_at_10=np.mean(recalls),
        mrr=np.mean(mrrs),
        latency_p50_ms=provider.get_latency_p50(),
        latency_p99_ms=provider.get_latency_p99(),
        cost_per_1k_embeddings=provider.cost_per_1k,
    )

Running evaluations with multiple providers

import { OpenAI } from "openai";

interface EmbeddingConfig {
  name: string;
  provider: "openai" | "cohere" | "voyage" | "local";
  model: string;
  dimensions: number;
  batchSize: number;
}

const CONFIGS: EmbeddingConfig[] = [
  {
    name: "openai-large",
    provider: "openai",
    model: "text-embedding-3-large",
    dimensions: 3072,
    batchSize: 2048,
  },
  {
    name: "openai-small",
    provider: "openai",
    model: "text-embedding-3-small",
    dimensions: 1536,
    batchSize: 2048,
  },
];

async function embedBatch(
  config: EmbeddingConfig,
  texts: string[],
): Promise<number[][]> {
  const client = new OpenAI();
  const batches: number[][][] = [];

  for (let i = 0; i < texts.length; i += config.batchSize) {
    const batch = texts.slice(i, i + config.batchSize);
    const response = await client.embeddings.create({
      model: config.model,
      input: batch,
      dimensions: config.dimensions,
    });
    batches.push(response.data.map((d) => d.embedding));
  }

  return batches.flat();
}

The dimensionality trade-off

Higher dimensions capture more nuance but cost more to store and search. OpenAI's Matryoshka embeddings let you truncate dimensions with graceful quality degradation:

DimensionsRecall@10Storage per 1MSearch Latency
30720.9412 GB45ms
15360.926 GB28ms
7680.893 GB18ms
2560.831 GB8ms

For most production systems, 1536 dimensions hits the sweet spot — 2% quality loss for 50% storage and latency savings.

Decision framework

Choose text-embedding-3-large when:

  • Maximum retrieval quality is non-negotiable
  • You're indexing <1M documents
  • Query latency budget is >100ms

Choose text-embedding-3-small when:

  • Cost is a constraint and quality at 1536d is acceptable
  • High query volume (>500K/day)
  • You need the OpenAI ecosystem integration

Choose Voyage-3 when:

  • Domain-specific content (legal, medical, scientific)
  • Documents are long (>2000 tokens)
  • You need best-in-class on specialized benchmarks

Choose self-hosted (nomic/bge) when:

  • Latency is critical (<20ms query embedding)
  • Data can't leave your infrastructure
  • You have GPU capacity available
  • Cost must be near-zero at scale

Key takeaways

  • Run domain-specific evaluations — MTEB rankings don't predict your production performance
  • Latency matters more than you think — 50ms embedding time adds up when you're doing re-ranking
  • Matryoshka truncation is an underused lever — test reduced dimensions before paying for full
  • Self-hosted models are viable for production if you have GPU infrastructure
  • The best model is the one you can afford to run at your query volume

Start with text-embedding-3-small for prototyping, benchmark against your domain data, then scale to what the quality metrics demand.

Comments

    No comments yet. Be the first to share your thoughts.