Choosing Embedding Models: Quality vs Latency vs Cost
A practical guide to selecting embedding models for production — benchmarking OpenAI, Cohere, Voyage AI, and open-source options across quality, speed, and cost.

The embedding model you choose determines the ceiling of your retrieval system. A perfect reranker can't fix embeddings that don't capture semantic similarity. After testing 12 embedding models across 4 production use cases, here's what actually matters.
Why embedding selection is hard
Leaderboard rankings (MTEB, BEIR) measure general performance across academic datasets. Production workloads are specific: your domain vocabulary, your document lengths, your query patterns. A model that ranks #1 on MTEB might rank #5 for your legal document retrieval task.
The three dimensions that matter:
- Retrieval quality — does it find the right documents?
- Latency — can you embed at query time without noticeable delay?
- Cost — what's the per-document and per-query expense at scale?
You're always trading between these three.
Models tested
| Model | Dimensions | Max Tokens | Provider |
|---|---|---|---|
| text-embedding-3-large | 3072 | 8191 | OpenAI |
| text-embedding-3-small | 1536 | 8191 | OpenAI |
| embed-v4 | 1024 | 512 | Cohere |
| voyage-3 | 1024 | 16000 | Voyage AI |
| e5-large-v2 | 1024 | 512 | Open-source |
| bge-large-en-v1.5 | 1024 | 512 | Open-source |
| nomic-embed-text-v1.5 | 768 | 8192 | Nomic (open) |
Quality benchmarks (domain-specific)
I tested against our production dataset: 500K technical documents with 2,000 labeled query-document pairs across four domains.
Recall@10 by domain
| Model | Engineering Docs | Legal | Medical | General |
|---|---|---|---|---|
| text-embedding-3-large | 0.94 | 0.91 | 0.89 | 0.96 |
| voyage-3 | 0.93 | 0.93 | 0.91 | 0.94 |
| embed-v4 | 0.91 | 0.90 | 0.88 | 0.93 |
| text-embedding-3-small | 0.89 | 0.87 | 0.85 | 0.92 |
| nomic-embed-text-v1.5 | 0.88 | 0.85 | 0.83 | 0.91 |
| bge-large-en-v1.5 | 0.87 | 0.86 | 0.84 | 0.90 |
| e5-large-v2 | 0.86 | 0.84 | 0.82 | 0.89 |
Voyage-3 excels on domain-specific content (legal, medical) due to its longer context window capturing more document structure. OpenAI's large model wins on engineering docs and general queries.
Latency benchmarks
Measured at query time — single embedding generation latency.
| Model | P50 (ms) | P99 (ms) | Self-hosted? |
|---|---|---|---|
| text-embedding-3-large | 45 | 120 | No |
| text-embedding-3-small | 28 | 75 | No |
| embed-v4 | 35 | 95 | No |
| voyage-3 | 52 | 140 | No |
| nomic-embed-text-v1.5 | 8 | 15 | Yes (GPU) |
| bge-large-en-v1.5 | 12 | 22 | Yes (GPU) |
| e5-large-v2 | 14 | 25 | Yes (GPU) |
Self-hosted models on a single A10G GPU are 3-5x faster at inference because there's no network round-trip. The quality trade-off is real but manageable for many use cases.
Cost analysis (1M documents, 500 tokens avg)
| Model | Embedding Cost | Monthly Query Cost (100K/day) | Storage (1M vectors) |
|---|---|---|---|
| text-embedding-3-large | $65 | $6.50/day | 12 GB |
| text-embedding-3-small | $10 | $1.00/day | 6 GB |
| embed-v4 | $5.50 | $0.55/day | 4 GB |
| voyage-3 | $60 | $6.00/day | 4 GB |
| nomic-embed-text (self-hosted) | $0* | $0* | 3 GB |
| bge-large (self-hosted) | $0* | $0* | 4 GB |
*Self-hosted: GPU instance cost ($0.50-$1.50/hr) amortized across all workloads.
The evaluation pipeline
Don't trust benchmarks — run your own evaluation. Here's the framework:
import numpy as np
from dataclasses import dataclass
from typing import Protocol
class EmbeddingProvider(Protocol):
async def embed(self, texts: list[str]) -> list[list[float]]: ...
@dataclass
class EvalResult:
model: str
recall_at_10: float
mrr: float
latency_p50_ms: float
latency_p99_ms: float
cost_per_1k_embeddings: float
async def evaluate_model(
provider: EmbeddingProvider,
queries: list[str],
relevant_docs: dict[str, list[str]], # query -> relevant doc IDs
corpus_embeddings: np.ndarray,
corpus_ids: list[str],
) -> EvalResult:
"""Evaluate an embedding model on your domain-specific data."""
recalls = []
mrrs = []
for query in queries:
query_embedding = await provider.embed([query])
query_vec = np.array(query_embedding[0])
# Cosine similarity
similarities = np.dot(corpus_embeddings, query_vec) / (
np.linalg.norm(corpus_embeddings, axis=1) * np.linalg.norm(query_vec)
)
top_10_indices = np.argsort(similarities)[-10:][::-1]
top_10_ids = [corpus_ids[i] for i in top_10_indices]
# Recall@10
relevant = set(relevant_docs[query])
retrieved = set(top_10_ids)
recalls.append(len(relevant & retrieved) / len(relevant))
# MRR
for rank, doc_id in enumerate(top_10_ids, 1):
if doc_id in relevant:
mrrs.append(1.0 / rank)
break
else:
mrrs.append(0.0)
return EvalResult(
model=provider.name,
recall_at_10=np.mean(recalls),
mrr=np.mean(mrrs),
latency_p50_ms=provider.get_latency_p50(),
latency_p99_ms=provider.get_latency_p99(),
cost_per_1k_embeddings=provider.cost_per_1k,
)
Running evaluations with multiple providers
import { OpenAI } from "openai";
interface EmbeddingConfig {
name: string;
provider: "openai" | "cohere" | "voyage" | "local";
model: string;
dimensions: number;
batchSize: number;
}
const CONFIGS: EmbeddingConfig[] = [
{
name: "openai-large",
provider: "openai",
model: "text-embedding-3-large",
dimensions: 3072,
batchSize: 2048,
},
{
name: "openai-small",
provider: "openai",
model: "text-embedding-3-small",
dimensions: 1536,
batchSize: 2048,
},
];
async function embedBatch(
config: EmbeddingConfig,
texts: string[],
): Promise<number[][]> {
const client = new OpenAI();
const batches: number[][][] = [];
for (let i = 0; i < texts.length; i += config.batchSize) {
const batch = texts.slice(i, i + config.batchSize);
const response = await client.embeddings.create({
model: config.model,
input: batch,
dimensions: config.dimensions,
});
batches.push(response.data.map((d) => d.embedding));
}
return batches.flat();
}
The dimensionality trade-off
Higher dimensions capture more nuance but cost more to store and search. OpenAI's Matryoshka embeddings let you truncate dimensions with graceful quality degradation:
| Dimensions | Recall@10 | Storage per 1M | Search Latency |
|---|---|---|---|
| 3072 | 0.94 | 12 GB | 45ms |
| 1536 | 0.92 | 6 GB | 28ms |
| 768 | 0.89 | 3 GB | 18ms |
| 256 | 0.83 | 1 GB | 8ms |
For most production systems, 1536 dimensions hits the sweet spot — 2% quality loss for 50% storage and latency savings.
Decision framework
Choose text-embedding-3-large when:
- Maximum retrieval quality is non-negotiable
- You're indexing <1M documents
- Query latency budget is >100ms
Choose text-embedding-3-small when:
- Cost is a constraint and quality at 1536d is acceptable
- High query volume (>500K/day)
- You need the OpenAI ecosystem integration
Choose Voyage-3 when:
- Domain-specific content (legal, medical, scientific)
- Documents are long (>2000 tokens)
- You need best-in-class on specialized benchmarks
Choose self-hosted (nomic/bge) when:
- Latency is critical (<20ms query embedding)
- Data can't leave your infrastructure
- You have GPU capacity available
- Cost must be near-zero at scale
Key takeaways
- Run domain-specific evaluations — MTEB rankings don't predict your production performance
- Latency matters more than you think — 50ms embedding time adds up when you're doing re-ranking
- Matryoshka truncation is an underused lever — test reduced dimensions before paying for full
- Self-hosted models are viable for production if you have GPU infrastructure
- The best model is the one you can afford to run at your query volume
Start with text-embedding-3-small for prototyping, benchmark against your domain data, then scale to what the quality metrics demand.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.