What does an embedding model actually do?
An embedding model maps text into a fixed-length numerical vector such that semantically similar text lands close together in that vector space — "cardiac arrest" and "heart attack" end up near each other, while "cardiac arrest" and "car arrest" don't, despite sharing letters. In a retrieval-augmented generation pipeline, every document chunk is embedded once at ingest time and stored in a vector database; every user query is embedded with the same model at query time, and nearest-neighbour search returns the chunks whose vectors are closest to the query's vector.
Most RAG retrieval is asymmetric: a short question needs to retrieve a longer passage that answers it. Several embedding models handle this with separate query/document input conventions rather than treating both the same way.
Four things to trade off when picking a model
The MTEB leaderboard is a shortlist tool, not a verdict — its own authors note no single method dominates across tasks, and BEIR has shown dense retrievers losing to plain keyword search (BM25) out of domain. Four axes decide the real-world choice:
| Axis | Why it matters | Risk if ignored |
|---|---|---|
| Retrieval quality | MTEB/BEIR scores are a shortlist signal, not a verdict for your corpus | Leaderboard winner can lose on your own eval set |
| Dimensions | Directly drives vector database storage and query cost | A 3072-dim model costs 3x the storage of a 1024-dim one at the same corpus size |
| Max context | Caps how big a chunk can be before silent truncation | 512-token models truncate long chunks without erroring |
| Price / self-hosting | API models bill per token; open models need GPU ops | Cheaper per-call can still cost more at scale than self-hosting |
Get the weekly AI engineering brief
RAG, agents, evals, and the tools worth using — one practical email a week. Plus the free roadmap PDF.
By subscribing you agree to receive emails from AI Engineer Insights. Unsubscribe anytime. See our Privacy Policy.
2026 shortlist: API and open-weight models
Filter the universe by whether you need to self-host, need multilingual or multimodal input, or just want a strong English-only default with zero ops work (RAG Stack Guide — Best Embedding Model for RAG 2026; How to Choose an Embedding Model in 2026):
| Model | Type | Dimensions | Max context | Notes |
|---|---|---|---|---|
| OpenAI text-embedding-3-large | API | 3072 (reducible) | 8,192 tokens | Strong general-purpose default, widely documented |
| Voyage voyage-3-large | API | 1024 (256-2048) | 32,000 tokens | Long-context, asymmetric query/document prefixes |
| Cohere Embed v4 | API | 256-1536 | ~128,000 tokens | Multimodal (text + images), int8/binary output |
| Google Gemini Embedding | API | 128-3072 | 2,048 tokens | 100+ languages, task_type for query vs document |
| BAAI BGE-M3 | Open (MIT) | 1024 | 8,192 tokens | Dense + sparse + multi-vector in one self-hosted model |
| Qwen3-Embedding-8B | Open (Apache-2.0) | 32-4096 | 32,000 tokens | Strong multilingual MTEB scores, self-hosted |
Prices and exact dimension options change often — treat this table as a shortlist to validate, not a final answer, and check each vendor's current docs before committing (TeachYou.ai — How to Choose an Embedding Model for RAG in 2026; dev.to — How to Choose an Embedding Model in 2026).
How to validate the choice on your own corpus
Before committing, hand-label 30-100 of your own queries with the chunk(s) that should be retrieved, embed your corpus with two or three candidate models, and measure recall@k or context recall for each. This takes an afternoon and routinely overturns the public leaderboard order, because a leaderboard averages across someone else's tasks (dev.to — the blunt rule; Medium — How to Choose a Vector Embedding Model for RAG). See our guide to RAG evaluation metrics for how context precision and context recall are computed.
- • Ingest and query must use the same model. Vectors from two different models are not comparable in the same space — mixing them silently breaks retrieval.
- • Use the model's input_type/task_type convention consistently. Query vs document prefixes exist precisely because RAG retrieval is asymmetric.
- • Test lower dimensions before paying for the largest option. Several current models support Matryoshka-style dimension reduction (e.g. 256 or 512 instead of 3072) at a small, measurable quality cost and a large storage saving.
- • Re-embedding the whole corpus is required to switch models — budget for that migration cost before swapping mid-project, not after.
Frequently Asked Questions
What is an embedding model in RAG?
An embedding model turns a chunk of text (or a query) into a fixed-length numerical vector such that semantically similar text ends up close together in vector space. In a RAG pipeline it runs at two points — once when documents are ingested into the vector database, and again on every user query — and the quality of that vector is what the retrieval step searches over, so a weak embedding model caps retrieval quality before generation ever runs.
What is the best embedding model for RAG?
There is no single best model — it depends on your corpus, language, chunk length, and budget. A reasonable 2026 shortlist is OpenAI text-embedding-3-large or Voyage voyage-3-large for English-heavy API use, Cohere Embed v4 or Gemini Embedding for multilingual or multimodal needs, and BGE-M3 or Qwen3-Embedding for self-hosted open-weight deployments. Treat the shortlist as a starting point and validate on your own labeled queries before committing.
Do embedding dimensions matter for RAG?
Yes — dimensions drive vector database storage and query cost directly, since every vector the database stores and searches is that many floats. Higher-dimensional embeddings are not automatically better retrieval; several current models support dimension reduction (e.g. via Matryoshka representation learning) so you can test lower dimensions against your own eval set before paying for the largest option.
Should ingest and query use the same embedding model?
Always. Vectors produced by two different embedding models are not comparable in the same vector space, so mixing models between ingestion and query silently breaks retrieval. The same applies to the input_type or task_type convention some models use to distinguish a query from a document — use it consistently or skip it consistently.
API embedding model or self-hosted?
API models (OpenAI, Voyage, Cohere, Gemini) need no infrastructure and get frequent updates, at the cost of per-token pricing and an external dependency on every ingest and query call. Self-hosted open-weight models (BGE-M3, Qwen3-Embedding) avoid that recurring cost and keep data in-house, but require GPU capacity and ops work. Teams with strict data-residency requirements or very high query volume are the ones where self-hosting usually pays off.
How do I evaluate embedding models for my own RAG pipeline?
Hold chunking and the vector database constant, embed the same corpus with two or three candidate models, and measure context precision and context recall against a hand-labeled set of 30-100 of your own queries. This routinely overturns the public MTEB leaderboard order because leaderboards average across tasks that are not yours. See our guide to RAG evaluation metrics for how those scores are computed.
References
- • TeachYou.ai — How to Choose an Embedding Model for RAG in 2026
- • RAG Stack Guide — Best Embedding Model for RAG: How to Choose in 2026
- • dev.to — How to Choose an Embedding Model in 2026 (RAG & Semantic Search)
- • Medium — How to Choose a Vector Embedding Model for RAG: A Practical Guide
- • heycc.cn — How to Choose an Embedding Model in 2026