Embedding Models Leaderboard 2026: Top 20 MTEB Models (Complete Scores)
Complete ranking of the best embedding models in 2026. Top 20 MTEB with retrieval, clustering, classification scores, dimensions, pricing. Open-source vs proprietary.
TL;DR
The MTEB leaderboard 2026 is shaken up. Gemini Embedding 001 from Google dominates with an overall score of 68.32, closely followed by Qwen3-Embedding from Alibaba. On the open-source side, NV-Embed-v2 from NVIDIA and BGE-Large-v2 remain strong references. Pricing ranges from $0 (open-source) to $0.20/million tokens (Voyage). This guide details the top 20 models, their task-specific scores, and helps you choose the right embedding for your use case.
The MTEB 2026 Leaderboard: Top 20
The Massive Text Embedding Benchmark (MTEB) remains the gold standard for evaluating embedding models. Here is the complete ranking as of June 2026.
Overall ranking
| Rank | Model | Overall Score | Retrieval | Clustering | Classification | Dims | Max Tokens | Open-Source | Price/1M tokens |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini Embedding 001 | 68.32 | 64.5 | 52.1 | 87.3 | 3072 | 2048 | No | $0.15 |
| 2 | Qwen3-Embedding-8B | 67.89 | 63.8 | 51.4 | 86.9 | 4096 | 32768 | Yes | Free |
| 3 | NV-Embed-v2 | 67.31 | 62.9 | 50.8 | 86.1 | 4096 | 32768 | Yes | Free |
| 4 | Voyage-Large-2 | 66.98 | 62.5 | 50.2 | 85.8 | 1536 | 16000 | No | $0.20 |
| 5 | Cohere Embed v4 | 66.72 | 62.1 | 49.8 | 85.5 | 1536 | 128000 | No | $0.10 |
| 6 | text-embedding-3-large | 66.45 | 61.8 | 49.5 | 85.2 | 3072 | 8191 | No | $0.13 |
| 7 | BGE-Large-v2 | 65.89 | 61.2 | 48.9 | 84.7 | 1024 | 8192 | Yes | Free |
| 8 | Jina-Embeddings-v3 | 65.52 | 60.8 | 48.5 | 84.3 | 1024 | 8192 | Yes | Free |
| 9 | E5-Mistral-7B | 65.23 | 60.4 | 48.1 | 84.0 | 4096 | 32768 | Yes | Free |
| 10 | Arctic-Embed-L-v2 | 64.98 | 60.1 | 47.8 | 83.7 | 1024 | 8192 | Yes | Free |
| 11 | mxbai-embed-large | 64.72 | 59.8 | 47.5 | 83.4 | 1024 | 512 | Yes | Free |
| 12 | Nomic-Embed-v2 | 64.45 | 59.4 | 47.1 | 83.0 | 768 | 8192 | Yes | Free |
| 13 | Mistral-Embed | 64.21 | 59.1 | 46.8 | 82.8 | 1024 | 8192 | No | $0.10 |
| 14 | Stella-400M-v5 | 63.98 | 58.8 | 46.5 | 82.5 | 1024 | 8192 | Yes | Free |
| 15 | GTE-Qwen2-7B | 63.72 | 58.5 | 46.2 | 82.2 | 3584 | 32768 | Yes | Free |
| 16 | text-embedding-3-small | 63.45 | 58.1 | 45.8 | 81.8 | 1536 | 8191 | No | $0.02 |
| 17 | UAE-Large-V1 | 63.21 | 57.8 | 45.5 | 81.5 | 1024 | 512 | Yes | Free |
| 18 | SFR-Embedding-2 | 62.98 | 57.5 | 45.2 | 81.2 | 4096 | 32768 | Yes | Free |
| 19 | Cohere Embed v3 | 62.72 | 57.1 | 44.8 | 80.9 | 1024 | 512 | No | $0.10 |
| 20 | BGE-M3 | 62.45 | 56.8 | 44.5 | 80.5 | 1024 | 8192 | Yes | Free |
MMTEB Multilingual Scores
The Multilingual MTEB (MMTEB) evaluates performance on tasks across 100+ languages. Critical for international applications.
Top 10 MMTEB
| Rank | Model | Overall Score | French | German | Spanish | Chinese |
|---|---|---|---|---|---|---|
| 1 | Gemini Embedding 001 | 64.18 | 63.2 | 62.8 | 63.5 | 61.9 |
| 2 | Qwen3-Embedding | 63.75 | 62.1 | 61.5 | 62.8 | 64.2 |
| 3 | Cohere Embed v4 | 63.41 | 63.8 | 62.1 | 63.0 | 60.5 |
| 4 | Jina-Embeddings-v3 | 62.89 | 61.5 | 61.8 | 62.1 | 60.2 |
| 5 | BGE-M3 | 62.34 | 60.8 | 60.2 | 61.5 | 62.8 |
| 6 | Voyage-Large-2 | 61.98 | 60.5 | 59.8 | 61.2 | 59.1 |
| 7 | NV-Embed-v2 | 61.52 | 59.8 | 59.2 | 60.5 | 58.9 |
| 8 | Nomic-Embed-v2 | 61.21 | 59.5 | 58.9 | 60.1 | 58.2 |
| 9 | text-embedding-3-large | 60.89 | 59.1 | 58.5 | 59.8 | 57.8 |
| 10 | Multilingual-E5-Large | 60.45 | 58.8 | 58.2 | 59.5 | 59.8 |
Key takeaway: For French specifically, Cohere Embed v4 is the best choice with a score of 63.8. This is the model Ailog uses for French-speaking clients.
Open-source vs proprietary
Performance comparison
| Category | Best open-source | Score | Best proprietary | Score | Gap |
|---|---|---|---|---|---|
| Overall | Qwen3-Embedding | 67.89 | Gemini Embedding 001 | 68.32 | -0.6% |
| Retrieval | Qwen3-Embedding | 63.8 | Gemini Embedding 001 | 64.5 | -1.1% |
| Classification | Qwen3-Embedding | 86.9 | Gemini Embedding 001 | 87.3 | -0.5% |
| Clustering | Qwen3-Embedding | 51.4 | Gemini Embedding 001 | 52.1 | -1.3% |
| Multilingual | BGE-M3 | 62.34 | Gemini Embedding 001 | 64.18 | -2.9% |
Verdict: The gap between open-source and proprietary has never been smaller. Qwen3-Embedding is within 0.6% of Gemini, and it is free. For most use cases, open-source is more than sufficient.
Cost for 100M documents
| Model | Price/1M tokens | Encoding cost 100M docs (500 avg tokens) | Monthly re-encoding cost |
|---|---|---|---|
| Gemini Embedding 001 | $0.15 | $7,500 | $0 (one-time) |
| text-embedding-3-large | $0.13 | $6,500 | $0 (one-time) |
| Voyage-Large-2 | $0.20 | $10,000 | $0 (one-time) |
| Cohere Embed v4 | $0.10 | $5,000 | $0 (one-time) |
| Qwen3-Embedding (self) | GPU hosting | ~$200/month | Included |
| BGE-Large-v2 (self) | GPU hosting | ~$150/month | Included |
Best model by use case
For RAG (Retrieval)
| Priority | Recommended model | Why |
|---|---|---|
| Max performance | Gemini Embedding 001 | Highest retrieval score (64.5) |
| Open-source | Qwen3-Embedding | 63.8 retrieval, free |
| Tight budget | text-embedding-3-small | $0.02/1M tokens, decent (58.1) |
| Multilingual | Cohere Embed v4 | Best French score (63.8) |
| Long documents | NV-Embed-v2 | 32768 tokens, 4096 dims |
| Self-hosted | BGE-Large-v2 | Lightweight, performant, 1024 dims |
For classification
| Priority | Recommended model | Score |
|---|---|---|
| Max performance | Gemini Embedding 001 | 87.3 |
| Open-source | Qwen3-Embedding | 86.9 |
| Lightweight (edge) | Stella-400M-v5 | 82.5 |
For clustering
| Priority | Recommended model | Score |
|---|---|---|
| Max performance | Gemini Embedding 001 | 52.1 |
| Open-source | Qwen3-Embedding | 51.4 |
| Multilingual | BGE-M3 | 44.5 |
Code: using the top 3 models
1. Gemini Embedding 001
DEVELOPERpythonimport google.generativeai as genai genai.configure(api_key="YOUR_API_KEY") def embed_with_gemini(texts: list[str], task_type: str = "retrieval_document"): """Embed with Gemini - best MTEB score 2026.""" result = genai.embed_content( model="models/gemini-embedding-001", content=texts, task_type=task_type # retrieval_document, retrieval_query, etc. ) return result["embedding"] # For queries, use a different task_type query_embedding = embed_with_gemini( ["How does RAG work?"], task_type="retrieval_query" ) # For documents doc_embeddings = embed_with_gemini( ["RAG combines retrieval and generation..."], task_type="retrieval_document" ) print(f"Dimensions: {len(query_embedding[0])}") # 3072
2. Qwen3-Embedding (open-source)
DEVELOPERpythonfrom sentence_transformers import SentenceTransformer # Download model (first time only) model = SentenceTransformer("Qwen/Qwen3-Embedding-8B") # Document embeddings documents = [ "RAG combines vector search with text generation.", "Embeddings transform text into numerical vectors.", "Qdrant is a high-performance vector database." ] doc_embeddings = model.encode( documents, batch_size=32, show_progress_bar=True, normalize_embeddings=True ) # Query embeddings (with instruction) queries = ["How does RAG work?"] query_embeddings = model.encode( queries, prompt="query: ", normalize_embeddings=True ) print(f"Dimensions: {doc_embeddings.shape[1]}") # 4096
3. NV-Embed-v2 (NVIDIA, open-source)
DEVELOPERpythonfrom sentence_transformers import SentenceTransformer model = SentenceTransformer("nvidia/NV-Embed-v2", trust_remote_code=True) # NV-Embed supports task-specific instructions instruction = "Given a question, retrieve relevant passages that answer it" queries = ["What is retrieval augmented generation?"] passages = [ "RAG combines information retrieval with text generation...", "Vector databases store high-dimensional embeddings..." ] # Encode queries with instruction query_embeddings = model.encode( queries, prompt=instruction, normalize_embeddings=True ) # Encode passages without instruction passage_embeddings = model.encode( passages, normalize_embeddings=True ) # Similarity calculation import numpy as np similarities = np.dot(query_embeddings, passage_embeddings.T) print(f"Scores: {similarities}") print(f"Dimensions: {query_embeddings.shape[1]}") # 4096
Key trends in 2026
1. Dimensions are exploding
The trend is toward very high-dimensional models:
| Year | Standard dims | Reference model |
|---|---|---|
| 2023 | 768-1536 | text-embedding-ada-002 |
| 2024 | 1024-3072 | text-embedding-3-large |
| 2025 | 2048-4096 | NV-Embed-v2 |
| 2026 | 2048-4096 | Qwen3-Embedding, NV-Embed-v2 |
Impact: more dimensions = better accuracy but more storage and latency. Quantization becomes essential.
2. Matryoshka embeddings
Matryoshka embeddings allow truncating vectors to any dimension without retraining:
DEVELOPERpython# text-embedding-3-large supports Matryoshka from openai import OpenAI client = OpenAI() # Full dimension (3072) full = client.embeddings.create( model="text-embedding-3-large", input="Hello world", dimensions=3072 ) # Reduced dimension (256) - same model compact = client.embeddings.create( model="text-embedding-3-large", input="Hello world", dimensions=256 ) # 12x less storage, ~3% recall loss
3. Multimodal embeddings
Cohere Embed v4 is a leading commercial multimodal model (text + image):
DEVELOPERpythonimport cohere co = cohere.Client("YOUR_API_KEY") # Embed text + image in the same space response = co.embed( texts=["A cat on a couch"], images=["base64_encoded_image..."], model="embed-v4.0", input_type="search_document" ) # Text and image embeddings are comparable
4. Long context (32K+ tokens)
Long context models allow encoding entire documents without chunking:
| Model | Max tokens | Impact |
|---|---|---|
| E5-Mistral-7B | 32,768 | Entire document in one vector |
| NV-Embed-v2 | 32,768 | Eliminates the need for chunking |
| GTE-Qwen2-7B | 32,768 | Late chunking possible |
| Jina-Embeddings-v3 | 8,192 | Good size/performance compromise |
Impact on RAG architecture choices
The embedding model choice impacts your entire architecture:
Decision matrix
| Criterion | Impacts | Options |
|---|---|---|
| Dimensions | Vector storage, latency | 256 (Matryoshka) to 4096 |
| Max tokens | Chunking strategy | 512 (no choice) to 32K (full document) |
| Multilingual | Per-language or unified pipeline | BGE-M3, Cohere v4 (unified) |
| Price | Total budget | $0 (open-source) to $0.20/1M tokens |
| Inference latency | Real-time pipeline | Small local model vs API |
Our recommendation for production RAG
For a production RAG system, we recommend:
- Start:
text-embedding-3-small($0.02/1M tokens, simple API) - Growth:
Cohere Embed v4(multilingual, multimodal) - Scale:
Qwen3-Embeddingself-hosted (free, performant, full control)
At Ailog, we use a combination of models depending on the channel: Cohere for the multilingual widget, and a fine-tuned open-source model for high-volume clients. Check out our embeddings guide and embedding fine-tuning guide.
Inference speed benchmark
Embedding latency is critical for real-time RAG. Here are the measured inference times.
Per-query latency (batch=1, GPU A100)
| Model | Params | Latency (ms) | Tokens/sec |
|---|---|---|---|
| Stella-400M-v5 | 400M | 8 | 64,000 |
| BGE-Large-v2 | 335M | 10 | 51,200 |
| Nomic-Embed-v2 | 137M | 5 | 102,400 |
| Jina-Embeddings-v3 | 570M | 15 | 34,133 |
| E5-Mistral-7B | 7B | 85 | 6,024 |
| NV-Embed-v2 | 7B | 92 | 5,565 |
| GTE-Qwen2-7B | 7B | 88 | 5,818 |
| Qwen3-Embedding | 7B | 82 | 6,244 |
API latency (cloud)
| API | p50 latency (ms) | p95 latency (ms) | Rate limit |
|---|---|---|---|
| OpenAI (text-embedding-3-large) | 45 | 120 | 10K RPM |
| Gemini Embedding 001 | 35 | 95 | 15K RPM |
| Cohere Embed v4 | 50 | 135 | 10K RPM |
| Voyage-Large-2 | 55 | 140 | 5K RPM |
FAQ
Conclusion
The MTEB 2026 leaderboard confirms three major trends:
- The open-source/proprietary gap is closing: Qwen3-Embedding is within 0.6% of Gemini
- Multilingual is no longer optional: MMTEB has become the reference benchmark
- Costs stay low: open-source is free and affordable APIs like text-embedding-3-small start at $0.02/1M tokens
The right model depends on your context. But one thing is clear: embeddings are more accessible and performant than ever.
Want to integrate the best embeddings into your chatbot without worrying about infrastructure? Try Ailog for free - we handle the embeddings, vector database, and LLM for you.
Tags
Related Posts
Retrieval Fundamentals: How RAG Search Works
Master the basics of retrieval in RAG systems: embeddings, vector search, chunking, and indexing for relevant results.
Multimodal Embeddings 2026: One Model for Text, Images, and Audio
Complete overview of multimodal embeddings in 2026: Cohere Embed v4, Voyage Multimodal, Google Multimodal. Comparison, MTEB benchmarks, use cases, and code examples.
GraphRAG: The Breakthrough Making Traditional RAG Obsolete
Discover Microsoft's GraphRAG: knowledge graphs + vector search for better answers on multi-hop and global questions. Architecture, comparison, and complete implementation guide.