Nemotron-3-Embed-8B
As of 14 Sep 2026, Nemotron-3-Embed-8B is the RightSignal pick for best text embeddings.
NVIDIA's launch materials put Nemotron-3-Embed-8B-BF16 top of the RTEB multilingual board as of 16 July 2026, on a headline 78.46 average NDCG@10 across the 16 public RTEB tasks its model card actually reports, plus 75.45 on MMTEB Retrieval.
Google's Gemini Embedding 2 is natively multimodal, mapping text, images, video, audio and documents into one vector space; it went to public preview on 10 March 2026 and reached general availability on 22 April 2026 through the Gemini API, Vertex AI and the Gemini Enterprise Agent Platform. Output defaults to 3,072 dimensions and is Matryoshka-truncatable from 128 up, with 768, 1,536 and 3,072 the recommended sizes, over an 8,192-token context, at list pricing of about $0.20 per million text tokens. It remains API-only, so it is not an option if you need weights you can host yourself.
Why it’s the challengercurrent reign
held off
reigns
Why Nemotron-3-Embed-8B is the pick
- NVIDIA's launch materials put Nemotron-3-Embed-8B-BF16 top of the RTEB multilingual board as of 16 July 2026, on a headline 78.46 average NDCG@10 across the 16 public RTEB tasks its model card actually reports, plus 75.45 on MMTEB Retrieval.
- The card confirms 34 evaluated languages, a 32,768-token maximum sequence length and 4,096-dimension vectors you can slice to 2,048 or 1,024 and re-normalise, with weights under OpenMDW-1.1 and 1B BF16 and NVFP4 checkpoints for cheaper serving.
- Mind the basis before you quote it: the live RTEB(beta) multilingual board lists 30 tasks over 22 languages, ranks by Borda count of per-task ranks rather than by mean score, and has had its private column temporarily removed since January 2026, so the placing rests on a vendor snapshot of an open-task view.
Evidence
The sources behind this title’s record.
- huggingface.co/nvidia/Nemotron-3-Embed-8B-BF16
- huggingface.co/blog/rteb
- blog.voyageai.com/2026/01/15/voyage-4/
- developers.googleblog.com/en/gemini-embedding-available...
- blog.voyageai.com/2025/01/07/voyage-3-large/
- blog.voyageai.com/2024/09/18/voyage-3/
- developers.openai.com/api/docs/models/text-embedding-3-large
- huggingface.co/BAAI/bge-large-en-v1.5
Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.
Caveats & challengers
- Microsoft's card for harrier-oss-v1-27b claims state-of-the-art on Multilingual MTEB v2 at release, with a 74.3 average across a 94-language family (0.6b 69.0, 270m 66.5), but it gives no task count, no rival scores and no Borda ranking — so treat any head-to-head against Tencent's KaLM-Embedding-Gemma3-12B-2511 as unconfirmed rather than settled.
- That card dates from late March 2026 and the model landed in Foundry in April, so pull the live MTEB table before repeating either number.
- Either way an MTEB v2 task mean and our pick's RTEB NDCG@10 answer different questions: one aggregates all task types, the other is retrieval only.
Google's Gemini Embedding 2 is natively multimodal, mapping text, images, video, audio and documents into one vector space; it went to public preview on 10 March 2026 and reached general availability on 22 April 2026 through the Gemini API, Vertex AI and the Gemini Enterprise Agent Platform. Output defaults to 3,072 dimensions and is Matryoshka-truncatable from 128 up, with 768, 1,536 and 3,072 the recommended sizes, over an 8,192-token context, at list pricing of about $0.20 per million text tokens. It remains API-only, so it is not an option if you need weights you can host yourself.
At a glance
- Title
- Best text embeddings
- Licence
- OpenMDW-1.1
- Pick since
- 16 Jul 2026
- Last reviewed
- 13 Sep 2026
- Title holders to date
- 8
- Official page
- huggingface.co/nvidia/Nemotron-3-Em...
Title history
Every change, on the record.
NVIDIA's launch materials put Nemotron-3-Embed-8B-BF16 top of the RTEB multilingual board as of 16 July 2026, on a headline 78.46 average NDCG@10 across the 16 public RTEB tasks its model card actually reports, plus 75.45 on MMTEB Retrieval.
60 daysas pick
The first production-grade MoE embedding model: 8.20% better general retrieval than both gemini-embedding-001 and Cohere Embed v4, 14.05%…
182 daysas pick
Top of the MTEB Multilingual board continuously since its March experimental launch, and genuinely deployable from GA: 100+ languages, 30…
185 daysas pick
9.74% ahead of OpenAI's v3-large across 100 datasets — and its 512-dimension binary vectors beat OpenAI's 3072-dimension floats outright …
188 daysas pick
Beat text-embedding-3-large by 7.55% on retrieval across domains at 2.2x lower price, with 1024 dimensions instead of 3072 and a 32k cont…
111 daysas pick
OpenAI's answer to the open leaders: a real MTEB jump over ada-002 plus Matryoshka dimension truncation, letting teams cut vector-databas…
237 daysas pick
The first open model you could swap in for ada-002 and simply win: MTEB average 64.23 against 60.99, retrieval 54.29 against 49.25, at 33…
135 daysas pick
as pick