Cohere shipped Embed 5 on Sep 30, 2026, and the most interesting part is not a benchmark score. It is the shape of the family: two tiers, embed-v5.0-pro and embed-v5.0-fast, that share one embedding space. You can index your corpus once with the premium Pro model and serve every query with the cheaper, faster Fast model - Cohere reports this mix keeps 98.4 out of 100 quality, about a 1.6% loss, across 40 development datasets - and if you swap tiers later, you never rebuild the index. Around that core idea, Embed 5 adds a 128K context window, multimodal inputs (text, image, and fused), support for 100+ languages, and Matryoshka dimensions from 2048 down to 256 in float, int8, or binary. Cohere says compression this deep cuts raw vector storage from roughly 819 GB to 3.2 GB per 100 million chunks.
One honest flag before the details: every benchmark number below is Cohere-published, scored on a metric Cohere introduced the very same day. And on Cohere's own multilingual table, Gemini Embedding 2 beats Embed 5 Pro on 9 of 10 non-European languages. This post translates the release into index-rebuild math, storage math, and a clear list of situations where you should not switch.
In short:
- Cohere released Embed 5 on Sep 30, 2026: two tiers, Pro and Fast, living in one shared embedding space.
- The recommended pattern is index with Pro, query with Fast - Cohere reports 98.4/100 quality (about a 1.6% loss), and no re-index when you change tiers later.
- Text pricing is $0.12 per 1M tokens for Pro and $0.08 for Fast; image tokens cost $0.40 per 1M.
- Storage compression runs from 8 KB per vector (2048-d float32) down to 32 bytes (256-d binary) - roughly 819 GB to 3.2 GB per 100M chunks.
- All benchmarks are Cohere-reported on Cohere's own new RCP-nDCG@10 metric, which measures reranking quality rather than first-stage retrieval.
What did Cohere actually ship?
Cohere's announcement post describes a family rather than a single model. Embed 5 Pro is the maximum-quality tier, which Cohere positions for "offline indexing; complex enterprise corpora." Embed 5 Fast is the latency-and-cost tier, aimed at "interactive search; high-volume RAG; agentic retrieval." Both tiers expose the identical capability surface:
- a 128K token context window
- text, image, and fused (text-plus-image) inputs
- 100+ languages
- output dimensions of 2048, 1536, 1024, 768, 512, or 256 - Matryoshka embeddings, so you can truncate to a smaller size without retraining
- float, int8, or binary output types
- self-hosting support, with both models servable via vLLM
The economics of Fast, per Cohere: it "costs a third less than Pro" ($0.08 vs $0.12 per million text tokens) and delivers "an average of 2.4x higher throughput than Pro." Note what Cohere did not publish: absolute latency figures in milliseconds. Only the relative throughput claim exists, so treat "fast" as a relative term until you measure it yourself.
Why does the two-tier split matter? Normally, a quality tier and a cheap tier are two different models with incompatible vectors. A team that wants cheap queries must either index with the weaker model (dragging quality down everywhere) or maintain two full indexes (doubling cost and operational pain). Embed 5's answer is to put both tiers in the same embedding space, which is the next section's story.
Availability, in one line: the models are generally available on the Cohere API and Model Vault, Microsoft Foundry, and Amazon SageMaker, and for private deployments in your own VPC or on-premises, both can be served with vLLM.
The short version: Embed 5 is two tiers with one capability surface, one shared embedding space, and a 128K context window.
Why does the shared embedding space change the index-rebuild math?
Here is the mechanism, in Cohere's own words: "Pro and Fast share a single embedding space, so vectors from either model can be compared directly." Cohere tested every corpus/query pairing across 40 development datasets spanning text, image, fused, and parsed-document retrieval, with scores normalized so an all-Pro system scores 100:
| Corpus + query pairing | Quality (normalized; Cohere reports) |
|---|---|
| Pro corpus + Pro query | 100 |
| Pro corpus + Fast query | 98.4 |
| Fast corpus + Pro query | 97.3 |
| Fast corpus + Fast query | 96.6 |
Cohere's summary of that table: the cross-model combinations "remain close to the same-model baselines (averaging just 1.6% and 2.7% losses for Fast and Pro queries, respectively), with no dataset showing a major failure."
In one sentence: the shared space lets you index once with Pro and swap the query tier later without re-embedding anything.

Why does this change the rebuild math? In a retrieval pipeline, the expensive path is embedding the corpus: every chunk, every document, processed once and stored as vectors. The query path is cheap by comparison - until you remember that agentic retrieval multiplies query calls, sometimes many times per user turn. With Embed 5, you do the expensive work once with Pro, then run every query (and every multiplied agentic call) on Fast at $0.08 per million tokens with 2.4x throughput. And if you later decide queries deserve Pro quality, you flip the query model with zero re-indexing. Our vector database explainer covers what an index rebuild involves in practice.
Two fine-print points matter. First, Cohere says both sides must use the same output dimension - but compatibility also holds with Matryoshka truncation and int8 quantization, so the Pro-index/Fast-query pattern works on compressed indexes too. Second, the shared space only helps inside the Embed 5 family. Embedding vectors are never compatible across vendors, so migrating from any non-Cohere model to Embed 5 still means a full re-embed and index rebuild.
What does Embed 5 cost: tokens and terabytes
The API pricing, as published by Cohere:
| Model | Text (per 1M tokens) | Image inputs (per 1M tokens) |
|---|---|---|
| Embed 5 Pro | $0.12 | $0.40 |
| Embed 5 Fast | $0.08 | $0.40 |
For dedicated capacity, Cohere's Model Vault lists Embed 5 instances (both Fast and Pro, Small and Medium sizes) at $3.00 to $5.00 per hour.
For context, competitor text-embedding prices: Voyage 4 Large is $0.12 per 1M tokens, OpenAI's text-embedding-3-large is $0.13, and Gemini Embedding 2 sits around $0.20 per 1M according to secondary price trackers (Google's own pricing page was not directly checked for this line, so treat that figure as reported, not confirmed). Cheaper floors also exist: Voyage 4 Lite is confirmed at $0.02 per 1M on Voyage's docs, and OpenAI's text-embedding-3-small is reported at $0.02 per 1M. In other words, Embed 5 is not the cheapest option on the market - the pitch rests on quality plus shared-space economics, not headline price.
An illustration, derived purely from the unit prices above: indexing 1 billion tokens once with Pro costs about $120, and querying 50M tokens per month with Fast costs about $4 per month. (All-Pro queries would be $6 per month at the same volume.) The illustration makes the real point visible: the bigger Fast win is throughput and latency on the query path, not the query bill itself.
Now the terabytes. Cohere notes that at enterprise scale, the vector index can cost more to operate than the model that generates it. The compression ladder looks like this:
| Vector format | Size per vector |
|---|---|
| 2048-d float32 | 8 KB |
| 1024-d int8 | 1 KB |
| 256-d binary | 32 bytes |
That is, per Cohere, "a 256x reduction" - across 100 million chunks, raw vector storage drops "from roughly 819 GB to 3.2 GB." Cohere says int8 retains near-full-precision retrieval quality in both tiers, and recommends 1,024-dimensional int8 vectors as "the ideal performance-efficiency point" for most deployments. Binary vectors are positioned as well suited to fast first-pass retrieval before higher-precision reranking. If the cost terms in this section are new to you, our plain-English RAG explainer grounds them in how retrieval pipelines actually spend money.

How much of the benchmark story should you believe?
On ViDoRe V3, a benchmark of visually rich documents across 8 enterprise domains, here is what Cohere reports on its new RCP-nDCG@10 metric:
| Model | ViDoRe V3 score (Cohere reports) |
|---|---|
| Embed 5 Pro | 85.8 |
| Embed 5 Fast | 84.5 |
| Voyage 4 Large | 83.7 |
| Gemini Embedding 2 | 83.2 |
| OpenAI text-embedding-3-large | 75.5 |
Cohere describes that 85.8 as "an impressive 8.8 gain from Embed 4," with Pro leading five of the eight domains outright and tying Voyage 4 Large on energy. On finance benchmarks, Cohere reports FinanceBench at 80.1 (Pro) and 80.0 (Fast), FinQA at 90.0 and 88.8, and ViDoRe V3 Finance at 85.0 and 83.9. On its parsed-PDF suite (service docs, corporate reports, SEC filings, manuals, privacy policies), Cohere reports Pro at 84.8.
Now the caveat, stated plainly. Every number above is scored on RCP-nDCG@10 - "Rubric-Calibrated Preferences nDCG@10" - a methodology Cohere introduced the same day in a companion post, and Embed 5 is the first model family evaluated with it. Cohere's own footnote says the metric "requires evaluating embedding models in a two-stage retrieval setup, using their similarity scores to reorder a fixed candidate set," so the scores "reflect reranking quality rather than first-stage retrieval performance." Cohere itself also warns that these models "may not always look strongest when judged only by legacy nDCG scores." The validation study behind the metric is vendor-run: 46 contracted annotators, 289 contests across 273 queries, and reviewers sided with RCP-nDCG 70% of the time where the two metrics disagreed. Independent replication is still pending.
The practical rule for a buyer: read every table as "Cohere reports," not as ground truth. If your own harness uses MTEB or BEIR-style nDCG, or if you run single-stage retrieval and care about Recall@k, run your own bake-off before believing the ViDoRe V3 story.
When should you not switch to Cohere Embed 5?
The case against switching is real, and some of it comes from Cohere's own data:
- Non-European language retrieval. This is the strongest reason. On Cohere's own multilingual table, Gemini Embedding 2 beats Embed 5 Pro on 9 of the 10 listed non-European languages - Japanese 90 vs 87, Hindi 84 vs 80, Telugu 91 vs 80, with similar gaps elsewhere. Pro leads only Chinese, 82 vs 81, essentially a tie. If your corpus is heavy in Asian or South Asian languages, the "best model" headline does not hold on Cohere's own numbers.
- Single-stage retrieval. The benchmark scores measure reranking quality over a fixed candidate set, not first-stage vector search. If your pipeline retrieves directly from the index with no reranker, test Recall@k on your own corpus before drawing conclusions.
- You already run a non-Cohere index. Migrating means a full re-embed of everything. The no-re-index benefit only covers future Pro/Fast tier swaps inside the Embed 5 family.
- Pure price floor. If "good enough" quality works for you, $0.02 per 1M options exist - Voyage 4 Lite (confirmed on Voyage's docs) and OpenAI's text-embedding-3-small (reported in secondary sources).
- Open-weight preference. Embed 5 supports self-hosting via vLLM, but Cohere's announcement does not state open-weight license terms for downloading the models - verify licensing before assuming parity with something like Voyage 4 Nano, which Voyage's own blog lists as Apache 2.0.
- You hoped 128K context was the upgrade. It is not a differentiator: Cohere's own Embed 4 already had a 128K context window. The upgrade case from Embed 4 rests on quality and the tier split, not context length.
So what should your team actually do?
A sensible adoption plan differs by starting point:
- Building a new index: index with Pro at 1,024-dimensional int8 (Cohere's recommended balance point), query with Fast, and run a bake-off on your own corpus before committing - especially if non-European languages matter to you.
- On Embed 4: the upgrade case rests on the reported +8.8 ViDoRe V3 jump and the new tier split. Plan one re-embed, and gain tier-swap freedom afterward.
- On Voyage, OpenAI, or Gemini: weigh the migration cost (a full re-embed) against the shared-space economics, and only move if the quality delta shows up on your own eval, not on Cohere's.
Whichever bucket you are in, the hygiene rule is the same: a vendor benchmark report should never be the decision. Run retrieval tests on your own data. If you want implementation details - dimensions, output types, and integrations with LangChain, Weaveate, Qdrant, Pinecone, and friends - Cohere's docs are the source of truth.
Bottom line
The genuinely novel, practical parts of Embed 5 are economic: index once with Pro, query cheap with Fast, swap tiers without re-indexing, and compress vectors 256x when storage hurts. The benchmark story is promising but single-vendor - treat ViDoRe V3 as a hypothesis to test on your own corpus, especially for non-European languages or single-stage retrieval. Starting fresh? Try the Pro-index, Fast-query pattern at 1,024-d int8 and evaluate on your data. On Embed 4? The quality jump may justify one re-embed. Elsewhere? Let your own eval decide.
Pricing and plan details are as published by the vendor around Oct 5 2026 and can change - confirm on the official site.
If you are re-planning your retrieval stack, the Toolbit AI tools directory lists vector databases and retrieval tooling that pair with the new models.




