Skip to main content
Back to timeline
Cohere LabsSource publication:

Cohere launches Embed 5 with Pro at 85.8 average on ViDoRe V3 and Fast at 2.4x throughput sharing one embedding space

Synopsis

Cohere released the Embed 5 family of embedding models in two tiers, Pro and Fast, which share a single embedding space and support 128K context, text/image/fused inputs, 100+ languages, and Matryoshka outputs from 2048 down to 256 dimensions; Pro averages 85.8 on ViDoRe V3 (an 8.8 gain over Embed 4), scores 80.1 on FinanceBench and 84.8 across the parsed-document suite, while Fast averages 84.5 on ViDoRe V3 with about 2.4x the document throughput of Pro, priced at $0.12 and $0.08 per million tokens respectively.

AI-generated editorial illustration: Introducing Embed 5—A New Family of Frontier Embedding Models

Interpretation

Embed 5 covers different retrieval settings with Pro and Fast tiers that share one embedding space, so documents can be indexed with Pro and queried with either model without rebuilding the index. Compared with rebuilding indexes per model, this decouples the quality tier from the latency tier; the text reports cross-model combinations stay close to same-model baselines, averaging just 1.6% and 2.7% losses for Fast and Pro queries, with no dataset showing a major failure. Based on testing every corpus/query pairing across 40 development datasets spanning text, image, fused, and parsed-document retrieval, reported as mean nDCG@10 normalized to Pro corpus plus Pro query = 100.

Embed 5 Pro reaches the highest averages reported in the text for visually rich enterprise documents and financial document retrieval. On ViDoRe V3 Pro averages 85.8, an 8.8 gain over Embed 4, ahead of Voyage 4 Large (83.7), Gemini Embedding 2 (83.2), and OpenAI text-embedding-3-large (75.5); across three public financial benchmarks Pro ranks first (FinanceBench 80.1, FinQA 90.0, ViDoRe V3 Finance 85.0) with Fast second on each. Based on ViDoRe V3, which samples documents across enterprise domains including financial filings, technical manuals, regulatory material, government reports, textbooks, and lectures, plus the FinanceBench, FinQA, and ViDoRe V3 Finance comparisons.

Embed 5 brings multimodal and multilingual retrieval into one model family and offers compressible vector outputs to control storage cost. On fused text-image corpora Pro averages 82.3 across five datasets (Gemini Embedding 2 at 61.3), and on page-image retrieval Pro averages 77.0 across five financial datasets; for multilingual, Pro averages 77 across German, French, Spanish, Italian, and Russian, about 7 points above Embed 4, with Farsi +13, Telugu +12, and Hindi +12; outputs support float, int8, and binary with Matryoshka dimensions, and the text states raw vector storage across 100 million chunks drops from roughly 819 GB to 3.2 GB. Based on averages over fused and page-image datasets, a per-language score table for ten languages, and the storage arithmetic that a 2048-dimensional float32 vector is 8 KB, a 1024-dimensional int8 vector is 1 KB, and a 256-dimensional binary vector is 32 bytes.

Embed 5 Fast targets latency- and cost-sensitive high-volume retrieval while keeping the same context, multimodal, and multilingual coverage at higher throughput. Fast is priced at $0.08 per million tokens, a third less than Pro, and delivers an average of 2.4x higher document throughput than Pro; on ViDoRe V3 it leads Voyage 4 Nano by almost seven points and Jina Embeddings v5 Text Small, Perplexity, and Microsoft Harrier 0.6B by ten or more, and outperforms Qwen3-VL-Embedding-2B by about 20 points despite being roughly half the size. Based on throughput measurements across context sizes, average-score comparisons on ViDoRe V3 and financial retrieval, and parsed PDF scores of 83.4 versus Gemini Embedding 2 at 80.8 and just below Voyage 4 Large at 83.6.

Perspective

The release targets teams needing enterprise-grade retrieval: Pro for offline indexing and quality-critical retrieval, Fast for interactive search, agent loops, and high-volume query workloads, with 1024-dimensional int8 recommended as the performance-efficiency point. The models support 128K context, text, image, and fused inputs, 100+ languages, outputs from 2048 to 256 dimensions in float, int8, and binary formats, and can be deployed through the Cohere API, Model Vault, Microsoft Foundry, Amazon SageMaker, or self-hosted vLLM, plugging into existing vector databases and frameworks. Cross-model mixing requires both sides to use the same output dimension, and remains compatible with Matryoshka truncation and int8 quantization.

All scores are vendor-run, so readers judging fit for their own corpora would still need to reproduce them on their own data. RCP-nDCG@10 uses a two-stage retrieval setup that reorders a fixed candidate set with similarity scores, so its scores reflect reranking quality rather than first-stage retrieval performance, which the text says is evaluated elsewhere against nDCG and Recall. In the multilingual table, Gemini Embedding 2 or Voyage 4 Large score higher than Embed 5 Pro on some languages (Japanese, Korean, Arabic, Bengali, Telugu, Indonesian, Thai), so the highest average applies to the average over the tested language set. The binary format carries some accuracy tradeoff and is suggested for fast first-pass retrieval before higher-precision reranking. Parse 5, the Compass Cloud private beta, and the October 8 online event are mentioned without results.

Sources