Cohere Releases Embed 5: Specs, Speed & Pricing
Cohere launches Embed 5 in Pro and Fast tiers, supporting direct image embedding and achieving up to 377.3 documents per second.

Stock photo for illustration only, not from the actual event
- Cohere introduces Embed 5 in Pro and Fast tiers, generally available now.
- The Fast model processes 377.3 documents per second compared to Pro's 159.7.
- Embed 5 leads benchmarks on ViDoRe V3 and financial NLP suites.
- Supports dimension truncation and lower-precision outputs to save storage.
Cohere has officially released its latest embedding model, Embed 5, available in both Pro and Fast tiers. The new models are generally available on the Cohere API, Model Vault, Microsoft Foundry, and Amazon SageMaker, with private VPC or on-premise serving supported through vLLM.
The API model identifiers are embed-v5.0-pro and embed-v5.0-fast according to model documentation. Both tiers output dimensions of 2048, 1536, 1024, 768, 512, or 256, with 2048 set as the default. Embeddings are returned as float, int8, or binary formats.
In terms of capability, Embed 5 can embed a page image directly and fuse an image with its metadata into a single vector. This feature is vital for scanned pages, slide decks, schematics, and charts where traditional text extraction typically loses critical information.

Stock photo for illustration only, not from the actual event
Cohere evaluated corpus and query pairings across 40 development datasets, recommending a pattern where developers index with Pro and query with Fast. This configuration scored 98.4 normalized to an all-Pro baseline of 100, provided both sides utilize the identical output dimension.
"Fast processed 377.3 documents per second versus 159.7 for Pro."
Cohere Team
On ViDoRe V3 benchmarks, Embed 5 Pro averaged 85.8—an 8.8-point increase over Embed 4—while Fast averaged 84.5. They outperformed competing models like Voyage 4 Large at 83.7, Gemini Embedding 2 at 83.2, and OpenAI text-embedding-3-large at 75.5. Pro also secured top ranks in financial benchmarks including FinanceBench, FinQA, and ViDoRe V3 Finance.
Modern embedding model development increasingly prioritizes low latency and storage efficiency through Matryoshka representation learning, enabling Retrieval-Augmented Generation (RAG) systems to scale effectively for AI agents. Cohere's split-tier approach highlights a deliberate strategy to decouple heavy indexing costs from the high-frequency query latency demanded by autonomous agent workloads.
Pricing for the Pro tier is set at $0.12 per 1 million text tokens, while the Fast tier costs $0.08. Image inputs are priced at $0.40 per 1 million tokens for both options. By leveraging Matryoshka training and lower-precision outputs, raw storage requirements for large-scale vector databases can be drastically reduced.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment