Cohere has released Embed 5, a new version of its embedding model. This release includes two distinct models: Embed 5 Pro and Embed 5 Fast. The primary design allows teams to use Embed 5 Pro for indexing data and Embed 5 Fast for querying, without needing to create separate indexes.
The company recommends this Pro-for-indexing, Fast-for-querying approach, especially for Retrieval Augmented Generation (RAG) and agent workloads. These applications often involve repeated searches, where latency can accumulate. The Embed 5 Fast model is designed to address this by offering faster and more cost-effective queries.
Embed 5 Fast delivers an average of 2.4 times the document throughput compared to Embed 5 Pro in Cohere's tests. Additionally, Pro costs $0.12 per million tokens, while Fast costs $0.08 per million tokens, representing a 33% cost reduction for queries. This cost and speed optimization is particularly beneficial for systems that query data more frequently than they index it.
Cohere's testing indicates a small trade-off in retrieval quality when using Fast queries against a Pro index. Across 40 datasets, this configuration scored 98.4 relative to a Pro-to-Pro baseline of 100. Using Fast for both indexing and queries resulted in a score of 96.6. Cohere states that no individual dataset showed a major drop in quality when Pro and Fast were combined.
Both Pro and Fast models share the same embedding space and produce compatible vectors with identical dimensions. This allows users to switch between models without re-embedding their data corpus. The models support six vector dimensions from 256 to 2,048, with float32, int8, and binary formats. Cohere suggests 1,024-dimensional int8 for most deployments to balance memory, storage, and retrieval quality, significantly reducing storage requirements compared to float32 vectors.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Cohere launched Embed 5, introducing Pro and Fast models that share an embedding space, allowing users to index data with the Pro model and query with the faster, cheaper Fast model. This setup aims to reduce latency and cost in RAG and agent workloads, with Cohere's tests showing minimal impact on retrieval quality.