← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

Cohere Releases Embed 5 with Pro and Fast Models for Optimized RAG Workloads

🔄 Updated 1d ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Cohere released Embed 5 with Pro and Fast models.
  • Pro for indexing, Fast for querying data in RAG.
  • Fast queries are 2.4x faster and 33% cheaper.
  • Retrieval quality reduction is minimal (98.4% of Pro-Pro baseline).

Cohere Embed 5 Introduction

Cohere has released Embed 5, a new version of its embedding model. This release includes two distinct models: Embed 5 Pro and Embed 5 Fast. The primary design allows teams to use Embed 5 Pro for indexing data and Embed 5 Fast for querying, without needing to create separate indexes.

Optimized for RAG and Agent Workloads

The company recommends this Pro-for-indexing, Fast-for-querying approach, especially for Retrieval Augmented Generation (RAG) and agent workloads. These applications often involve repeated searches, where latency can accumulate. The Embed 5 Fast model is designed to address this by offering faster and more cost-effective queries.

Performance and Cost Benefits

Embed 5 Fast delivers an average of 2.4 times the document throughput compared to Embed 5 Pro in Cohere's tests. Additionally, Pro costs $0.12 per million tokens, while Fast costs $0.08 per million tokens, representing a 33% cost reduction for queries. This cost and speed optimization is particularly beneficial for systems that query data more frequently than they index it.

Minimal Impact on Retrieval Quality

Cohere's testing indicates a small trade-off in retrieval quality when using Fast queries against a Pro index. Across 40 datasets, this configuration scored 98.4 relative to a Pro-to-Pro baseline of 100. Using Fast for both indexing and queries resulted in a score of 96.6. Cohere states that no individual dataset showed a major drop in quality when Pro and Fast were combined.

Shared Embedding Space and Vector Compression

Both Pro and Fast models share the same embedding space and produce compatible vectors with identical dimensions. This allows users to switch between models without re-embedding their data corpus. The models support six vector dimensions from 256 to 2,048, with float32, int8, and binary formats. Cohere suggests 1,024-dimensional int8 for most deployments to balance memory, storage, and retrieval quality, significantly reducing storage requirements compared to float32 vectors.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Cohere launched Embed 5, introducing Pro and Fast models that share an embedding space, allowing users to index data with the Pro model and query with the faster, cheaper Fast model. This setup aims to reduce latency and cost in RAG and agent workloads, with Cohere's tests showing minimal impact on retrieval quality.