← All stories
● Covered by 1 source · 1 reportMedium impact1 positive

Mercury 2.5 LLM Released with 40% Intelligence Increase and Enhanced Performance

🔄 Updated 12h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Mercury 2.5 offers a 40% intelligence increase over Mercury 2.
  • It maintains low-latency and low-cost serving profiles.
  • The model achieves 1,107 tokens per second on NVIDIA GPUs.
  • It supports 260K tokens context and includes tunable reasoning.

Introduction of Mercury 2.5

Mercury 2.5, the latest production model, has been released, marking a significant improvement in quality compared to Mercury 2. This new version maintains the same low-latency and low-cost serving profile as its predecessor.

Performance and Capabilities

Mercury 2.5 demonstrates a 40% increase in intelligence over Mercury 2, with capabilities comparable to other cost-optimized frontier models. It processes 1,107 tokens per second on NVIDIA GPUs and supports a context window of 260K tokens. The model is priced at $0.20 per million input and $0.75 per million output, with an introductory offer of 80% off.

New capabilities include tunable reasoning, parallel tool calls, and schema-aligned JSON output. These features enhance the model's flexibility and integration into various applications.

Impact in Production Environments

Since its launch, Mercury 2 has been adopted by thousands of developers and dozens of enterprises, with usage growing significantly. Mercury 2.5 is designed to further support latency-sensitive workloads in areas like search, voice, and coding products.

In search applications, Mercury 2.5 can manage multiple model calls for tasks such as query rewriting, result reranking, and summarization, ensuring fast user interactions. For voice agents, the model helps reduce response latency, as demonstrated by OpenCall, which saw median model response latency drop to approximately 170 milliseconds.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~8 min · 6 stories · Sep 08

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

The Mercury 2.5 large language model has been released, offering a 40% increase in intelligence over its predecessor, Mercury 2, while maintaining low latency and cost. This update provides improved capabilities for applications requiring fast, efficient AI processing, such as search agents and voice applications.