Mercury 2.5, the latest production model, has been released, marking a significant improvement in quality compared to Mercury 2. This new version maintains the same low-latency and low-cost serving profile as its predecessor.
Mercury 2.5 demonstrates a 40% increase in intelligence over Mercury 2, with capabilities comparable to other cost-optimized frontier models. It processes 1,107 tokens per second on NVIDIA GPUs and supports a context window of 260K tokens. The model is priced at $0.20 per million input and $0.75 per million output, with an introductory offer of 80% off.
New capabilities include tunable reasoning, parallel tool calls, and schema-aligned JSON output. These features enhance the model's flexibility and integration into various applications.
Since its launch, Mercury 2 has been adopted by thousands of developers and dozens of enterprises, with usage growing significantly. Mercury 2.5 is designed to further support latency-sensitive workloads in areas like search, voice, and coding products.
In search applications, Mercury 2.5 can manage multiple model calls for tasks such as query rewriting, result reranking, and summarization, ensuring fast user interactions. For voice agents, the model helps reduce response latency, as demonstrated by OpenCall, which saw median model response latency drop to approximately 170 milliseconds.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
The Mercury 2.5 large language model has been released, offering a 40% increase in intelligence over its predecessor, Mercury 2, while maintaining low latency and cost. This update provides improved capabilities for applications requiring fast, efficient AI processing, such as search agents and voice applications.