← All stories
● Covered by 1 source · 2 reportsMedium impact1 negative

Ramp Router Launches for Efficient AI Model Routing

New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • 2.75 trillion tokens routed monthly
  • 30% cost reduction with 99.999% routing success rate
  • Supports models from OpenAI, Anthropic, and others

Introduction to Ramp Router

Ramp has introduced the Ramp Router, a new API designed to optimize the usage of AI models for various applications. The Router chooses from over 100 different models based on cost, quality, and availability, streamlining how developers access artificial intelligence in their projects.

Cost and Performance Benefits

The Ramp Router reportedly reduces costs by around 30% while maintaining high performance with an additional 30 milliseconds of latency. Success rate for routes is advertised at 99.999%, ensuring reliable service for users.

Functionality and Features

The router handles various tasks such as caching, compaction, and semantic attribution, applying over 100 optimizations to each request. It consolidates multiple AI models into a single API endpoint, simplifying integration for developers.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

Manifest, a company providing an LLM gateway, has deprecated its LLM router feature, which was designed to reduce inference costs by dynamically selecting models based on request complexity. The company found that prompt complexity is difficult to deduce accurately and that caching is a more effective cost-reduction strategy, leading to the decision to remove the router.

Ramp has launched its Router, an API that routes requests to the most cost-effective AI models while optimizing for performance. It aims to reduce operational costs by approximately 30% and offers a reliable infrastructure for AI use cases across various providers.