← All stories
● Covered by 2 sources · 2 reportsMedium impact2 neutral

Meta Launches Muse Voice Transcribe for Real-Time Multilingual Speech Recognition

🔄 Updated 1h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Meta launched Muse Voice Transcribe, a real-time audio perception model.
  • It provides multilingual, streaming transcription with speaker diarization.
  • The model supports over 70 languages and distinguishes more than 20 speakers.
  • Available via Meta Model API, Meta AI for Mac, and Muse Code.
  • Achieved a 3.1% word error rate on the AA-WER Streaming benchmark for English.

Real-Time Speech Recognition for Multiple Platforms

Meta's Superintelligence Labs has introduced Muse Voice Transcribe, its first real-time audio perception model. This new model delivers multilingual, streaming transcription capabilities to Meta AI for Mac, Muse Code, and developers through the Meta Model API. It integrates streaming automatic speech recognition with speaker diarization and endpointing, allowing for transcription as speech occurs, separation of over 20 speakers, and identification of when a speaker finishes talking without post-processing.

Multilingual Support and Advanced Features

Muse Voice Transcribe was trained across more than 70 languages, with 25 validated at launch. It supports audio longer than an hour and features native code-switching within or between sentences. The model also incorporates language, keyword, and context biasing to improve recognition accuracy. Unlike traditional models, it uses an "adaptive delay" mechanism, adjusting the listening duration before committing each word to balance speed and accuracy.

Performance Benchmarks and Availability

On the Artificial Analysis's AA-WER Streaming speech-to-text accuracy benchmark for English, Muse Voice Transcribe achieved a 3.1% word error rate, outperforming other comparable real-time speech processing models from companies like OpenAI and Google. The model is accessible through the Meta Model API, priced at $3.00 per 1,000 audio minutes. Meta has stated that, unlike its Muse Glimmer models, the open weights for Muse Voice Transcribe will not be made available.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~24 min · 20 stories · Sep 01

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

Meta's Superintelligence Labs released Muse Voice Transcribe, a new real-time speech recognition model that shows a 3.1% word error rate on the AA-WER Streaming benchmark for English, surpassing models from OpenAI and Google. This model supports over 70 languages, distinguishes more than 20 speakers, and is available via Meta Model API, Meta AI for Mac, and Muse Code.

Meta released Muse Voice Transcribe, a real-time audio perception model offering multilingual, streaming transcription for Meta AI on Mac, Muse Code, and developers via the Meta Model API. This development provides advanced dictation capabilities and speaker diarization, impacting applications requiring accurate and immediate speech-to-text conversion.