← All stories
● Covered by 4 sources · 5 reportsMedium impact5 neutral

Meta Launches Muse Voice Transcribe for Real-Time Multilingual Speech Recognition

🔄 Updated 1d ago — new reporting from VentureBeat
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Meta launched Muse Voice Transcribe, a real-time audio perception model.
  • It provides multilingual, streaming transcription with speaker diarization.
  • The model supports over 70 languages and distinguishes more than 20 speakers.
  • Available via Meta Model API, Meta AI for Mac, and Muse Code.
  • Achieved a 3.1% word error rate on the AA-WER Streaming benchmark for English.
  • Muse Voice Transcribe is the first real-time audio model from Meta Superintelligence Lab (MSI).
  • The model was released on September 1, 2026.
  • Muse Voice Transcribe costs $0.18 per hour of audio processing.
  • The model offers endpoint detection.
  • It supports long audio exceeding an hour.
  • It supports seamless multilingual code-switching.
  • It supports language and keyword biasing.
  • Diarization does not require a separate post-processing pipeline.
  • 25 languages were extensively validated for the initial release.

Real-Time Speech Recognition for Multiple Platforms

Meta's Superintelligence Labs has introduced Muse Voice Transcribe, its first real-time audio perception model. This new model delivers multilingual, streaming transcription capabilities to Meta AI for Mac, Muse Code, and developers through the Meta Model API. It integrates streaming automatic speech recognition with speaker diarization and endpointing, allowing for transcription as speech occurs, separation of over 20 speakers, and identification of when a speaker finishes talking without post-processing.

Multilingual Support and Advanced Features

Muse Voice Transcribe was trained across more than 70 languages, with 25 validated at launch. It supports audio longer than an hour and features native code-switching within or between sentences. The model also incorporates language, keyword, and context biasing to improve recognition accuracy. Unlike traditional models, it uses an "adaptive delay" mechanism, adjusting the listening duration before committing each word to balance speed and accuracy.

Performance Benchmarks and Availability

On the Artificial Analysis's AA-WER Streaming speech-to-text accuracy benchmark for English, Muse Voice Transcribe achieved a 3.1% word error rate, outperforming other comparable real-time speech processing models from companies like OpenAI and Google. The model is accessible through the Meta Model API, priced at $3.00 per 1,000 audio minutes. Meta has stated that, unlike its Muse Glimmer models, the open weights for Muse Voice Transcribe will not be made available.

Updates

🕒 2026-09-03 · new reporting from VentureBeat
  • Muse Voice Transcribe costs $0.18 per hour of audio processing.
  • The model offers endpoint detection.
  • It supports long audio exceeding an hour.
  • It supports seamless multilingual code-switching.
  • It supports language and keyword biasing.
  • Diarization does not require a separate post-processing pipeline.
  • 25 languages were extensively validated for the initial release.
🕒 2026-09-02 · new reporting from Engadget
  • Muse Voice Transcribe is the first real-time audio model from Meta Superintelligence Lab (MSI).
  • The model was released on September 1, 2026.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~19 min · 16 stories · Sep 04

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Primary sources

GitHub QwenAudio/Fun-ASR

How outlets covered it

Microsoft AI launched MAI-Transcribe-2, a new speech-recognition model priced at $0.10 per hour of audio, a 72% reduction from its previous version. This release expands language support to 60 languages and includes features like speaker diarization and word-level timestamps, positioning Microsoft to compete in AI transcription without relying on OpenAI's technology.

Meta Superintelligence Labs introduced Muse Voice Transcribe, a new audio perception model offering real-time speech-to-text, endpoint detection, and speaker diarization for over 20 speakers. The service is priced at $0.18 per hour of audio processing, aiming to compete in the enterprise market for meeting systems and call analytics.

Meta has launched Muse Voice Transcribe, a real-time audio model capable of transcribing and dictating for over 20 speakers and handling multiple languages, including code-switching. This model is available through the Meta AI Mac app, Muse Code, and Meta's Model API, offering advanced speech-to-text capabilities to developers and users.

Meta's Superintelligence Labs released Muse Voice Transcribe, a new real-time speech recognition model that shows a 3.1% word error rate on the AA-WER Streaming benchmark for English, surpassing models from OpenAI and Google. This model supports over 70 languages, distinguishes more than 20 speakers, and is available via Meta Model API, Meta AI for Mac, and Muse Code.

Meta released Muse Voice Transcribe, a real-time audio perception model offering multilingual, streaming transcription for Meta AI on Mac, Muse Code, and developers via the Meta Model API. This development provides advanced dictation capabilities and speaker diarization, impacting applications requiring accurate and immediate speech-to-text conversion.