Meta's Superintelligence Labs has introduced Muse Voice Transcribe, its first real-time audio perception model. This new model delivers multilingual, streaming transcription capabilities to Meta AI for Mac, Muse Code, and developers through the Meta Model API. It integrates streaming automatic speech recognition with speaker diarization and endpointing, allowing for transcription as speech occurs, separation of over 20 speakers, and identification of when a speaker finishes talking without post-processing.
Muse Voice Transcribe was trained across more than 70 languages, with 25 validated at launch. It supports audio longer than an hour and features native code-switching within or between sentences. The model also incorporates language, keyword, and context biasing to improve recognition accuracy. Unlike traditional models, it uses an "adaptive delay" mechanism, adjusting the listening duration before committing each word to balance speed and accuracy.
On the Artificial Analysis's AA-WER Streaming speech-to-text accuracy benchmark for English, Muse Voice Transcribe achieved a 3.1% word error rate, outperforming other comparable real-time speech processing models from companies like OpenAI and Google. The model is accessible through the Meta Model API, priced at $3.00 per 1,000 audio minutes. Meta has stated that, unlike its Muse Glimmer models, the open weights for Muse Voice Transcribe will not be made available.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Meta's Superintelligence Labs released Muse Voice Transcribe, a new real-time speech recognition model that shows a 3.1% word error rate on the AA-WER Streaming benchmark for English, surpassing models from OpenAI and Google. This model supports over 70 languages, distinguishes more than 20 speakers, and is available via Meta Model API, Meta AI for Mac, and Muse Code.
Meta released Muse Voice Transcribe, a real-time audio perception model offering multilingual, streaming transcription for Meta AI on Mac, Muse Code, and developers via the Meta Model API. This development provides advanced dictation capabilities and speaker diarization, impacting applications requiring accurate and immediate speech-to-text conversion.