Grok has announced the release of Grok Voice Transcribe 2.0, an updated speech-to-text model. This new version is reported to be twice as accurate as Grok Voice Transcribe 1.0, maintaining the same pricing structure. The model is built on the audio foundation model that powers Grok Voice, which is currently used for customer support calls, video narration transcription, and voice agents in products like Tesla vehicles.
Grok Voice Transcribe 2.0 is designed to handle complex real-world audio, which often includes background noise, multiple speakers, accents, and spoken credentials like phone numbers or email addresses. The model was trained on a diverse dataset of live, noisy, multilingual audio. This training approach aims to provide more accurate transcriptions compared to models optimized for clean, single-speaker audio.
On the public Artificial Analysis leaderboard, Grok Voice Transcribe 2.0 holds the top position for accuracy among 32 streaming models. Internally, Grok measured improvements across four production traffic sets: telephony audio from customer support, conversations with Grok, spoken credentials, and short multilingual voice commands. The new version improved on Grok Voice Transcribe 1.0 in all four categories, leading all tested models in telephony transcription.
A key improvement in Grok Voice Transcribe 2.0 is its enhanced multilingual accuracy. The model can transcribe dozens of languages, automatically detect the language, and follow language switches within a single recording pass. For short phrases, such as in-car commands, the word error rate decreased from 20.6% to 6.8%.
Existing Speech-to-Text API integrations will automatically benefit from the accuracy improvements of Grok Voice Transcribe 2.0 without requiring code changes. The model also supports advanced configuration and controls. This update is significant for industries relying on accurate transcription, such as customer service, content creation, and automotive voice assistants, by providing more reliable speech-to-text capabilities in challenging audio environments.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Grok has released Voice Transcribe 2.0, a new speech-to-text model that is twice as accurate as its predecessor and ranks first on a public streaming model leaderboard. This update improves transcription accuracy for real-world audio, including noisy and multilingual environments, which is significant for applications like customer support and voice assistants.