Google's technologies and products now support interactions in over 300 languages, covering more than 7 billion people globally. This expansion addresses the historical underrepresentation of many languages in the digital world, a goal Google has pursued since launching Google Translate in 2006.
Google is moving beyond text-based translation to native audio intelligence, training models like Gemini to process audio directly. This approach aims to capture elements of human communication such as tone, pacing, emotion, and context, which are often lost in traditional multi-step speech recognition systems.
This direct audio processing allows for better handling of natural speech patterns, including laughter, hesitations, and code-switching between multiple languages within a single sentence.
Gemini 3.5 Live Translate now provides real-time spoken translation across 70 languages and over 2,000 language pairs. This tool is designed to naturally capture code-switching and emotional cues during conversations.
Gemini 3.5 Transcribe is Google's most precise speech-to-text model to date. It converts raw audio into formatted text, functioning effectively in noisy environments and with complex jargon. This technology also powers features like Rambler on Android Gboard, which refines spoken input by removing filler words, correcting grammar, and allowing voice-based editing and language switching.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Google is advancing its AI-powered language technologies, including real-time spoken translation across 70 languages and improved speech-to-text models. These developments aim to better capture the nuances of human communication and support a wider range of languages.