← All stories
● Covered by 5 sources · 9 reportsMedium impact6 neutral3 positive

Google DeepMind Introduces Gemini 3.8 Live Models for Voice Agents

🔄 Updated 4d ago — new reporting from Engadget
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Gemini 3.8 Live and 3.8 Live Extended Thinking were introduced on September 17, 2026.
  • The models enhance near real-time reasoning for voice agents.
  • Gemini 3.8 Live is for scale and cost efficiency, with conversational intelligence and visual grounding.
  • Gemini 3.8 Live Extended Thinking handles high-complexity tasks and multi-step reasoning.
  • The models are available through the Gemini API and Google AI Studio.
  • Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS are new text-to-speech models.
  • Gemini 3.8 Flash TTS is for creative direction and character design.
  • Gemini 3.8 Flash-Lite TTS is for high-volume, cost-efficient scale.
  • The new models are part of the Gemini Audio family.
  • Gemini 3.5 Live Translate and 3.5 Transcribe are also part of the Gemini Audio family.
  • Gemini 3.8 Flash TTS and Flash-Lite TTS allow custom voice creation.
  • Developers can describe or record a voice to save and reuse it.
  • Voice replication uses a new Voices endpoint (POST /v1beta/voices).
  • Two clean recordings (10-30 seconds) from the same speaker are required.
  • A separate consent recording is needed for voice replication.
  • Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise.
  • Live Avatar was first previewed at Google Cloud Next 2026.
  • Live Avatar generates video avatars with synchronized lip-syncing.
  • Custom avatar feature is available via allowlist only.
  • Live Avatar brings interactive visual presence to conversational video agents across web, mobile, and interactive kiosks.
  • Live Avatar supports 97 languages with consistent video fidelity.
  • Live Avatar can display information on-screen while speaking.
  • Live Avatar offers a library of preset avatars.
  • Live Avatar integrates near real-time visual personas with live dialogue capabilities.
  • Live Avatar allows AI agents to communicate using facial expressions and dynamic visuals.
  • Live Avatar expands application in customer service and sales support.
  • Live Avatar offers both realistic and cartoon-like avatars.
  • Live Avatar features natural expressions and fluid turn-taking.
  • ChatGPT Voice launched in 2022.
  • ChatGPT Voice was the first real-time voice chat feature.
  • ChatGPT Voice uses more natural conversational markers.
  • ChatGPT Voice is more likely to search the internet for current information.
  • Gemini Live has a flatter delivery.
  • Gemini Live relies more on its existing knowledge base.

New Models for Voice AI

Google DeepMind introduced two new AI models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, on September 17, 2026. These models are designed to advance near real-time reasoning for voice agents, making AI conversations more intuitive and intelligent.

Model Capabilities

Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is designed for high-complexity tasks, offering increased intelligence and multi-step reasoning. It achieved the top spot on Artificial Analysis' Speech to Speech Quality Index with 82.6 and leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark.

Developer and Enterprise Tools

These models provide developers and enterprises with building blocks for reliable, production-ready voice agents. They also aim to make speaking with Gemini more fluid and collaborative across the Gemini app, Google Workspace, and Search, enabling users to tackle complex tasks using voice.

Comparison with Other Models

The Gemini 3.8 Live models approach real-time voice agent functionality differently from other offerings. Gemini 3.8 Live Extended Thinking keeps reasoning within the voice model, allowing it to continue speaking while executing asynchronous tool calls. This contrasts with approaches that separate real-time conversation from backend reasoning, pushing more orchestration to the application layer.

Updates

🕒 2026-09-27 · new reporting from Engadget
  • ChatGPT Voice launched in 2022.
  • ChatGPT Voice was the first real-time voice chat feature.
  • ChatGPT Voice uses more natural conversational markers.
  • ChatGPT Voice is more likely to search the internet for current information.
  • Gemini Live has a flatter delivery.
  • Gemini Live relies more on its existing knowledge base.
🕒 2026-09-25 · new reporting from Engadget
  • Live Avatar integrates near real-time visual personas with live dialogue capabilities.
  • Live Avatar allows AI agents to communicate using facial expressions and dynamic visuals.
  • Live Avatar expands application in customer service and sales support.
  • Live Avatar offers both realistic and cartoon-like avatars.
  • Live Avatar features natural expressions and fluid turn-taking.
🕒 2026-09-24 · new reporting from The Verge
  • Live Avatar supports 97 languages with consistent video fidelity.
  • Live Avatar can display information on-screen while speaking.
  • Live Avatar offers a library of preset avatars.
🕒 2026-09-24 · new reporting from Google Cloud Blog, Google DeepMind
  • Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise.
  • Live Avatar was first previewed at Google Cloud Next 2026.
  • Live Avatar generates video avatars with synchronized lip-syncing.
  • Custom avatar feature is available via allowlist only.
  • Live Avatar brings interactive visual presence to conversational video agents across web, mobile, and interactive kiosks.
🕒 2026-09-24 · new reporting from The New Stack
  • Gemini 3.8 Flash TTS and Flash-Lite TTS allow custom voice creation.
  • Developers can describe or record a voice to save and reuse it.
  • Voice replication uses a new Voices endpoint (POST /v1beta/voices).
  • Two clean recordings (10-30 seconds) from the same speaker are required.
  • A separate consent recording is needed for voice replication.
🕒 2026-09-23 · new reporting from Google DeepMind
  • Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS are new text-to-speech models.
  • Gemini 3.8 Flash TTS is for creative direction and character design.
  • Gemini 3.8 Flash-Lite TTS is for high-volume, cost-efficient scale.
  • The new models are part of the Gemini Audio family.
  • Gemini 3.5 Live Translate and 3.5 Transcribe are also part of the Gemini Audio family.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

ChatGPT Voice and Gemini Live offer real-time voice chat capabilities, but differ in conversational expressiveness and internet search behavior. ChatGPT Voice uses more natural conversational markers and is more likely to search the internet for current information, while Gemini Live has a flatter delivery and relies more on its existing knowledge base.

Google has added Live Avatar to its Gemini 3.8 Live models, integrating near real-time visual personas with live dialogue capabilities. This development allows AI agents to communicate using facial expressions and dynamic visuals, expanding their application in customer service and sales support.

Google released Gemini 3.8 Live, which includes a "Live Avatar" feature allowing real-time conversations with an animated AI persona that lip-syncs and shows facial expressions. This feature is currently exclusive to Gemini Enterprise customers and supports 97 languages with consistent video fidelity.

Gemini 3.8 Live now includes Live Avatar, a feature that pairs near real-time video generation with speech to create dynamic visual personas for AI models. This allows enterprises to offer more interactive virtual experiences with visual presence, precise lip-syncing, and natural expressions.

Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise, offering conversational video agents with synchronized lip-syncing and fluid dialogue. This release provides enterprises with advanced AI capabilities for customer interactions across web, mobile, and interactive kiosks.

Google released Gemini 3.8 Flash TTS and Flash-Lite TTS, allowing developers to create custom voices through the Gemini API and Google AI Studio. This update enables users to describe or record a voice, then save and reuse it across applications, moving beyond pre-existing text-to-speech options.

Google's Gemini has released two new text-to-speech (TTS) models, 3.8 Flash TTS and 3.8 Flash-Lite TTS, expanding its audio family. These models offer advanced voice generation capabilities for creative applications and high-volume audio content, providing more dynamic and expressive audio experiences for developers and enterprises.

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, new voice models designed to reduce latency in voice agents by allowing them to continue speaking during background processing. This release follows OpenAI's GPT-Live-1, with both companies offering different architectural approaches to address the latency challenge in conversational AI.

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new AI models designed to improve near real-time reasoning for voice agents. These models aim to make AI conversations more intuitive and intelligent, offering capabilities for both scalable, cost-efficient applications and high-complexity tasks requiring multi-step reasoning.