NewsStocksGoogle DeepMind Launches Gemini 3.8 Live and 3.8 Live Extended Thinking for Natural Voice AI

Google DeepMind Launches Gemini 3.8 Live and 3.8 Live Extended Thinking for Natural Voice AI

Author: Google DeepMind Blog·

Key Takeaways

  • Google DeepMind launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, describing them as its most advanced live dialogue models to date.
  • Gemini 3.8 Live Extended Thinking ranked first overall on Artificial Analysis' Speech to Speech Quality Index, scoring 68.6% on the τ-Voice agentic benchmark and 97.7% on Big Bench Audio.
  • The models process visual inputs in near real time, switch automatically among 97 supported languages mid-conversation, and execute tools and API calls in the background while conversations continue.
  • Gemini 3.8 Live is positioned for scale and cost efficiency, while Extended Thinking is designed for high-complexity tasks with simultaneous reasoning and speaking, including live progress narration.
  • Both models are available starting today through the Gemini API and Google AI Studio, with enterprise and consumer rollouts across Gemini Enterprise, Google Workspace, Search Live, and the Gemini app.
Google DeepMind Launches Gemini 3.8 Live and 3.8 Live Extended Thinking for Natural Voice AI

Google DeepMind announced on September 15, 2026, the launch of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which the company describes as its most advanced live dialogue models to date. According to the announcement, the two models bring advancements in near real-time reasoning to more effectively enable voice agents, with major upgrades in intelligence and parallel reasoning that make them more intuitive to collaborate with and better suited to executing complex tasks using voice.

The post was authored by Tom Ouyang, Principal Engineer, and Malini Jaganathan, Member of Technical Staff, writing on behalf of the Gemini Audio Team.

Google said the launch is intended to make voice interactions more natural, fluid, and intelligent. The new models handle complex reasoning, real-time visual context, and background task execution without interrupting an ongoing conversation. They can manage interruptions, switch between languages, and explain their thought process while they work — behavior Google characterizes as a step toward AI that feels like a real conversation partner. Whether users are solving a complex problem or simply chatting, the company says the experience is designed to feel as though the AI is listening and thinking along with them. The features are available starting today through the Gemini API, Google Workspace, and the Gemini app.

The two models target distinct use cases:

  • Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding.
  • Gemini 3.8 Live Extended Thinking is built for high-complexity tasks, with increased intelligence and multi-step reasoning.

For developers and enterprises, Google positions the models as the building blocks for reliable, production-ready voice agents. The company also says they make speaking with Gemini across the Gemini app, Google Workspace, and Search more fluid and collaborative, helping users tackle complex tasks using just their voice.

Benchmarks and performance

Gemini 3.8 Live Extended Thinking delivers what Google describes as enterprise-grade task completion and intelligence, capturing the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index with a score of .6. The model leads in agentic task completion, scoring 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark. It also provides strong reasoning capabilities, scoring 97.7% on Big Bench Audio, while maintaining a highly competitive price point compared with other frontier models.

Gemini 3.8 Live has shown high preference among users, securing second place in the Speech Agent Arena. Beyond that result, Google says it remains highly cost-effective, providing developers and enterprises with a capable and efficient model built for scale.

On ServiceNow's EVA-Bench, a benchmark for evaluating voice agents, Google said its models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality. The company notes the evaluation was run on the Live API on the Gemini Enterprise Agent Platform.

Visual context, 97 languages, and background execution

Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses. It automatically detects and transitions between its 97 supported languages mid-conversation. The model also executes tools and API calls in the background while the conversation continues, allowing it to acknowledge requests and keep chatting while tasks finish behind the scenes.

Video demonstrations published with the announcement show the model guiding employee onboarding in real time, using visual context to answer live questions, and playing chess in near real-time by combining visual context, reasoning, and natural conversational flow.

Reasoning and speaking simultaneously

For tasks that require deeper reasoning, Gemini 3.8 Live Extended Thinking reasons and speaks at the same time. Google says the model delivers increased intelligence for complex workflows while maintaining an uninterrupted conversational flow. It uses early verbal cues, such as "Let me check that…," to acknowledge prompts naturally, and provides live progress narration to walk users through multi-step background tasks as those tasks progress.

Demonstrations include Gemini 3.8 Live Extended Thinking transforming raw sketches and near real-time voice feedback into functional React components, and coordinating multi-step bookings and asynchronous function calls without interrupting natural live conversation. Google also showed Gemini 3.8 Live building complete business plans and custom marketing toolkits on the fly through natural speech.

Google Workspace and Search integration

Across Google Workspace and Search, the Live models are designed to deliver more intuitive, collaborative experiences, particularly when tackling the most complex tasks. Gemini 3.8 Live Extended Thinking can be used in Google Workspace through Docs Live, Gmail Live, and Keep Live. Search Live offers step-by-step, real-time troubleshooting help powered by Gemini 3.8 Live.

Developer and enterprise voice ecosystem

Using the Gemini Live API, developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to build and deploy high-performance, voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.

Google added that it is partnering with companies including Salesforce, Genspark, and Lumeris that are excited about 3.8 Live and 3.8 Live Extended Thinking, highlighting the models' latency, fluidity, and tool-calling capabilities.

SynthID watermarking for transparency

All audio generated by Google's AI products is watermarked with SynthID. The company says this imperceptible watermark is woven directly into the audio output, ensuring AI-generated content remains detectable to help prevent misinformation. For details on the company's approach to safety and responsibility, Google points readers to the model card.

Availability

Both models begin rolling out today, across developer, enterprise, and consumer surfaces.

Gemini 3.8 Live:

  • For developers: in the Gemini API and Google AI Studio.
  • For enterprises: in private preview in Gemini Enterprise, and coming soon to Gemini Enterprise for Customer Experience.
  • For everyone: in Search Live.

Gemini 3.8 Live Extended Thinking:

  • For developers: in the Gemini API and Google AI Studio.
  • For enterprises: in private preview in Gemini Enterprise, and coming soon to Gemini Enterprise for Customer Experience and Google Workspace business customers.
  • For everyone: in Gemini Live, for Google AI Pro and Ultra subscribers in Workspace in Docs, and for all Google AI subscribers in Gmail and Keep.

The full announcement, including video demonstrations, is available on the Google DeepMind blog.