Google Launches Gemini 3.8 Live Audio Models for Real-Time Conversational AI
Key Takeaways
- •Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, audio models designed for building voice agents that reason and execute tasks during live conversations.
- •The models can perform API and tool calls while streaming audio responses, accept live visual inputs, and support more than 97 languages.
- •The Extended Thinking version includes configurable thinking, a feature that lets users turn a model's internal reasoning on or off.
- •The architecture separates interactive latency from reasoning latency, allowing agents to keep a conversation moving while background reasoning continues rather than pausing for an answer.
- •Analysts see Google catching up to vendors such as Apple while advancing real-time AI interaction, with enterprise applications including customer support chatbots, sales processes, and multimodal agentic applications.

Google has introduced two new audio models that enable enterprises to deploy intelligent AI agents with deeper reasoning, more interactivity and conversational speed.
Launched on September 15, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking let developers build voice agents that can reason and tasks while maintaining the flow of a conversation, Google said. The models support real-time interaction without the traditional prompt-response structure, allowing users to converse naturally rather than issue discrete prompts.
The models' key capabilities include performing API and tool calls — the mechanisms agents use to connect with outside systems and execute tasks — while continuing to stream audio responses to users, accepting live visual inputs that help agents understand both what users say and what they see, and supporting more than 97 languages, a range relevant to enterprises deploying voice agents across regions. The Extended Thinking version adds configurable thinking, a feature that lets users turn an AI model's internal reasoning on or off.
The release comes two weeks after Google initially introduced the Gemini 3.8 Flash and 3.8 Cyber foundation models.
With 3.8 Live's capabilities, Google appears to be catching up to vendors such as Apple, which already offers real-time transcription in productivity apps such as Notes and Voice Memos, said Bradley Shimmin, an analyst at Futurum Group. Still, Google's technology takes it to the next level by changing how users interact with AI, he said.
Real Conversations
"The way it's architected … you can customize this to behave in a lot of different ways," Shimmin said. He noted that while enterprises can use the models for traditional translation use cases, Google has also designed them so users don't interact in a prompt-response way; rather, they can have a real conversation.
"It's just a natural conversation with the ability to interrupt it in real time to inject," Shimmin said. "You can inject context into the conversation without interrupting what it's doing."
With Google's advanced speech technology, the new models don't just answer a question at face value, said Sid Nag, founder of Tekonyx.
"The model is designed to do background reasoning," Nag said. "It's doing it during the conversation so you can keep the conversation alive while it's reasoning."
The technology is an advancement in which interactive latency — the delay between a user performing an action and the system responding — is separated from reasoning latency, the total time it takes an AI model to process information. That separation is what allows an agent to keep a conversation moving while reasoning continues in the background, rather than pausing the interaction until an answer is ready.
"It's doing it in real time, rather than halting and then coming back with another answer," Nag said.
The speech models also change how an AI agent responds and interacts because "an AI agent doesn't necessarily have to choose between being fast and being thoughtful," Nag continued. "It can actually maintain a real-time interaction."
Enterprise applications for the new models include customer support chatbots, sales processes and multimodal agentic applications that combine text, voice and video, Nag said. Agentic applications are software in which AI agents reason and execute tasks rather than simply respond to prompts — a pattern that relies on mid-conversation tool calling and background reasoning of the kind the new models provide.
Reporting by Esther Shittu, AI Business. Original report: Gemini 3.8 Live Transforms Conversational AI, published September 17, 2026.