Google DeepMind introduces Gemini 3.8 Live with Live Avatar
Key Takeaways
- •Live Avatar enables enterprise agents to listen, see and respond through synchronized audio and video in near real time.
- •The system can execute tools and retrieve information in the background without interrupting an active conversation.
- •Live Avatar supports multilingual speech-to-speech synchronization across 97 languages, including language switching during a conversation.
- •Organizations can select preset avatars, while custom avatar creation from reference images is limited to allowlisted enterprises.
- •SynthID watermarking is embedded in generated audio and video to help identify AI-produced content.

Google DeepMind introduced Gemini 3.8 Live with Live Avatar on Sept. 24, 2026, adding near-real-time visual presence to its conversational AI. The feature combines Gemini’s native live dialogue capabilities with low-latency streaming video to create more natural and intuitive interactions for enterprises and their users.
The announcement was authored by Shuo-yiin Chang, a research scientist, and CJ Zheng, a software engineer, on behalf of the Gemini Audio Team. It follows last week’s launch of Gemini 3.8 Live.
According to Google DeepMind, Live Avatar pairs near-real-time video generation with speech so the system can listen, see and speak through a dynamic visual persona. The feature is designed with precise lip-syncing, natural facial expressions and fluid turn-taking, allowing to make virtual services more interactive. Potential uses described by the company include customer service and interactive walkthroughs.
Gemini 3.8 Live with Live Avatar is available starting today in Gemini Enterprise. The original announcement is available on the Google DeepMind Blog.
Multimodal conversations
Google DeepMind said conversation is inherently multimodal, involving listening, visual attention, speech and facial expressions. Live Avatar brings these capabilities to enterprise agents by processing visual and audio inputs simultaneously and generating responses that combine expressive audio and video.
The company presented the feature as capable of taking in what it sees and hears in near real time, then responding through audio and video to support more natural conversations. For the customer-service and interactive-walkthrough scenarios Google DeepMind cited, the practical effect is a virtual agent that reacts to what it is shown and answers with voice and facial cues in one conversation.
Background tool execution
Live Avatar also uses Gemini’s reasoning capabilities and asynchronous tool calling. It can initiate tool calls and retrieve data in the background while continuing an active conversation. This allows it to handle complex tasks without interrupting the dialogue.
Google DeepMind gave checking in a guest at a hotel as an example of a task in which tools can run in the background while the conversation continues. For enterprises, the value is continuity: customers experience an unbroken conversation while data retrieval and task execution happen out of sight.
Support for 97 languages
The feature includes native multilingual speech-to-speech synchronization. Google DeepMind said Live Avatar can dynamically adapt its lip-syncing and facial expressions while switching between 97 languages, without degrading video fidelity or causing visual drift.
The company also demonstrated switching languages in the middle of a conversation, with the avatar’s lip-sync and expressions adapting across languages. For businesses serving users across regions, that coverage — and the ability to switch languages without visual artifacts — speaks directly to global customer interactions.
Customizable avatars
Organizations can choose from a library of preset avatars with different appearances, voices and expressive styles. They can also customize Live Avatars to reflect a brand or character identity.
Using a high-quality reference image, developers can generate a fully animated and responsive avatar while preserving the reference likeness, brand styling or character identity. Google DeepMind said custom avatar creation is currently available only through enterprise allowlisting, leaving the preset library as the option for organizations outside that list.
Watermarking and safety
Google DeepMind said Live Avatar includes safeguards intended to respect identity and keep AI-generated content transparent. Output generated by its AI products is watermarked with SynthID, an imperceptible watermark embedded directly in audio and video.
The company said the watermark helps keep AI-generated content detectable and is intended to help minimize misinformation and misattribution. For enterprises putting expressive video personas in front of customers, the watermark is what keeps AI-generated interactions identifiable as synthetic. Further information about the company’s safety and responsible-deployment approach is available in its model card.
Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise. Google DeepMind also directed developers to its API documentation for getting started. For teams that want brand-specific avatars beyond the preset library, the enterprise allowlist for custom avatar creation is the access detail to track.