Google DeepMind Launches Gemini 3.8 Flash TTS and Flash-Lite TTS, Its Most Expressive Text-to-Speech Models Yet
Key Takeaways
- •Google DeepMind launched two text-to-speech models: Gemini 3.8 Flash TTS, built for expressive character design and scene direction, and Gemini 3.8 Flash-Lite TTS, built for high-volume, cost-efficient speech generation.
- •Users can design original voices through natural language prompts, choose from a library of more than 2,000 production-ready voices spanning over 100 languages and dialects, and replicate a voice from a 30-second audio sample.
- •Both models provide line-by-line control over pacing, emotion, nonverbal cues, and two-speaker scene staging, and can generate hours of continuous audio with minimal speaker drift.
- •Google reports that Gemini 3.8 Flash TTS placed first on Hume AI's Voice Design Benchmark with a score of 71.4, and the two models took the top two positions on Hume AI's Overall Quality Index.
- •Voice replication requires verified consent from the voice owner, all generated audio carries SynthID watermarking, and replication through AI Studio is unavailable in regions including Illinois, Texas, the EEA, the United Kingdom, Switzerland, and India.

Google DeepMind announced on September 23, 2026 the release of 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models the company describes as its most expressive audio generation models to date. The models generate custom character voices and fully directed, scene-based dialogue, and they are rolling out across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Text-to-speech has become a rapidly developing corner of generative AI, with audiobook production, gaming, dubbing, and real-time voice agents all driving demand for synthetic speech that sounds convincingly human.
The announcement, published the Google DeepMind blog, was authored by Leland Rechis, Group Product Manager, and Alan Cowen, Director of Research Science, writing on behalf of the Gemini Audio Team.
In its summary, Google says users can create custom voices from scratch or replicate existing ones with simple natural language prompts, direct their audio line by line to control pacing, emotion, and even realistic conversational sounds, and use the models for high-quality audiobooks, podcasts, or real-time voice agents at scale. Built-in safety tools, including watermarking, are included to help keep generated audio secure.
Two models for different workloads
Google frames the launch as transforming voice generation "from static presets into a dynamic creative studio," enabling creators, developers, and enterprises to build richer, more expressive audio experiences. The company adds that the models also enable improved user experiences in its own products, including Gemini Notebook and Google Vids.
Gemini 3.8 Flash TTS is built for deep creative direction and character design. It can create entirely new voices from scratch using natural language prompts to bring characters to life across gaming, immersive audiobooks, podcasts, and interactive media. Users can direct every performance line by line, with granular control over acting cues, pacing, dialect shifts, and backchanneling.
Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient scale. Google says it is optimized for high-volume dubbing, audio content creation, and expressive voice agents, with fine-grained control over tone, pacing, and expressive nuance. The Flash and Flash-Lite split follows the naming convention Google already uses across the Gemini family, where Lite variants serve as lighter, cost-focused options for high-throughput workloads.
The two models complement a fast-growing Gemini Audio family that previously included 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.
Create and customize your own voices
With the new models, Google is scaling up from 30 original voices to what it calls an infinite library. Whether the goal is an entirely original character voice or a consistent brand ambassador, the company says Gemini 3.8 Flash TTS powers a full vocal studio, letting users create and deploy expressive, natural-sounding voices for every moment while developers and enterprises build custom audio experiences on top of them.
The voice-creation capabilities announced today include:
Generative voice design. Users can create bespoke voices from scratch with Gemini 3.8 Flash TTS by customizing role, accent, and voice characteristics across more than 100 languages and dialects through natural language prompting — whether, in Google's examples, bringing a dramatic, fire-breathing dragon to life or crafting a charismatic narrator with a distinct regional cadence. Demo clips released alongside the post show the model generating a high-energy DJ voice from Melbourne, a super-tinny, monotone robot voice, and a Japanese dragon.
Expansive voice library. The platform offers more than 2,000 production-ready voices with broad language coverage, including regional varieties such as Mexican Spanish, Quebec French, and Scots English.
Voice replication. Users can recreate consistent vocal profiles from a 30-second audio sample of their own voice, or of a voice they have the rights to use. The feature is backed by built-in consent verification, SynthID watermarking, and C2PA credentials — an open industry standard for recording where a piece of media came from and how it was produced — to protect both developers and their vocal talent.
Save and scale. Custom voices can be saved and managed to ensure consistent performance and minimal drift across ongoing projects.
Voice remixing. Coming soon, users will be able to pick a voice from the library and fine-tune its timbre, pitch, pace, and accent, using prompts to dial in characteristics such as "add subtle Southern US accent" or "soften the delivery."
Direct the performance, line by line
After voices are selected, both TTS models give users precise control over how each line is delivered. Users can write their own stage directions or let Gemini steer delivery with natural script cues — from a calm customer service agent to a whispered suspense scene. Google's demos show the model powering natural, highly expressive conversations for interactive voice agents, as well as granular script control used to build deeply engaging, immersive audio experiences and fully performed dialogue scenes with natural turn-taking.
Other performance capabilities include:
Long-form generation. The models maintain high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift, which Google positions as ideal for podcasts and audiobooks.
Native two-speaker scene staging. Multi-turn conversations can be directed seamlessly from a single script — whether for a podcast or dramatic storytelling — while keeping both voices distinctly separated with natural conversational turn-taking.
Scripted vocal bursts and backchanneling. Nonverbal cues such as <laughs>, <sigh>, and <gasp>, together with active-listening interjections like |mhm| or |yeah|, add realistic conversational texture for precise comedic timing and reaction beats.
Benchmarks and multilingual reach
Google reports that Gemini 3.8 Flash TTS delivers leading voice customization capabilities, securing the #1 overall spot on Hume AI's Voice Design Benchmark, a third-party evaluation from the voice-focused AI research company Hume AI, with a score of 71.4, while also leading in accent modeling at 60.8. The company adds that Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS enable expressive performances without sacrificing reliability, securing the #1 and # spots, respectively, on Hume AI's Overall Quality Index.
Compared with Gemini 3.1 Flash TTS, Google says the new model shows major improvements across a wide range of use cases, including long-form content and dual-speaker screenplay control.
In blind human preference evaluations on Voice Arena, where listeners compare anonymized samples without knowing which model produced them, Gemini 3.8 Flash and Flash-Lite TTS secured top positions among competitors in key global languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish, and Hindi. With support for more than 100 languages, the company says the models empower creators, developers, and enterprises to build high-quality, multilingual voice experiences worldwide.
Safety, consent, and transparency
Google says it built the voice creation and replication capabilities with strict safeguards to help protect voice talent, respect identity, and ensure content transparency. For voice replication, the system leverages consent verification: users must provide a verbal consent recording the voice owner that matches the reference speaker before a voice can be created. The safeguards speak to one of the most closely watched risks in generative audio: the misuse of cloned voices for impersonation and fraud, which has made consent and verification workflows a focal point of AI policy debates.
More broadly, every audio clip generated by the Gemini Audio models is watermarked with SynthID, DeepMind's watermarking technology for identifying AI-generated content. Google describes the watermark as imperceptible and woven directly into the audio output, keeping AI-generated speech detectable in order to help prevent misinformation. Additional detail on the company's approach to safety and responsibility is available in the Gemini 3.8 audio model card.
Google also notes a regional restriction: voice replication through AI Studio is not available in Illinois, Texas, the European Economic Area (EEA), the United Kingdom, Switzerland, and India.
A voice design workspace in Google AI Studio
Starting today, developers can try the new speech generation capabilities in Google AI Studio. Built like a voice design workspace, the tool lets users prompt entirely new vocal identities from scratch or replicate their own voice, then bring those voices directly into a dual-speaker screenplay editor to direct line-by-line delivery.
Developer platforms and launch partners
Through the Gemini API, developer platforms including Agora, LiveKit, Pipecat, and Vercel — a group spanning real-time communications infrastructure, voice-agent frameworks, and deployment tooling — enable developers to build and deploy high-performance speech generation experiences with ease.
Google is also partnering with companies including Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang, which are integrating the latest TTS models to help accelerate global dubbing, localize media with nuanced regional accents, and power conversational voice agents at scale.
Availability
Both models are rolling out starting today.
Gemini 3.8 Flash TTS:
- For developers: available in the Gemini API and Google AI Studio.
- For enterprises: coming soon via API in Gemini Enterprise.
- For consumers: available in Gemini Notebook.
Gemini 3.8 Flash-Lite TTS:
- For developers: available in the Gemini API and Google AI Studio.
- For enterprises: coming soon via API in Gemini Enterprise.
- For consumers: available in Google Vids.
The capabilities Google has marked as coming soon — voice remixing and API access through Gemini Enterprise — are the next milestones to watch as the rollout extends beyond today's launch partners and early access points.
With the release, the Gemini Audio lineup spans live translation and transcription, real-time voice conversation through 3.8 Live and 3.8 Live Extended Thinking, and now expressive speech generation, giving creators, developers, and enterprises a single model family for building voice-driven products.