Google Launches Gemini 3.8 Text-to-Speech With Advanced Custom Voices
Google introduces Gemini 3.8 Flash TTS and Flash-Lite TTS, bringing natural language voice generation, line-by-line direction, and built-in safety controls.
Google Expands Audio Capabilities With Gemini 3.8 TTS Models
Google has officially introduced two new text-to-speech models to the Gemini family: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The release shifts voice generation away from static presets into a dynamic studio experience designed to let creators, developers, and enterprises craft realistic audio assets across more than 100 languages and dialects.
The standard Flash TTS variant targets deep creative direction and character design, allowing users to build custom voices from scratch using natural language prompts. Meanwhile, the Flash-Lite TTS variant is optimized for high-volume, cost-efficient scaling, making it well-suited for mass dubbing, localized media, and expressive voice agents.
Granular Performance Direction and Multilingual Scale
Both models introduce line-by-line performance direction, giving users the ability to apply script cues, pacing changes, dialect shifts, and non-verbal conversational textures like laughs, sighs, and active-listening interjections. The system also supports native two-speaker scene staging for multi-turn conversations from a single script without notable speaker drift over long-form durations.
According to vendor disclosures, Gemini 3.8 Flash TTS secured top placement on Hume AI’s Voice Design Benchmark, alongside leading scores for accent modeling and overall quality indices. Support spans over 100 languages, with initial rollouts appearing in Google AI Studio and the Gemini API for developers, alongside integrations within Gemini Notebook and Google Vids.
Safeguards, Verification, and What Changes for Developers
To address concerns surrounding synthetic media misuse, Google has integrated mandatory safety protocols into the platform. Voice replication features require verbal consent verification matching a reference speaker before a custom profile can be generated. Additionally, every audio clip produced by the models features embedded SynthID watermarking and C2PA credentials to help maintain content transparency.
For practitioners and developers building conversational interfaces or media localization pipelines, these models offer an expanded toolkit with deep parameter control. However, geographic restrictions apply to specific features such as voice replication, which remains unavailable in regions including the EEA, UK, Switzerland, and select U.S. states at launch.
Source
Official announcement or documentation
Some links on this page may be affiliate links. If you buy through them we may earn a commission at no extra cost to you. See our affiliate disclosure.