Google 2015 logo (Wikimedia Commons, public domain)
Google brings custom voice generation to Gemini 3.8 TTS
Gemini 3.8 Flash TTS can generate new voices from prompts, replicate voices from 30 second samples and direct performances line by line.
Google has launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash Lite TTS, two new text to speech models designed for expressive voice generation and large scale audio production.
Gemini 3.8 Flash TTS allows developers and creators to generate entirely new voices using natural language prompts. Users can specify characteristics including role, accent and vocal style across more than 100 languages and dialects.
The model can also replicate a voice from a 30 second audio sample when the user has permission to use it. Google said the feature includes consent verification, SynthID watermarking and C2PA credentials.
Both models support line by line performance direction, including changes to pacing, tone, whispers, pauses and other acting cues. They can also maintain character consistency across long form audio and stage two speaker conversations from a single script.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
Google said Gemini 3.8 Flash TTS ranked first on Hume AI’s Voice Design Benchmark with a score of 71.4 and led its accent modeling category with 60.8. Flash and Flash Lite also took the top two positions on Hume AI’s Overall Quality Index.
The models are available through the Gemini API and Google AI Studio. Gemini 3.8 Flash TTS is also rolling out to Gemini Notebook, while Flash Lite is being added to Google Vids.