Google brings custom voice generation to Gemini 3.8 TTS

Google 2015 logo (Wikimedia Commons, public domain)

Google brings custom voice generation to Gemini 3.8 TTS

Gemini 3.8 Flash TTS can generate new voices from prompts, replicate voices from 30 second samples and direct performances line by line.

Google has launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash Lite TTS, two new text to speech models designed for expressive voice generation and large scale audio production.

Gemini 3.8 Flash TTS allows developers and creators to generate entirely new voices using natural language prompts. Users can specify characteristics including role, accent and vocal style across more than 100 languages and dialects.

Advertisement

The model can also replicate a voice from a 30 second audio sample when the user has permission to use it. Google said the feature includes consent verification, SynthID watermarking and C2PA credentials.

Both models support line by line performance direction, including changes to pacing, tone, whispers, pauses and other acting cues. They can also maintain character consistency across long form audio and stage two speaker conversations from a single script.

Google said Gemini 3.8 Flash TTS ranked first on Hume AI’s Voice Design Benchmark with a score of 71.4 and led its accent modeling category with 60.8. Flash and Flash Lite also took the top two positions on Hume AI’s Overall Quality Index.

The models are available through the Gemini API and Google AI Studio. Gemini 3.8 Flash TTS is also rolling out to Gemini Notebook, while Flash Lite is being added to Google Vids.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
Google brings custom voice generation to Gemini 3.8 TTS
Google brings custom voice generation to Gemini 3.8 TTS

Gemini 3.8 Flash TTS can generate new voices from prompts, replicate voices from 30 second samples and direct performances line by line.

Share

Add us on Google

Google 2015 logo (Wikimedia Commons, public domain)

Google has launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash Lite TTS, two new text to speech models designed for expressive voice generation and large scale audio production.

Gemini 3.8 Flash TTS allows developers and creators to generate entirely new voices using natural language prompts. Users can specify characteristics including role, accent and vocal style across more than 100 languages and dialects.

Advertisement

The model can also replicate a voice from a 30 second audio sample when the user has permission to use it. Google said the feature includes consent verification, SynthID watermarking and C2PA credentials.

Both models support line by line performance direction, including changes to pacing, tone, whispers, pauses and other acting cues. They can also maintain character consistency across long form audio and stage two speaker conversations from a single script.

Google said Gemini 3.8 Flash TTS ranked first on Hume AI’s Voice Design Benchmark with a score of 71.4 and led its accent modeling category with 60.8. Flash and Flash Lite also took the top two positions on Hume AI’s Overall Quality Index.

The models are available through the Gemini API and Google AI Studio. Gemini 3.8 Flash TTS is also rolling out to Gemini Notebook, while Flash Lite is being added to Google Vids.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.