Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS
Google rolls out Gemini 3.8 Flash TTS and Flash-Lite TTS models featuring prompt-based voice design via Gemini API and AI Studio.

Stock photo for illustration only, not from the actual event
- Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS models
- Supports prompt-based voice design and stage directions in scripts
- Allows voice replication from a 30-second reference audio sample
- Embeds SynthID watermarks and C2PA credentials in generated clips
Google has officially released its latest text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, introducing advanced prompt-based voice design and enhanced directional controls. Both models are currently rolling out and can be accessed through the Gemini API and Google AI Studio.
Deployment is handled strictly via API access, with no open weights provided for self-hosting at this time. Enterprise API access through Gemini Enterprise is listed as coming soon. Within AI Studio, developers can utilize the playground links using the specific model identifiers gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.

Stock photo for illustration only, not from the actual event
While previous Gemini TTS iterations offered a set of 30 original voices, the 3.8 release transitions to a substantially expanded voice system. Both models accept stage directions written directly into the script, allowing Gemini to steer vocal delivery smoothly using natural script cues.
The integration of script-based stage directions and prompt-driven voice design highlights a major shift in generative audio technology, moving beyond monotonous text reading toward context-aware speech synthesis. This capability allows developers to tailor emotional delivery and pacing directly through text inputs, significantly streamlining the creation of realistic voiceovers for media and interactive applications.
Furthermore, the voice replication feature builds a consistent vocal profile utilizing just a 30-second audio sample. Google mandates that the reference sample must belong to the user or someone they have legal rights to use, ensuring compliance through rigorous verification safeguards.
To prevent misuse, voice replication requires a recorded verbal consent from the voice owner, which is matched against the reference speaker. For safety and provenance, every audio clip generated by Gemini Audio models features an imperceptible SynthID watermark embedded directly into the output. Additionally, replicated voices carry C2PA content credentials, aligning with Google's broader safety framework outlined in the Gemini 3.8 Audio model card.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment