Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two new text-to-speech models designed to generate and customise AI voices for applications ranging from podcasts and audiobooks to dubbing, interactive media and conversational voice agents.
Gemini 3.8 Flash TTS is positioned for voice design and creative direction, allowing users to generate new voices using natural language prompts. Developers can specify characteristics such as role, accent and vocal style across more than 100 languages and dialects. The model also provides access to more than 2,000 production-ready voices.
The model can replicate a consistent vocal profile from a 30-second audio sample, provided users have the rights to the voice. Google said the feature incorporates consent verification, requiring a verbal consent recording from the voice owner that matches the reference speaker before a replicated voice can be created.
Voice replication through Google AI Studio, however, is not available in India, along with Illinois, Texas, the European Economic Area, the UK and Switzerland.
Gemini 3.8 Flash-Lite TTS is aimed at high-volume and cost-efficient applications, including dubbing, audio content production and voice agents. Both models allow users to control elements such as pacing, tone and expressive delivery.
Google has also introduced line-by-line performance controls that enable creators and developers to add directions to scripts. The models can handle long-form audio while maintaining voice characteristics and can stage two-speaker conversations from a single script. Users can also incorporate non-verbal cues such as laughs, sighs and gasps into generated speech.
As part of its safeguards around synthetic audio, Google said all audio generated through Gemini Audio models carries its SynthID watermark. Voice replication also uses C2PA credentials alongside consent verification.
The company is working with platforms and companies including Figma, HeyGen, Linguana, Wondercraft and 99.co to integrate the models for applications such as media localisation, dubbing and conversational voice agents.
Both Gemini 3.8 Flash TTS and Flash-Lite TTS are rolling out through the Gemini API and Google AI Studio. Flash TTS is also available through Gemini Notebook, while Flash-Lite TTS is being introduced in Google Vids. Enterprise access for both models is expected to arrive through the Gemini Enterprise API.
The launch expands Google's Gemini Audio portfolio as technology companies increasingly develop generative AI systems capable of producing speech and powering voice-based digital experiences.
Disclaimer: This article may include information derived from interviews, press releases, public statements, research, company communications and other publicly available or third-party sources. Such material may be summarised, paraphrased or contextualised for journalistic and editorial purposes. All rights in third-party content remain with their respective owners.