Google launches Gemini text-to-speech models with voice replication
Google has introduced new Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models, enabling developers to create custom voices and replicate existing ones.

Google has launched two new Gemini text-to-speech (TTS) models: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. These models are designed to allow developers to generate custom speech voices, direct dialogue delivery, and replicate a speaker's voice.
The technology is available through Google AI Studio and the Gemini application programming interface (API). It is also being integrated into Google's own products, including Gemini Notebook and Google Vids. The new models support over 100 languages and are suitable for various applications, such as producing audiobooks, games, and podcasts, as well as creating dialogue scenes.
According to Google, the use of voice replication requires written consent from the voice owner. The company emphasizes that all content generated by Gemini Audio models will carry a SynthID watermark, which helps to identify AI-generated material.
The release expands Google's existing speech technology offerings, which already include Cloud Text-to-Speech and Chirp 3 HD voices. However, the company's documentation does not detail how the new Gemini models are positioned relative to previous products.