Google launches benchmark-topping text-to-speech models
Gemini 3.8 Flash and Flash-Lite TTS now available on Google Cloud, supporting 130 and 101 languages respectively.
What to know
- Google released Gemini 3.8 Flash TTS and Flash-Lite TTS, dual text-to-speech models optimized for quality and cost efficiency respectively, now available on Google Cloud.
- Both models ranked first and second on independent audio quality benchmarks and support customization through 2,000+ prepackaged voices or synthetic voice generation from audio samples.
- Google embedded SynthID watermarking and C2PA metadata in all generated speech to enable detection of AI-generated content and document its creation history.
“Flash-Lite TTS is optimized for cost-cost efficiency and inference speed. Flash TTS offers better audio quality for a higher price.”
SiliconANGLE, Reporter · SiliconANGLE
Google LLC AI model developerLeland Rechis Google stafferAlan Cowen Google staffer
How it unfolded 1 development · click the chart to see its coverage articles
-
1
Google releases Gemini 3.8 Flash and Flash-Lite TTS models
Google made two new text-to-speech models available through its cloud platform. Flash TTS supports 130 languages and prioritizes audio quality, while Flash-Lite TTS supports 101 languages and is optimized for cost efficiency and speed. Both offer access to over 2,000 prepackaged voices and allow developers to create custom voices using natural language prompts or 30-second audio samples.
“These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids.”
— Leland Rechis and Alan Cowen -
first by SiliconANGLE, 2d ago
-
-
background
Google embeds watermarking and metadata in generated speech — The models use SynthID technology to embed inaudible audio watermarks detectable by AI tools, and attach C2PA records to all generated audio files to document creation time, modification status, and related metadata.
-
background
Models rank first and second on audio quality benchmarks — Flash TTS and Flash-Lite TTS earned the top two positions on an audio quality benchmark developed by Hume AI Inc., and outperformed competing models on multiple versions of the Voice Arena benchmark that measures speech quality based on human feedback.