Google launches Gemini 3.8 Flash TTS with custom voice design and 100+ languages
New text-to-speech models let users design custom character voices, direct scene dialogue, and clone voices with consent verification.
What to know
- Google released two new TTS models — Gemini 3.8 Flash TTS and Flash-Lite TTS — with custom voice design, over 100 language support, and 2,000+ ready-made voices.
- Voice replication requires verbal consent from the voice owner as a safeguard against misuse.
- The models rank #1 or #2 on several third-party benchmarks (Hume AI, Artificial Analysis) for quality and pronunciation, though they trail the fastest competing TTS models on raw generation speed.
- Availability spans Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
Google Developer of Gemini modelsGoogle DeepMind AI research lab behind the models@officiallogank Google AI product figureArtificial Analysis Independent AI benchmarking tracker
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
1
Artificial Analysis publishes independent speed and quality benchmarks
Third-party AI benchmarking tracker Artificial Analysis reported Gemini 3.8 Flash TTS processes 44.1 characters per second (about 2.7x realtime) versus 40.2 for Flash-Lite, placing Google's models behind faster competitors like Falcon 2, while noting the Flash model debuts at #1 on its Pronunciation Robustness Benchmark and #2 on its Provider Voice Arena Leaderboard.
“Gemini 3.8 Flash TTS processes 44.1 characters per second, compared to 40.2 characters per second for Gemini 3.8 Flash-Lite TTS, approximately 2.7x and 2.4x faster than realtime, respectively.”
— @artificialanlys -
The new Gemini 3.8 TTS models are super-cheap and can generate conversations between multiple voices (from 2,000+, or you can clone your own) - I built a little UI for it, then had Claude knock up a script where two pelicans debate moving to Pacifica Pier
2 more of the top 3 · 31 posts in this stretch
-
Here's a video of my Emotive Audiobook Creator, KeenLore, a locally hosted web app:https://www.youtube.com/watch?v=WAeHgE94rVoNo cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character…
-
H
Gemini 3.8 text-to-speech Link: https:// blog.google/innovation-and-ai/ models-and-research/gemini-models/gemini-3-8-text-to-speech/ Discussion: https:// news.ycombinator.com/item?id=4 9817615
-
-
background
Google touts #1 benchmark rankings for the new models — Google's blog post says Gemini 3.8 Flash TTS took the #1 overall spot on Hume AI's Voice Design Benchmark and led in accent modeling, while both new models topped Hume AI's Overall Quality Index and ranked highly in blind Voice Arena preference tests across languages including Japanese, Brazilian Portuguese, Vietnamese, MSA, Mexican Spanish and Hindi.
-
2
Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS
Google and Google DeepMind announced two new text-to-speech models offering custom voice design, voice replication with consent verification, scene direction, and support for over 100 languages, available via Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
“Introducing Gemini 3.8 Flash and Flash-Lite TTS, our new SOTA text to speech model with: - a new voice design experience - 2,000+ production ready voices - voice replication - support for 100 languages - voice remixing (soon) - #1 spot on Hume AI's voice benchmarks and more!!”
— @officiallogank -
7 outlets first by Blockchain.News, 14h ago · also The Next Web, Google, Simon Willison, Simon Willison's Weblog, The Decoder, MarkTechPost · read ↗
-
3 outlets Gemini 3.8 text-to-speech
first by HN Best, 14h ago · also HN Frontpage, Google DeepMind
1 more headline
- Gemini 3.8 text-to-speech says hello HN Frontpage · 14h ago
-
What people are saying 19 voices from 3 sites · best of 31 · verbatim
- Today
-
I gave 3.8 a whirl today, replacing Eleven v3 TTS in an internal application that uses TTS to provide a listening function. The Google model produces extremely expressive output. To my ears, it’s on par with Eleven v3, which was already amazing.Nice work. It’s awesome to have these capabilities so close at hand and so trivially easy to integrate…
- Yesterday
-
N
Gemini 3.8 text-to-speech: https:// blog.google/innovation-and-ai/ models-and-research/gemini-models/gemini-3-8-text-to-speech/ Discussion: http:// news.ycombinator.com/item?id=4 9817615
-
An education account I have lists 3.1 Pro, and 3.6 Thinking and Flash as the available models in the app.My personal account, 3.6 Flash Lite and 3.5 Thinking.Meanwhile, I can go hog wild and drain my bank account on GCP. I don’t though, because my family has to eat.
-
Google is spreading too thin, as gemini isn't really that intelligent.They are creating gemini SOTA (not really any more), flash versions, text-to-speech, video (omni), etc.I can see they want to create an ecosystem, but I see no focus in any one area.
-
Pet peeve on Google's AI rollouts: there's no alignment across the three platforms they have, consumer, prosumer, cloud. Scroll to the end of every release, including this one, and you'll see different availabilities. The fun part is the models don't even have the same capabilities across platforms! Omni Flash, last I tried and read the docs, is…
-
I vibe coded a playground UI for trying this out. The conversation mode is neat, and it's very expensive - most of my experiments have cost less than a cent.
-
My primary use case for TTS is converting written content (blogs, articles, etc.) in to clips I can listen to on the go.Is there a good browser extension that does this with a flexible TTS backend? I know Qwen, Kokoro, and VibeVoice all have decent quality..
-
It's giving me an error when I try to generate a voice with Voice Design in AI Studio. It also says voice replication isn't available in my region.Also weird that there are no "neutral gender" voices in the English language. There's also limited "use cases," like the "Gaming" use case is empty?And there's no pricing listed anywhere.I don't know, I…
-
Related, for embedding small models, this lib is incredible.Having a voice under 1Mo is crazy, even if it sounds robotic.
-
I direct my own extended daydream Star Trek fanfic (okay, I'm on season 2 episode 17) and recently I looked to see if I could have each scene file be read aloud a la an audiobook or radio drama.Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough…
-
> Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.I guess voice cloning is widely enough available now from other…
-
Forget the leaderboard wins. The real headline in Google's Gemini 3.8 TTS launch is that cloning a voice from a sample lasting 30 seconds is now a standard developer feature. Google's safeguard is a matching spoken consent recording from the voice owner, plus SynthID watermarks and C2PA credential...
-
We're launching Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS ⚡️ Our most expressive audio models yet let you create custom voices across 100+ languages or pick from 2,000+ ready-to-use ones. You can direct back-and-forth conversations, guide the delivery line-by-line, and add natural cues li...
-
Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Gemini 3.8 Flash TTS debuts at #1 on our Pronunciation Robustness Benchmark and #2 on our Provider Voice Arena Leaderboard Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are @GoogleDeepMind's latest Text to Speech models, ...
-
I was invited to test Gemini 3.8 Flash TTS. This is currently my favorite model. The voice outputs from this model are extraordinary. I never really rated AI voice generation because of the uncanniness associated with low quality. This Gemini IMO crosses that chasm, like Opus 4.5 did with coding,...
-
introducing Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, our most expressive audio generation models yet these models enable creators, developers, and enterprises to create richer, more expressive audio experiences try them via the Gemini API and in AI Studio: https://aistudio.google.com/ ...
-
Create and deploy custom audio with our new text-to-speech models: 🔵 Gemini 3.8 Flash TTS: Design unique voices with distinct accents and characteristics. 🔵 Gemini 3.8 Flash-Lite TTS: Built for efficiency and scale, choose from your created styles or our expansive production-ready library.
-
Rolling out starting today: — Developers: Both models in @GoogleAIStudio and the Gemini API — Consumers: Gemini 3.8 Flash TTS in @Gemini_Notebook and Gemini 3.8 Flash-Lite TTS in Google Vids — Coming soon: Both models in Gemini Enterprise
-
Can you hear that? Our Gemini Audio family is getting louder 🔊 We're introducing two of our most expressive audio generation models yet from @GoogleDeepMind: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.