Google launches Gemini 3.8 Flash TTS and Flash‑Lite TTS for richer voice output
The new models promise more natural, controllable voices at scale, topping third‑party quality benchmarks.
Google has expanded its Gemini lineup with two new text-to-speech models, Gemini 3.8 Flash TTS and 3.8 Flash‑Lite TTS, plus a Live Avatar upgrade for Gemini 3.8 Live aimed at richer multimodal experiences. Flash TTS targets highly directed, character-style audio with custom voices created via prompts or short reference clips, while Flash‑Lite is tuned for cheaper, large‑scale use cases like dubbing and voice agents. Flash TTS can generate or replicate voices across more than 100 languages, offers access to a library of over 2,000 ready-made voices, and supports long-form, multi-speaker scripts with fine-grained control over delivery and non-verbal cues. Google says Flash and Flash‑Lite lead third‑party benchmarks such as Hume AI’s Voice Design Benchmark and Overall Quality Index, and perform strongly in blind human preference tests on Voice Arena in several major languages. All of these audio and Live Avatar outputs are watermarked with SynthID, and voice cloning requires explicit consent from the voice owner, with the new capabilities available through Google AI Studio and Gemini Enterprise APIs.
Why it matters
For teams working with synthetic voices, this raises the ceiling on both quality and control. Flash TTS can be steered like a human performer, from accent and character to non‑verbal cues, while keeping long recordings stable across hours of audio. Paired with Flash‑Lite, which is tuned for cheaper high‑volume output, it gives media and product builders options across premium and large‑scale use cases, backed by strong scores on independent benchmarks and quality indices.
Signal or noise?
Does this story matter, or is it hype? Decide before you see what everyone else thinks.