On September 23, 2026, Google DeepMind announced on its official blog Gemini 3.8 Flash TTS and Flash-Lite TTS, calling them the most expressive audio generation models to date. The new models support designing voices from scratch with natural-language prompts, cloning a voice from a 30-second sample, directing performances line by line, generating long-form audio, and orchestrating two-speaker dialogue — across more than 100 languages.
Start with how they differ from the previous generation of speech models. Old TTS was about "picking a voice": choose one from a preset library, tweak speed and pitch. The new models are about "creating voices" and "directing performances": write "a warm, slightly raspy late-night radio voice" and the model generates it; hand it a 30-second recording and it clones someone's voice; long reads can carry per-line emotion direction, and two-speaker scenes get their turn-taking orchestrated by the model. Flash-Lite is the lighter variant, optimized for low latency and low cost in real-time scenarios.
Where they land, and who benefits first
Both models are available across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Three groups will feel the impact first: podcast and audiobook producers, whose production pipeline maps directly onto long-form generation plus line-by-line direction; video creators, who can swap in custom voices for narration inside Google Vids; and corporate training and support teams, where two-speaker orchestration fits conversational teaching content. Notably, on the same day, Alibaba's Qwen released its Qwen-Audio-3.1 family of speech models (our coverage (/article/2056-a-li-qwen-fa-bu-qwen-audio-3-1-yi-kou-qi-wu-kuan-yu-yin-mo-xing-asr-jiang-jia-9)) — the voice track is clearly accelerating this week.
The stronger the capability, the clearer the boundaries need to be
A 30-second voice clone should first raise the question of risk: cloning someone's voice without consent touches voice and likeness rights in most jurisdictions, so the authorization chain must be confirmed before any commercial use. And "most expressive yet" is the official claim — real quality still needs testing across languages, especially whether rhythm and pauses sound natural in long Chinese texts. Those details usually only surface once developers get their hands on it.