Google announced on its official blog that it has made important upgrades to the Gemini 2.5 Flash and Gemini 2.5 Pro text-to-speech (TTS) preview models. This update focuses on improving the diversity of emotion and tone, compliance with style instructions, and consistency in multi-character dialogue scenarios, aiming to give developers more granular control over the style and listening feel of synthetic voices.
In terms of rhythm control, the new version can adjust the speed of speech with context-aware speech based on the content of the text, such as automatically slowing down when explaining complex content, speeding up the pace in tense or high-energy paragraphs, and better executing clear rhythm and pause commands. The multi-speaker capability has also been enhanced to maintain the tone, tone, and mood of different characters in scenarios such as podcasts, virtual interviews, and narration, and maintain consistent performance across the 24 languages the model already supports.
Currently, Gemini 2.5 Flash TTS (emphasis on low latency) and 2.5 Pro TTS (emphasis on sound quality and expressiveness) are open to developers in Google AI Studio through the Gemini API, replacing the earlier TTS version released in May this year. The industry expects that this highly controllable TTS technology will further automate audiobooks, online education, localized marketing, and creator content production, but it also places higher requirements for copyright compliance and speech synthesis abuse prevention.
FAQ
Q: What features have been mainly updated in Gemini 2.5 Flash and Pro TTS this time?
A: This update focuses on enhancing emotional and tone diversity, context-aware speed and rhythm control, and vocal consistency and stability during multi-speaker and multi-character conversations.
Q: What use cases are Gemini 2.5 Flash TTS and 2.5 Pro TTS suitable for?
A: Flash TTS is aimed at low-latency interaction scenarios, such as assistant conversations and interactive applications; Pro TTS is more suitable for content production such as audiobooks, plot narration, and marketing videos that pursue high sound quality and strong expressiveness.
Q: What languages and regions is currently supported for Gemini 2.5 TTS?
A: Gemini 2.5 Flash and Pro TTS are available to developers in the Gemini API as preview models, covering 24 supported languages, and specific languages and regions can be viewed in the relevant documentation.
Q: How can developers access the TTS capabilities of this update in their products?
A: Developers can select a 2.5 Flash or 2.5 Pro model with TTS capabilities in Google AI Studio through the Gemini API, configure voice style, tempo, and multi-speaker parameters in the playground or sample code and integrate them into their own applications.
Q: Can the previous Gemini TTS model still be used?
A: Google has made it clear that the new version of Gemini 2.5 Flash and Pro TTS will replace the old TTS model released in May this year, and the subsequent capability iteration will be based on the new model, and it is recommended that developers migrate the configuration and call logic as soon as possible.