Natural-language voice design
Gemini 3.8 Flash TTS can create bespoke voices from prompts that specify a role, accent, and other vocal characteristics. The source describes support across more than 100 languages and dialects.
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are audio-generation models for creating expressive voices, directing spoken performances, and producing conversational audio. They support creative teams, developers, and enterprises working across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
Gemini 3.8 text-to-speech is a pair of audio-generation models for converting written scripts into expressive speech. Gemini 3.8 Flash TTS focuses on creative voice design and detailed performance direction, while Gemini 3.8 Flash-Lite TTS is optimized for high-volume, cost-efficient generation. Together, they support custom voices, controlled delivery, multi-speaker scenes, and conversational audio workflows.
The models are intended for creators, developers, and enterprises producing audiobooks, podcasts, games, interactive media, dubbing, and voice agents. The supplied source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids as product environments where the models can be used.
Gemini 3.8 Flash TTS can create bespoke voices from prompts that specify a role, accent, and other vocal characteristics. The source describes support across more than 100 languages and dialects.
Users can select from more than 2,000 production-ready voices, including regional varieties such as Mexican Spanish, Quebec French, and Scots English.
The models can recreate a consistent vocal profile from a 30-second audio sample when the user has rights to use that voice. The workflow includes consent verification, SynthID watermarking, and C2PA credentials as described by Google.
Scripts can include stage directions and natural-language cues for controlling pacing, tone, emotion, dialect shifts, and conversational delivery.
The models are designed to maintain voice quality, pacing, and character timbre across extended audio, and can stage two-speaker scenes with distinct voices and conversational turn-taking.
Scripts can include nonverbal cues such as laughs, sighs, and gasps, as well as active-listening interjections such as “mhm” and “yeah”.
Use long-form generation and line-level direction to produce narrated chapters, dramatic scenes, or multi-speaker podcast segments while maintaining consistent character voices.
Create original character voices and direct individual lines for role-specific delivery, regional accents, emotional changes, and reactive dialogue.
Use the voice library’s language and regional coverage, along with expressive delivery controls, for high-volume dubbing and localized spoken content.
Build expressive voice-agent experiences with control over tone, pacing, conversational reactions, and backchanneling rather than relying only on fixed voice presets.
Use the models through Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, or Google Vids to create custom spoken content and improve audio experiences in those environments.
Gemini 3.8 Flash TTS is positioned for deep creative direction, custom character voices, and detailed performance control. Gemini 3.8 Flash-Lite TTS is positioned for high-volume, cost-efficient generation, including dubbing, audio content creation, and expressive voice agents.
Yes. Gemini 3.8 Flash TTS can design voices from natural-language prompts by specifying characteristics such as role and accent. The source also describes voice replication from a 30-second sample when the user has the rights to use the voice.
The source describes native two-speaker scene staging from a single script, with distinct voices and natural conversational turn-taking.
Users can direct speech line by line with stage directions and script cues for pacing, tone, emotion, dialect shifts, and conversational texture. Scripts can also include sounds such as laughs, sighs, and gasps, plus backchanneling such as “mhm” and “yeah.”
The supplied source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. It does not specify access requirements, pricing, quotas, or platform-specific limitations.
speechify.ai
SpeechifyAI is a developer API for expressive text-to-speech and voice cloning. It provides streaming Simba models, catalog and cloned voices, SSML support, and a free starting tier for building speech into applications.
voicemaker.in
Voicemaker is a browser-based text-to-speech platform for creating and fine-tuning AI voice output in multiple languages and formats.
twinelauncher.com
Android launcher and AI assistant with ChatGPT-style help, notes, tasks, reminders, and message drafting.
xiaoyi.huawei.com
小艺 is Huawei’s AI assistant for Q&A, writing, documents, coding, and image recognition.
altered.ai
AI voice changer and voice content creation platform for media and live use.
mmaudio.net
AI voice generation for turning videos into audio