Natural-language voice design
Gemini 3.8 Flash TTS can create bespoke voices from prompts that specify a role, accent, and other vocal characteristics. The source describes support across more than 100 languages and dialects.
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are audio-generation models for creating expressive voices, directing spoken performances, and producing conversational audio. They support creative teams, developers, and enterprises working across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
Gemini 3.8 text-to-speech is a pair of audio-generation models for converting written scripts into expressive speech. Gemini 3.8 Flash TTS focuses on creative voice design and detailed performance direction, while Gemini 3.8 Flash-Lite TTS is optimized for high-volume, cost-efficient generation. Together, they support custom voices, controlled delivery, multi-speaker scenes, and conversational audio workflows.
The models are intended for creators, developers, and enterprises producing audiobooks, podcasts, games, interactive media, dubbing, and voice agents. The supplied source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids as product environments where the models can be used.
Gemini 3.8 Flash TTS can create bespoke voices from prompts that specify a role, accent, and other vocal characteristics. The source describes support across more than 100 languages and dialects.
Users can select from more than 2,000 production-ready voices, including regional varieties such as Mexican Spanish, Quebec French, and Scots English.
The models can recreate a consistent vocal profile from a 30-second audio sample when the user has rights to use that voice. The workflow includes consent verification, SynthID watermarking, and C2PA credentials as described by Google.
Scripts can include stage directions and natural-language cues for controlling pacing, tone, emotion, dialect shifts, and conversational delivery.
The models are designed to maintain voice quality, pacing, and character timbre across extended audio, and can stage two-speaker scenes with distinct voices and conversational turn-taking.
Scripts can include nonverbal cues such as laughs, sighs, and gasps, as well as active-listening interjections such as “mhm” and “yeah”.
Use long-form generation and line-level direction to produce narrated chapters, dramatic scenes, or multi-speaker podcast segments while maintaining consistent character voices.
Create original character voices and direct individual lines for role-specific delivery, regional accents, emotional changes, and reactive dialogue.
Use the voice library’s language and regional coverage, along with expressive delivery controls, for high-volume dubbing and localized spoken content.
Build expressive voice-agent experiences with control over tone, pacing, conversational reactions, and backchanneling rather than relying only on fixed voice presets.
Use the models through Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, or Google Vids to create custom spoken content and improve audio experiences in those environments.
Gemini 3.8 Flash TTS is positioned for deep creative direction, custom character voices, and detailed performance control. Gemini 3.8 Flash-Lite TTS is positioned for high-volume, cost-efficient generation, including dubbing, audio content creation, and expressive voice agents.
Yes. Gemini 3.8 Flash TTS can design voices from natural-language prompts by specifying characteristics such as role and accent. The source also describes voice replication from a 30-second sample when the user has the rights to use the voice.
The source describes native two-speaker scene staging from a single script, with distinct voices and natural conversational turn-taking.
Users can direct speech line by line with stage directions and script cues for pacing, tone, emotion, dialect shifts, and conversational texture. Scripts can also include sounds such as laughs, sighs, and gasps, plus backchanneling such as “mhm” and “yeah.”
The supplied source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. It does not specify access requirements, pricing, quotas, or platform-specific limitations.
speechify.ai
SpeechifyAI is a developer API for expressive text-to-speech and voice cloning. It provides streaming Simba models, catalog and cloned voices, SSML support, and a free starting tier for building speech into applications.
voicemaker.in
基于浏览器的文本转语音平台,支持多语言 AI 语音、编辑、下载和 API 自动化。
twinelauncher.com
适用于 Android 的启动器与 AI 助手,提供类似 ChatGPT 的帮助、笔记、任务、提醒和消息撰写功能。
xiaoyi.huawei.com
小艺是华为自主研发的 AI 智慧助手,支持问答、写作、文档阅读、代码辅助和识图。
altered.ai
面向媒体制作和实时语音应用的 AI 变声与语音内容创作平台。
mmaudio.net
将视频转换为音频的 AI 语音生成工具