Gemini 3.8 text-to-speech logo

Gemini 3.8 text-to-speech

Freemium
訪問

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are audio-generation models for creating expressive voices, directing spoken performances, and producing conversational audio. They support creative teams, developers, and enterprises working across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.

Gemini 3.8 text-to-speechとは?

Gemini 3.8 text-to-speech is a pair of audio-generation models for converting written scripts into expressive speech. Gemini 3.8 Flash TTS focuses on creative voice design and detailed performance direction, while Gemini 3.8 Flash-Lite TTS is optimized for high-volume, cost-efficient generation. Together, they support custom voices, controlled delivery, multi-speaker scenes, and conversational audio workflows.

The models are intended for creators, developers, and enterprises producing audiobooks, podcasts, games, interactive media, dubbing, and voice agents. The supplied source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids as product environments where the models can be used.

Gemini 3.8 text-to-speechでできること

Natural-language voice design

Gemini 3.8 Flash TTS can create bespoke voices from prompts that specify a role, accent, and other vocal characteristics. The source describes support across more than 100 languages and dialects.

Production voice library

Users can select from more than 2,000 production-ready voices, including regional varieties such as Mexican Spanish, Quebec French, and Scots English.

Voice replication with provenance controls

The models can recreate a consistent vocal profile from a 30-second audio sample when the user has rights to use that voice. The workflow includes consent verification, SynthID watermarking, and C2PA credentials as described by Google.

Line-by-line performance direction

Scripts can include stage directions and natural-language cues for controlling pacing, tone, emotion, dialect shifts, and conversational delivery.

Long-form and two-speaker generation

The models are designed to maintain voice quality, pacing, and character timbre across extended audio, and can stage two-speaker scenes with distinct voices and conversational turn-taking.

Scripted vocal sounds and backchanneling

Scripts can include nonverbal cues such as laughs, sighs, and gasps, as well as active-listening interjections such as “mhm” and “yeah”.

利用シーン

“Audiobooks and podcasts”

Use long-form generation and line-level direction to produce narrated chapters, dramatic scenes, or multi-speaker podcast segments while maintaining consistent character voices.

“Games and interactive media”

Create original character voices and direct individual lines for role-specific delivery, regional accents, emotional changes, and reactive dialogue.

“Dubbing and localized audio”

Use the voice library’s language and regional coverage, along with expressive delivery controls, for high-volume dubbing and localized spoken content.

“Voice agents”

Build expressive voice-agent experiences with control over tone, pacing, conversational reactions, and backchanneling rather than relying only on fixed voice presets.

“Creative and productivity workflows”

Use the models through Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, or Google Vids to create custom spoken content and improve audio experiences in those environments.

よくある質問

What is the difference between Gemini 3.8 Flash TTS and Flash-Lite TTS?

Gemini 3.8 Flash TTS is positioned for deep creative direction, custom character voices, and detailed performance control. Gemini 3.8 Flash-Lite TTS is positioned for high-volume, cost-efficient generation, including dubbing, audio content creation, and expressive voice agents.

Can Gemini 3.8 text-to-speech create custom voices?

Yes. Gemini 3.8 Flash TTS can design voices from natural-language prompts by specifying characteristics such as role and accent. The source also describes voice replication from a 30-second sample when the user has the rights to use the voice.

Can the models generate conversations with more than one speaker?

The source describes native two-speaker scene staging from a single script, with distinct voices and natural conversational turn-taking.

What kinds of performance controls are supported?

Users can direct speech line by line with stage directions and script cues for pacing, tone, emotion, dialect shifts, and conversational texture. Scripts can also include sounds such as laughs, sighs, and gasps, plus backchanneling such as “mhm” and “yeah.”

Where are the models available?

The supplied source lists Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. It does not specify access requirements, pricing, quotas, or platform-specific limitations.

クイック情報

Product type
Text-to-speech and expressive audio-generation models
Models
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS
Platforms named by Google
Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids
Voice coverage
More than 100 languages and dialects for generative voice design; 2,000+ production-ready voices
Primary workflows
Audiobooks, podcasts, games, interactive media, dubbing, and voice agents
Safety and provenance features
Consent verification, SynthID watermarking, and C2PA credentials are described for voice replication

Gemini 3.8 text-to-speechの代替品

SpeechifyAI logo

SpeechifyAI

speechify.ai

SpeechifyAI is a developer API for expressive text-to-speech and voice cloning. It provides streaming Simba models, catalog and cloned voices, SSML support, and a free starting tier for building speech into applications.

Voicemaker logo

Voicemaker

voicemaker.in

多言語AI音声の生成・編集・ダウンロード・API自動化に対応するブラウザ型TTSプラットフォーム。

Twine AI Launcher logo

Twine AI Launcher

twinelauncher.com

ChatGPTのような支援をホーム画面で使えるAndroidランチャー兼AIアシスタント。メモやタスク、リマインダー、メッセージ作成に対応。

小艺 logo

小艺

xiaoyi.huawei.com

小艺はHuawei独自開発のAIアシスタントで、質問応答や文章作成、文書読解、コード支援、画像認識に対応します。

Altered logo

Altered

altered.ai

メディア制作とライブ音声向けのAIボイスチェンジャー・音声制作プラットフォーム。

MMaudio logo

MMaudio

mmaudio.net

動画を音声に変換するAI音声生成ツール