Skip to main content
AI Audio Models
Google
Gemini TTS

Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS supports documented multi-speaker dialogue and voice controls, but a reliable maximum output-audio estimate is not yet available. Its details remain available while generation stays disabled.

Online use is currently unavailable

This page is for reviewing the model's capabilities, use cases, and public references. The model does not appear in the home workbench, and you cannot submit generation tasks with it.

Current Status
Blocked by Preauthorization
Input Format
Speakers + Ordered Dialogue
Call Boundary
No Task Creation or Credit Deduction
This is not a Coming Soon model; its preauthorization requirements are incomplete.
Multi-speaker input, ordered dialogue, accent, style, and pacing fields have been verified.
The model does not appear in selectors, create tasks, or deduct credits.
30-second overview
Gemini TTS
Currently unavailable
What it does best
A Gemini TTS candidate designed for faster speech generation and controlled multi-character dialogue.
Best for
People tracking podcast dialogue, narration, and character-voice trends who want to prepare future evaluation scripts in advance.
Popular searches
Is Gemini 3.1 Flash TTS available nowGemini 3.1 Flash TTS multi-speaker voiceGemini 3.1 Flash TTS pricing

At a glance

What is this model like?

This model is neither Coming Soon nor discontinued. The current restriction comes from this site's preauthorization safety boundary: the speaker and dialogue arrays have no total item limit, and there is no provable worst-case conversion from input text to output audio tokens. The model is therefore excluded from selectors, and every request is rejected before any credits are deducted.

What it does best
A Gemini TTS candidate designed for faster speech generation and controlled multi-character dialogue.
Best for
People tracking podcast dialogue, narration, and character-voice trends who want to prepare future evaluation scripts in advance.
Why use it on Kyeo AI
This page publishes the verified capabilities and the precise reason generation is blocked. The model will not enter the workbench or callable interface until a safe preauthorization bound is established.

Key facts

Quickly assess whether this model fits your use case.

Developer
Google
Model Family
Gemini TTS
Verified Input
Speaker Configuration and Ordered Dialogue
Text per Turn
Up to 10,000 Characters
Current Gap
Maximum Output Audio Token Bound

Also known as

The same model may appear under different names across documentation and community discussions; this list keeps them easy to search and compare.

Gemini Flash TTS
Gemini 3.1 Text to Speech

Common questions

These practical questions focus on the task, input conditions, and result requirements you should confirm before choosing.

Is Gemini 3.1 Flash TTS available now
Gemini 3.1 Flash TTS multi-speaker voice
Gemini 3.1 Flash TTS pricing
How to use Gemini TTS
Gemini TTS alternative models

Selection guide

Use these decision points when choosing a model.

1
You need multi-speaker dialogue now

Use ElevenLabs Dialogue V3 first, and test speaker separation with a short script.

2
You need fast single-speaker narration now

Compare ElevenLabs Turbo 2.5 with Multilingual V2 first.

3
You are preparing a future evaluation

Save three non-sensitive scripts: narration, a two-person interview, and a short character scene.

Popular comparisons

Compare common alternatives on the same task to clarify differences in inputs, controls, and cost.

Gemini 3.1 Flash TTS vs ElevenLabs Dialogue V3
Why is Gemini 3.1 Flash TTS temporarily unavailable
Which voice controls does Gemini 3.1 Flash TTS support

Model comparison

Compare the current model with alternatives at a glance.

Practical usage insights

Practical guidance based on public sources, current on-site limits, and representative tasks.

Low latency does not establish a cost ceiling

Even if a model is positioned for speed, missing maximum output audio tokens means preauthorization coverage cannot be guaranteed.

Multi-character input expands worst-case usage

When the arrays have no total item limit, a 10,000-character cap on each turn does not bound the whole request.

Capabilities

Multi-Speaker Configuration

The contract assigns an identifier, voice, and accent to each speaker.

Ordered Dialogue

Dialogue is produced in sequence, with every turn linked to a declared speaker.

Voice Direction

The input can describe the scene, overall tone, audio profile, style, and pacing.

Use cases

Podcast Dialogue Planning

Prepare role-specific scripts for a host and guest.

Character Scene Preparation

Plan sequential dialogue and emotional pacing across multiple characters.

Future Side-by-Side Testing

Once generation opens, compare speaker separation, latency, and actual credit usage.

Prompt tips

Keep Speaker IDs Consistent

Every dialogue turn must reference a speaker that has already been declared.

Validate with a Short Script First

Check proper nouns, pauses, and speaker distinction before expanding the text.

Do Not Estimate Audio Length

Text length alone cannot prove the maximum number of output audio tokens.

Why choose it

Speaker, dialogue, and voice-control fields are clearly defined.
The documented capabilities cover both solo narration and multi-character dialogue.
Public model information remains strictly separated from actual generation.

What to know first

Voice tasks cannot currently be created, and credits cannot be deducted.
Each dialogue turn is capped at 10,000 characters, but the speaker and dialogue arrays have no total item limit.
Input token pricing cannot prove the worst-case cost of output audio tokens.

Currently unavailable

This page is for reviewing the model's capabilities, use cases, and public references. The model does not appear in the home workbench, and you cannot submit generation tasks with it. A Gemini TTS candidate designed for faster speech generation and controlled multi-character dialogue.

No additional settings are available on this page.

FAQ

Related models

Compare these similar candidates before deciding.

ElevenLabs Dialogue V3

ElevenLabs Dialogue V3 provides line-by-line dialogue generation: enter each turn, choose a voice, and pay by total dialogue length. It is not a real-time voice agent, and ElevenLabs notes that users may need multiple generations to find a usable result. Start with a sample to check speaker changes, audio tags, language code, output format, and long-script completeness.

AI audio model
14 credits / 1000 characters

ElevenLabs Turbo 2.5

ElevenLabs Turbo 2.5 is a retained single-speaker TTS workflow on Kyeo AI. It exposes a voice choice plus stability, similarity boost, style, speed, timestamps, surrounding text, and language code. The workflow remains selectable, but ElevenLabs now recommends Flash v2.5 instead. The Turbo name does not make low latency, language support, timestamp accuracy, voice authorization, or output quality a Kyeo guarantee.

AI audio model
6 credits per 1,000 characters

ElevenLabs Multilingual V2

ElevenLabs Multilingual V2 is Kyeo AI's current single-speaker, high-naturalness voiceover page. It shares a similar form with Turbo 2.5, but the task boundary differs: Turbo is a fast voiceover baseline, while Multilingual V2 is aimed at long-form narration, cross-language content, and sustained brand voice. Kyeo still uses a fixed voice list and a 5,000-character entry, so the practical decision is whether the listening experience justifies twice Turbo's per-character cost.

AI audio model
12 credits / 1000 characters

Sources

Content is based on model documentation, feature references, and the settings available on Kyeo AI.

Source note

The speaker, dialogue, voice-control, and token-rate contract for Gemini 3.1 Flash TTS has been verified. However, the maximum number of output audio tokens cannot be proven before a request is created, so generation is not currently available on this site.

Last updated: 2026-08-28
Gemini 3.1 Flash TTS — Generation Temporarily Unavailable
Review