Google

Gemini 3.1 Flash TTS

The Gemini 3.1 Flash TTS Preview model provides powerful, low-latency speech generation with natural outputs, steerable prompts, and new expressive audio tags for precise narration control.

Unknown N/A context 16,384 tokens output

Model Overview

High-signal model metadata in a structured two-column overview table.

Provider

The entity that provides this model.

Google

Input Context Window

The number of tokens supported by the input context window.

N/A tokens

Maximum Output Tokens

The number of tokens that can be generated by the model in a single request.

16,384 tokens tokens

Open Source

Whether the model's code is available for public use.

No

Release Date

When the model was first released.

Unknown

Knowledge Cut-off Date

When the model's knowledge was last updated.

Unknown

API Providers

The providers that offer this model. This is not an exhaustive list.

Google

Modalities

Types of data this model can process.

N/A

Pricing for Gemini 3.1 Flash TTS

Primary API pricing shown in the same “quick compare” spirit as the reference page.

Price Comparison

Additional usage-cost dimensions synced into the project for this model.

maxResponseSize 16,384 tokens

API Access & Providers

Places where this model is available, based on the synced detail-page metadata.

Google

Configuration & Parameters

The configurable options currently documented for this model.

Voice

Select

Prebuilt voice preset to use.

Default: Kore
Zephyr (bright) Puck (upbeat) Charon (informative) Kore (firm) Fenrir (excitable) Leda (youthful) Orus (firm) Aoede (breezy) Callirhoe (easy-going) Autonoe (bright) Enceladus (breathy) Iapetus (clear) Umbriel (easy-going) Algieba (smooth) Despina (smooth) Erinome (clear) Algenib (gravelly) Rasalgethi (informative) Laomedeia (upbeat) Achernar (soft) Alnilam (firm) Schedar (even) Gacrux (mature) Pulcherrima (forward) Achird (friendly) Zubenelgenubi (casual) Vindemiatrix (gentle) Sadachbia (lively) Sadaltager (knowledgeable) Sulafat (warm)

Style Instruction

Prompt

Optional natural-language direction for delivery (e.g. "Say cheerfully:", "Whisper softly:", "Narrate dramatically:"). Prepended to the input before synthesis. Leave blank for a neutral read. You can also embed expressive audio tags directly in your input text like [happy], [whisper], [laughing].

Supported Request Parameters

Parameters currently listed by OpenRouter or the local catalog for this model.

Voice Style Instruction

Compare Gemini 3.1 Flash TTS with related models

Jump straight into the most relevant side-by-side comparison pages for this model.

Related Daily Briefs

Recent daily stories tied to Gemini 3.1 Flash TTS through direct model mentions or provider-level coverage.

Community discussion

What people think about Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS discussions are most active in r/GeminiAI, r/Bard, r/GoogleGeminiAI. The strongest match in this snapshot has 257 upvotes and 50 comments.

r/TextToSpeech 1 upvotes 13 comments April 21, 2026
gemini-3.1-flash-tts-preview is slow?

Hey, I am playing around with the new flash TTS preview and it seems very slow.
Generating TTS for

"It’s a bright, sunny day with clear blue skies stretching across the horizon, and a gentle breeze that keeps the air feeling fresh. The temperature is pleasantly warm, making it comfortable to be outside, whether you’re walking, relaxing, or enjoying time in nature."

takes over **12** seconds, while elevenlabs with a **cloned** voice takes less than **2** seconds.

Am I misinterpreting the "flash" and "low latency" part of the model?

Open Reddit thread
r/GeminiAI 62 upvotes 1 comments April 15, 2026
Google Launches Gemini 3.1 Flash TTS Text-to-Speech Model

Gemini 3.1 Flash TTS introduces audio tags for controlling vocal style, delivery, and pace with natural language commands, scene direction, speaker-level specificity, and more natural expressive voices. The model supports over 70 languages including Hindi, Japanese, and German, with features like SynthID watermarking and multi-speaker audio. It is available in preview via the Gemini API, Google AI Studio, Vertex AI, and rolling out in Google Workspace via Google Vids.

Open Reddit thread
r/aicuriosity 16 upvotes 2 comments April 15, 2026
Google Launches Gemini 3.1 Flash TTS With Audio Tags For Voice Control

Google AI just rolled out Gemini 3.1 Flash TTS, their most expressive text to speech model yet. The real highlight is audio tags, simple commands you drop straight into your text to tweak the voice style, speed or tone on the fly.

Want it to whisper, yell, sound excited, sarcastic or even reflective? Just add the tag and it follows naturally. The demo walks through these shifts in real time, from a calm hello to a laughing multilingual bit that switches languages without missing a beat.

It covers over 70 languages with strong quality in 24 of them including Hindi, Japanese and Arabic. You can try it right now in Google Vids or grab it in preview through the Gemini API and Google AI Studio. Solid option if you create videos, presentations or any narration that needs to feel more alive.

Open Reddit thread
View more discussions →

More models from Google

Continue browsing adjacent models from the same provider.

← All AI Models