Gemini 3.1 Flash TTS
by Google DeepMind
Google's most expressive text‑to‑speech, with audio tags and multi‑speaker dialogue.

Key features
Technical specifications
Multilingual
Style, pace, and accent control across many languages
Up to 2
Two distinct voices in one multi-speaker generation
Audio tags
Natural-language notes plus inline bracket cues
SynthID
Imperceptible AI-provenance watermark on output
Use cases
Video narration and voice-over
Add natural narration to AI or live-action video, with the tone and pacing set in plain language.
Character dialogue
Voice two-speaker scenes for shorts, games, and explainers, each character with its own voice.
Localized voice-over
Narrate the same script across many languages with native pacing and accent control.
Audiobook and long-form
Keep delivery natural and consistent across long passages of narration.
Explainers and tutorials
Clear, directable narration for product walkthroughs, lessons, and how-tos.
Ad reads and promos
Expressive, on-brand voice reads with the energy and emphasis you direct.
Prompt examples
Warm narration
Say this warmly and slowly, like comforting a child: The storm has passed. You're safe now.
Edit promptOverview
Gemini 3.1 Flash TTS is Google's text-to-speech model, announced on April 15, 2026, and integrated into Morphic as an audio model for speech. It turns text into expressive, natural narration that you direct with plain-language instructions and inline audio tags, voices multi-speaker dialogue, and watermarks every clip with SynthID.
What Gemini 3.1 Flash TTS does differently
Most text-to-speech reads your words back at a fixed delivery. Gemini 3.1 Flash TTS takes direction. Write an instruction in plain language before a line to set the tone, pace, or accent, and drop inline cues in square brackets, like [laughs] or [whispering], exactly where you want them. The model performs the direction instead of reading it aloud, so a single script can shift from a hushed aside to a full-energy read.
Gemini 3.1 Flash TTS on Morphic
On Morphic, Gemini 3.1 Flash TTS sits in the audio model picker for Speech, alongside ElevenLabs. Switch the prompt bar to Audio, choose Speech, pick Gemini 3.1 Flash TTS, then write your script with any direction or tags, choose a voice and language, and generate. The audio drops straight into Canvas next to your video clips.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
1100 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Pro Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
ChatGPT Images 2.5
OpenAI
OpenAI's image model, released September 2026. Sharper detail, precise edits, faster.
Lyria 3.5
Google DeepMind
Google's best-sounding music model. Full songs with structure, vocals, and lyrics.
MiniMax H3 Max Turbo
fal.ai
The fast tier of fal's post-trained H3. Twice the speed, nearly the same look.
Inworld TTS 2
Inworld
Ninety-five voices, 100+ languages. Speech that starts in under a second.