Inworld TTS 2
by Inworld
Ninety‑five voices, 100+ languages.
Speech that starts in under a second.

Key features
Technical specifications
95
Ready-made voices across fourteen native languages
100+
Crosslingual, with one voice identity held across all of them
2,000 chars
Maximum script length per generation
Sub-200ms
Median time to first audio
MP3, 48 kHz
Full-rate audio, ready to drop onto a timeline
2
Inworld TTS 2 and Inworld TTS 2 Flash
Use cases

Interactive characters
Speech that starts fast enough to hold a conversation. The latency is the feature here, not a footnote.

Narration and voice-over
Cast a narrator once and keep the same voice across every cut, pickup, and revision.

Localized voice-over
Run the same script across many languages and keep one voice, so a brand survives translation.

Audiobooks and long-form
Work through a book a passage at a time with the narrator holding steady across chapters.

Game and NPC dialogue
Cast a cast. Ninety-five voices is enough to fill a village without reusing the same one twice.

Explainers and tutorials
Regenerate a single line when the product changes, instead of rebooking a session.
Prompt examples






Overview
Inworld TTS 2 is Inworld's Realtime TTS-2 text-to-speech model, released on August 31, 2026 and available on Morphic the same day. It reads a script in one of ninety-five ready-made voices, holds that voice across more than 100 languages, and returns audio fast enough to sit inside a live conversation.
What Inworld TTS 2 does differently
Two things separate it from the rest of the speech shelf. The first is latency: median time to first audio is under 200 milliseconds, which moves text-to-speech out of the render-and-wait category and into interactive use, where a character has to answer while the player is still listening.
The second is that a voice is not tied to a language. Any voice in the catalog can speak any supported language, and the identity holds across a switch inside a single generation. A bilingual script comes back sounding like one speaker rather than two takes stitched together, which is what makes a single brand voice survive localization.
Inworld TTS 2 on Morphic
On Morphic, Inworld TTS 2 sits in the audio model picker for Speech, alongside ElevenLabs and Gemini 3.1 Flash TTS. Switch the prompt bar to Audio, choose Speech, pick Inworld TTS 2, choose a voice, paste your script, and generate. Scripts run up to 2,000 characters per generation and come back as 48 kHz MP3, straight onto your Canvas.
Two tiers are available. Inworld TTS 2 is the default. Inworld TTS 2 Flash draws on the same ninety-five-voice catalog at a faster, lower-cost setting, which suits volume work and drafts. Switching between them does not mean recasting the voice.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
1100 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Pro Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
ChatGPT Images 2.5
OpenAI
OpenAI's image model, released September 2026. Sharper detail, precise edits, faster.
Lyria 3.5
Google DeepMind
Google's best-sounding music model. Full songs with structure, vocals, and lyrics.
MiniMax H3 Max Turbo
fal.ai
The fast tier of fal's post-trained H3. Twice the speed, nearly the same look.
Gemini Omni Flash 1.1
Google DeepMind
Google's video model where the edit is a sentence. Now with 4K output and 40-second scenes.