Seed Audio 1.0
by ByteDance
ByteDance's all‑in‑one TTS model.
Voice, music, and sound effects in one generation.

Key features
Hear the range
A warm, measured documentary voice-over.
A hushed, tense line read, close and intimate.
A layered open-air market sound bed.
A rolling storm building to a distant thunderclap.
A short rising cue for strings and brass.
A relaxed beat with soft keys and vinyl crackle.
Technical specifications
ByteDance
Developed by ByteDance's Seed research team.
20
English, Chinese, Japanese, Korean, Spanish, French, German and more.
Up to 2 min
Maximum two minutes of generated audio per pass.
3,000 chars
Room for a full scene brief, dialogue included.
Up to 3
Up to three reference clips, each up to 30 seconds.
Non-streaming
Renders the complete track, not a realtime stream.
Use cases
Audiobooks
Narration, character voices, and sound design for a full book. ByteDance puts the cost near a tenth of studio recording.
Video dubbing
Describe the voice or upload a character image, then use timestamps to land each line exactly where the picture needs it.
Game audio
Character barks, scripted performances, and environmental sound effects for immersive scenes, generated from the script.
One-pass video audio
Give a video clip its narration, sound design, and score in one generation, with no separate mixing step afterward.
Ads and promos
A spoken line, sound effects, and music as one ready-to-use track, made for short-form content.
Dialogue and audio drama
Multiple characters, each with a distinct voice and delivery, in one scene with matching ambience and timing.
Prompt examples
Timestamp control
Ryan (warm, breathless): '[5.5s:8.0s] Maya! Wait, you're leaving tonight?'
Edit promptAudiobook scene
Rain on a library window. Narrator, low and unhurried: 'She read it twice.'
Edit promptSports commentary
Packed stadium, roaring crowd. Commentator, exhilarated: 'OH, HE SCORES!'
Edit promptOverview
Seed Audio 1.0 is ByteDance's all-in-one audio model, built as the next step for text-to-speech rather than a departure from it. Traditional audio workflows use separate tools for voice, music, and sound effects, each generated independently and then mixed. Seed Audio 1.0 produces all three together from a single text prompt, and the output is a complete, ready-to-use audio track.
From text-to-speech to reference-to-audio
Ordinary TTS takes a voice and a block of text. Seed Audio 1.0 takes a scene: the weather and the room, the music underneath, the effects around the action, what each character looks and sounds like, and the lines they speak. Everything you can describe becomes part of the same generation, which is why a prompt can run to 3,000 characters without being padding.
The model understands scene context. A café conversation prompt gets ambient chatter and the right acoustic space, not generic room tone. An action cue gets sharp effects and a dynamic score, not background music pasted over silence. The three layers come out proportioned and mixed for the scene, not assembled from separate exports.
Four modes
The simplest is text-to-speech: pick a voice, enter text, and the model reads it. Voice cloning adds one step, upload a single audio clip and the cloned voice becomes available for that same text-to-speech. The other two modes are where the model separates itself from ordinary TTS.
In text prompt to audio (T2A), every voice comes from your description. Name a character's age, accent, emotion, tone, and speed, and the model casts it.
In text prompt plus audio to audio (TA2A), you upload up to three reference clips of up to 30 seconds each and tag each one to a character in the prompt. The generated voice then follows the recording rather than a description. Reference clips can be uploaded for a single job or kept in an asset library and reused across a series.
Twenty languages
Seed Audio 1.0 generates in English, Chinese, Japanese, Korean, Mexican Spanish, Castilian Spanish, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, and Swedish. Prompt language and script language should match: write the brief in the language the characters speak.
Timing you can specify to the second
Seed Audio 1.0 accepts a timestamp on any line, written inline as [5.5s:8.0s], and fits the delivery to that exact window. For dubbing, this is the difference between audio that roughly fits and audio that lands on the cut. Generation is non-streaming: the model renders the finished track in a single pass, up to two minutes long.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
1100 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Pro Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
ChatGPT Images 2.5
OpenAI
OpenAI's image model, released September 2026. Sharper detail, precise edits, faster.
Lyria 3.5
Google DeepMind
Google's best-sounding music model. Full songs with structure, vocals, and lyrics.
MiniMax H3 Max Turbo
fal.ai
The fast tier of fal's post-trained H3. Twice the speed, nearly the same look.
Inworld TTS 2
Inworld
Ninety-five voices, 100+ languages. Speech that starts in under a second.