Audio generation

Seed Audio 1.0

by ByteDance

ByteDance's all‑in‑one TTS model.
Voice, music, and sound effects in one generation.

Seed Audio 1.0

Key features

Hear the range

Documentary narrationSpeech

A warm, measured documentary voice-over.

0:00
0:12
Thriller voice-overSpeech

A hushed, tense line read, close and intimate.

0:00
0:12
Spice-market ambienceSound effects

A layered open-air market sound bed.

0:00
0:12
ThunderstormSound effects

A rolling storm building to a distant thunderclap.

0:00
0:12
Orchestral cueMusic

A short rising cue for strings and brass.

0:00
0:12
Lo-fi beatMusic

A relaxed beat with soft keys and vinyl crackle.

0:00
0:12

Technical specifications

ByteDance

Developed by ByteDance's Seed research team.

20

English, Chinese, Japanese, Korean, Spanish, French, German and more.

Up to 2 min

Maximum two minutes of generated audio per pass.

3,000 chars

Room for a full scene brief, dialogue included.

Up to 3

Up to three reference clips, each up to 30 seconds.

Non-streaming

Renders the complete track, not a realtime stream.

Use cases

Audiobooks

Narration, character voices, and sound design for a full book. ByteDance puts the cost near a tenth of studio recording.

Video dubbing

Describe the voice or upload a character image, then use timestamps to land each line exactly where the picture needs it.

Game audio

Character barks, scripted performances, and environmental sound effects for immersive scenes, generated from the script.

One-pass video audio

Give a video clip its narration, sound design, and score in one generation, with no separate mixing step afterward.

Ads and promos

A spoken line, sound effects, and music as one ready-to-use track, made for short-form content.

Dialogue and audio drama

Multiple characters, each with a distinct voice and delivery, in one scene with matching ambience and timing.

Prompt examples

Narrated explainer

Calm narrator, soft kitchen ambience: 'Combine the flour and butter.'

Edit prompt

Timestamp control

Ryan (warm, breathless): '[5.5s:8.0s] Maya! Wait, you're leaving tonight?'

Edit prompt

Audiobook scene

Rain on a library window. Narrator, low and unhurried: 'She read it twice.'

Edit prompt

Sports commentary

Packed stadium, roaring crowd. Commentator, exhilarated: 'OH, HE SCORES!'

Edit prompt

Audio drama

Detective, tense: 'Don't move.' Footsteps stop, a door creaks, a siren.

Edit prompt

Game moment

Deep narrator: 'The ancient seal has broken.' Stone grinding, a dark hum.

Edit prompt

Overview

Seed Audio 1.0 is ByteDance's all-in-one audio model, built as the next step for text-to-speech rather than a departure from it. Traditional audio workflows use separate tools for voice, music, and sound effects, each generated independently and then mixed. Seed Audio 1.0 produces all three together from a single text prompt, and the output is a complete, ready-to-use audio track.

From text-to-speech to reference-to-audio

Ordinary TTS takes a voice and a block of text. Seed Audio 1.0 takes a scene: the weather and the room, the music underneath, the effects around the action, what each character looks and sounds like, and the lines they speak. Everything you can describe becomes part of the same generation, which is why a prompt can run to 3,000 characters without being padding.

The model understands scene context. A café conversation prompt gets ambient chatter and the right acoustic space, not generic room tone. An action cue gets sharp effects and a dynamic score, not background music pasted over silence. The three layers come out proportioned and mixed for the scene, not assembled from separate exports.

Four modes

The simplest is text-to-speech: pick a voice, enter text, and the model reads it. Voice cloning adds one step, upload a single audio clip and the cloned voice becomes available for that same text-to-speech. The other two modes are where the model separates itself from ordinary TTS.

In text prompt to audio (T2A), every voice comes from your description. Name a character's age, accent, emotion, tone, and speed, and the model casts it.

In text prompt plus audio to audio (TA2A), you upload up to three reference clips of up to 30 seconds each and tag each one to a character in the prompt. The generated voice then follows the recording rather than a description. Reference clips can be uploaded for a single job or kept in an asset library and reused across a series.

Twenty languages

Seed Audio 1.0 generates in English, Chinese, Japanese, Korean, Mexican Spanish, Castilian Spanish, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, and Swedish. Prompt language and script language should match: write the brief in the language the characters speak.

Timing you can specify to the second

Seed Audio 1.0 accepts a timestamp on any line, written inline as [5.5s:8.0s], and fits the delivery to that exact window. For dubbing, this is the difference between audio that roughly fits and audio that lands on the cut. Generation is non-streaming: the model renders the finished track in a single pass, up to two minutes long.

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

1100 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3625 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6350 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24650 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is Seed Audio 1.0?
Seed Audio 1.0 is ByteDance's all-in-one text-to-speech model. From one text prompt it produces voice, instrumental music, and sound effects together as a finished, mixed track. It works in two headline modes: text prompt to audio (T2A), where everything comes from your description, and text prompt plus audio to audio (TA2A), where you add reference clips to cast specific voices.
What languages does Seed Audio 1.0 support?
Seed Audio 1.0 supports 20 languages: English, Chinese, Japanese, Korean, Mexican Spanish, Castilian Spanish, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, and Swedish. For best results, write the prompt in the same language as the lines you want spoken.
How does voice reference work in Seed Audio 1.0?
In TA2A mode you supply up to three reference clips of up to 30 seconds each, then tag them in the prompt so each character maps to a recording. The model takes the vocal character and emotion from the reference and carries it across the generation. You can also define a voice from a text description or a character image instead of a recording.
Can Seed Audio 1.0 clone a voice?
Yes. Alongside T2A and TA2A there is a voice cloning mode: upload a single audio clip, and the cloned voice becomes available for straight text-to-speech. ByteDance documents it as a one-clip clone. When the voice has to sit inside a full scene with music, effects, and other characters, use TA2A and its three reference clips instead.
Can Seed Audio 1.0 control the timing of each line?
Yes. Seed Audio 1.0 supports accurate time control: put a timestamp such as [5.5s:8.0s] at the start of a line and the model fits that line to the exact window. This is what makes it usable for dubbing, where dialogue has to match picture.
Can Seed Audio 1.0 generate multiple speakers at once?
Yes. Write a scene with several characters and describe each voice inline. Seed Audio 1.0 gives each speaker a distinct voice, emotion, and pacing in a single generation, along with the ambience and effects around them.
How long can a Seed Audio 1.0 generation be?
Seed Audio 1.0 generates up to two minutes of audio in a single pass, from a prompt of up to 3,000 characters. Longer productions are built by generating scene by scene.
How is Seed Audio 1.0 different from ordinary text-to-speech?
Ordinary text-to-speech picks a voice and reads text aloud. Seed Audio 1.0 goes from text-to-speech to reference-to-audio: one prompt describes the environment, the score, the sound effects, and every character's voice, and the model returns the whole scene mixed together. The difference is scope, a finished audio production versus only the voice.