Audio generation
Available now

ElevenLabs

by ElevenLabs

Six audio modes in one place.
Voice, music, sound design, dialogue, and a voice changer.

ElevenLabs

Key features

Hear the range

Documentary narrationSpeech

A warm, unhurried read at documentary pace.

0:00
0:12
Song with sung vocalsMusic

Written lyrics, sung back as an indie folk track.

0:00
0:12
Evening ambienceSound effects

Crickets, a far-off dog, one car passing.

0:00
0:12
Two-hander sceneDialogue

Two speakers, distinct voices, one generation.

0:00
0:12

Technical specifications

6

Speech, music, SFX, dialogue, voice changer, transcripts

3,000 chars

Per text-to-speech generation

5s – 5min

Instrumental, or sung from your lyrics

0.5 – 22s

Per generated effect

5,000 chars

Every speaker in a single pass

Up to 5 min

Of source audio per conversion

Use cases

A ribbon of light swelling and narrowing like a spoken waveform

Narration and voice-over

Narrate AI-generated or live-action footage, choosing the voice by gender, age, accent and use case instead of scrolling a list.

Concentric arcs of gold light expanding like a struck chord

Original songs and score

Score a video with an instrumental bed, or write the lyrics and get a finished song sung back in the style you asked for.

Luminous filaments radiating from a single bright impact

Sound design and ambience

Build foley, room tone and weather beds for a cut without licensing anything, and set the exact length each moment needs.

Two ribbons of light, one amber and one slate, braiding together

Podcasts and audio drama

Produce intros, narration segments and full multi-speaker scenes, each character with its own voice and delivery.

A single thread of warm light receding far into the dark

Audiobook narration

Long-form reading with consistent pacing and tone across chapters, generated in passes you can direct chapter by chapter.

One ribbon of light crossing from warm amber into cool silver

Dubbing and transcripts

Re-voice a recorded take without a re-record, and turn any clip into a transcript that marks who is speaking.

Prompt examples

Voice-over

Warm documentary narrator: 'The tide has never once been late.'

0:00
0:12
Edit prompt

Song with vocals

Slow indie folk, brushed drums, one close female vocal singing my lyrics.

0:00
0:12
Edit prompt

Sound design

A heavy oak door dragging open over stone, hinges protesting. 8 seconds.

0:00
0:12
Edit prompt

Ambience

Late evening: crickets shimmering, a distant dog, one car far off. 20s.

0:00
0:12
Edit prompt

Two-hander dialogue

Ana (flat): 'You said nine.' Theo (fast): 'The bridge was closed, I ran.'

0:00
0:12
Edit prompt

Character read

Lighthouse keeper, gravelled and fond: 'Forty winters, this light.'

0:00
0:12
Edit prompt

Overview

ElevenLabs covers the audio side of Morphic end to end. Six modes run on it: speech, music, sound effects, dialogue, the voice changer, and Scribe transcription. That range is the point. A single cut usually needs a narrator, a bed of music, a few effects and sometimes a second character, and all four come from one place, in the same project as the footage.

The controls are what separate it from ordinary text-to-speech. Stability, style, resemblance and speed turn a generation into a direction rather than a lottery. Music takes your lyrics and sings them. Sound effects take a length, so a hit lands on the frame it was written for. And speech to speech means a take you already recorded can carry a different voice without losing the performance you captured.

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

1100 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3625 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6350 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24650 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

What is ElevenLabs on Morphic?
ElevenLabs is the audio engine behind six of Morphic's audio modes: text to speech, text to music, text to sound effects, text to dialogue, speech to speech (the voice changer), and Scribe transcription. You generate audio in the same workspace as your images and video, so nothing has to be exported and re-imported.
What kinds of audio can ElevenLabs generate?
Spoken narration and voice-over, original music that can be instrumental or sung, sound effects and ambience, multi-speaker dialogue scenes, and a converted version of a recording you already have. Scribe also runs the other direction, turning audio or video into a transcript with speaker labels.
Can I control how an ElevenLabs voice performs a line?
Yes. Each text-to-speech generation exposes stability, style, resemblance and speed. Stability trades consistency against expressiveness, style pushes the delivery further from neutral, resemblance controls how closely the output tracks the chosen voice, and speed adjusts pacing. Dialogue generations carry a stability control of their own.
Can ElevenLabs music include lyrics and vocals?
Yes. The music tool has a vocals toggle and a lyrics field. Paste your own lyrics and the model sings them, or leave the field empty and it writes them for you. Tracks run from five seconds to five minutes, so the same tool covers a short sting and a full song.
What does the ElevenLabs voice changer do?
Speech to speech takes a recording and re-performs it in a different voice while keeping the original delivery, pacing and emotion. It accepts up to five minutes of source audio per conversion and works on English and multilingual input. It is voice conversion, not voice cloning: it does not build a reusable copy of a specific person's voice.
How long can ElevenLabs audio be?
Music runs from five seconds to five minutes. Sound effects run from half a second to twenty-two seconds. Speech is bounded by prompt length rather than a clock, at up to 3,000 characters per generation, and a dialogue script can run to 5,000 characters. Longer pieces are generated in passes and assembled.
How does ElevenLabs audio work with AI video on Morphic?
Generate the narration, score and sound design with ElevenLabs, then bring them into Canvas alongside clips from any video model on Morphic. Audio and picture live in the same project, so a cut can be scored and narrated without leaving the workspace.