MiniMax H3
by MiniMax
MiniMax's multimodal video model.
2K, native stereo audio, clips up to 15s.

Key features
Technical specifications
5 to 15s
One generation, with multiple shots possible inside it.
Up to 2K
1440-pixel short edge; a 768p mode is announced as coming.
24 FPS
The cadence film and broadcast are shot at.
3 modes
Text-to-video, first and last frame, and omni reference.
Use cases
Brand films and commercials
Trailers, TV spots, and premium brand films, one of the production categories MiniMax builds the model's examples around.
Vertical drama
Short-form drama in 9:16, where dialogue, performance, and shot-reverse-shot coverage carry a 15 second hook.
Product and e-commerce
Product showcases, feature demos, and performance-marketing assets built from a still of the real product.
Motion design and titles
Title sequences, typography, and VFX packages, with the type animation directed as its own layer of the brief.
Game and interface concepts
Game UI, web UI, interaction demos, and feature walkthroughs, timed beat by beat across the length of the clip.
Stylized animation
Claymation, anime promos, and 3D character films, with an identity reference holding the design across shots.
Prompt examples
Overview
MiniMax H3, widely referred to as Hailuo 3.0 or Hailuo 03, is MiniMax's open-weight, general-purpose multimodal video model. It generates 5 to 15 seconds at up to 2K and 24fps, and every generation comes back with native stereo audio. MiniMax positions it as a move away from separate models for generating, editing, and referencing, toward one model that reads text, images, video, and audio as a single context.
What MiniMax H3 brings
The practical change is how much you can hand it. A single generation accepts up to 9 images, 3 video clips, and 3 audio files, 12 files in total, and the model reads them for different things: who the character is, how they move, how the camera behaves, what the edit rhythm sounds like, what voice comes out of their mouth. Sound is generated with the picture rather than after it, so dialogue, effects, and room tone arrive together. And when a clip is close but wrong in one place, you name that element and change it, instead of re-rolling the shot and losing what already worked.
Where MiniMax H3 sits
H3 follows Hailuo 2.3, which produces 1080p clips of roughly ten seconds from a prompt or a single image. H3 raises that to 2K and 15 seconds, but the wider gap is in what it accepts and what it returns: mixed references instead of one image, stereo sound instead of silence, and targeted editing instead of regeneration. MiniMax has said the weights will be released; in the meantime the model runs through the Hailuo AI app, the MiniMax Hub desktop app, the MiniMax Open Platform API, and hosting providers that have picked it up. Aspect ratios run 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, with an auto mode that lets the model choose when references are driving the shot.
MiniMax H3 on Morphic
MiniMax H3 runs on Morphic, alongside Veo, Kling, Seedance, Vidu, and the earlier Hailuo models. Pick it in the prompt bar, write the shot as a brief, and attach the references you want it to hold. Because everything renders onto the Canvas, you can run the same brief through H3 and another model and compare the takes side by side before committing to one.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
1100 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Pro Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
More about MiniMax H3
MiniMax H3: complete guide, references, editing, and audio
The complete MiniMax H3 guide: omni references, native audio, instruction-based editing, and the brief structure that gets the most out of the model.
MiniMax H3 prompt library: video briefs that work
A growing library of MiniMax H3 video prompts with example clips. Each is a full shot brief with reference roles, timed beats, and directed sound.
ChatGPT Images 2.5
OpenAI
OpenAI's image model, released September 2026. Sharper detail, precise edits, faster.
Lyria 3.5
Google DeepMind
Google's best-sounding music model. Full songs with structure, vocals, and lyrics.
MiniMax H3 Max Turbo
fal.ai
The fast tier of fal's post-trained H3. Twice the speed, nearly the same look.
Inworld TTS 2
Inworld
Ninety-five voices, 100+ languages. Speech that starts in under a second.