Wan 3.0
by Alibaba
Alibaba's latest video model, in public beta.
Native 30s clips from prompts, docs or webpages.

Key features
Features confirmed from Alibaba's Wan 3.0 public-beta launch, August 2026.
Technical specifications
Up to 30s
Native 30 seconds in a single run, with intelligent duration control.
480p to 1080p
Three tiers: 480p, 720p, and 1080p. There is no 4K output.
Public beta
Live on Alibaba Cloud Model Studio and Qwen Cloud as wan3.0-video.
API only
No open weights. Alibaba's open Wan line stops at Wan 2.2.
Use cases
Report to video
Hand Wan 3.0 a deck, spreadsheet, or pdf and it builds a video sequence from the data and text inside it.
Full ad spot in one take
A 30-second clip covers a broadcast spot in a single generation, with no cuts to stitch or match afterwards.
Reference-locked series
Omni-Reference holds a character, product, or set across every shot, so a sequence reads as one piece.
Prompt examples
Reference continuity
The same courier crosses three city blocks, morning haze to midday glare
Edit promptCharacter performance
A violinist finishes a phrase, lowers the bow, and exhales in low light
Edit promptOverview
Wan 3.0 is the latest model in Alibaba Tongyi Lab's Wan video line, launched in public beta on 6 August 2026. Alibaba positions it as a jump "from generating a frame to telling a whole story": native 30-second clips in one pass, audio generated alongside the picture, and an input surface that reaches past prompts into documents, decks, and webpages. It runs on Alibaba Cloud Model Studio and Qwen Cloud as the model wan3.0-video, with per-second API pricing across three resolution tiers.
Video from documents is the headline
The capability that sets Wan 3.0 apart is Omni-Reference. Beyond text, image, audio, and video, it reads structured files, doc, xls, ppt, pdf, txt, key, pages, numbers, and md, along with any webpage URL, and builds a video sequence from what it finds inside. Alibaba frames it as turning "static, text-heavy data into reality-grade video content." Hand it a quarterly deck or a product spec sheet and the output is a finished clip, not a storyboard. No other mainstream video model reads documents this way today.
The long take, in one pass
Thirty seconds in a single continuous run is about double what the line did at Wan 2.7, and the jump changes the unit of work rather than just the number. A full ad spot, a short scene, or a slow reveal fits inside one generation instead of being assembled from several clips that then have to be matched for light, pace, and continuity. Intelligent duration control matches the length to the brief, so a short beat is not padded out with empty seconds.
A long shot also rewards a different kind of prompt. Thirty seconds needs an arc, a beginning, a turn, and an end, so the brief should describe how the frame evolves rather than only what it contains. Length without a turn is just a slow clip, which is the main way long-form generation gets wasted.
Reference control and precision editing
Omni-Reference accepts up to 20 reference assets and holds a character, product, or set consistent across the clip and from shot to shot, which is what makes a sequence read as one piece of work. Labelling each reference in the prompt matters as much as supplying it.
Editing is now part of the same model rather than a separate tool. Wan 3.0 supports instruction-based and reference-based editing: you can select a time interval and regenerate only that part, leaving the rest untouched, and adjust the visuals, the action, and even a character's lines. Chinese coverage described the shift as moving "from gacha to production," a controllable pass rather than endless rerolls.
Rendering, audio, and text
Alibaba calls the visual target "reality-grade rendering," with a clear focus on people: detailed faces and skin, restrained natural emotion, and micro-expressions tied to gestures, alongside portrait customization down to bone structure and eye detail. Audio is generated in the same pass as the picture, so pacing belongs in the prompt: name the sound bed, place the beat, and picture and sound land together. On-screen text has also improved, with long-form text rendering across 12 languages and cleaner charts, formulas, and infographics, though reviewers note audio quality and in-frame text still have room to grow.
Resolution and the 4K myth
Wan 3.0 outputs at 480p, 720p, and 1080p. There is no 4K tier. The "Wan 3.0 does 4K" claims circulating in search results are not from Alibaba; its own model listing prices exactly three tiers and mentions 4K nowhere. Treat 4K, and any "60B parameter" or "Apache 2.0 open weights" spec tables, as fabricated until Alibaba publishes a model card.
On open weights
Wan 3.0 is a closed, API-only model. No weights have appeared on the Wan-AI Hugging Face organization, the Wan-Video GitHub organization, or ModelScope. Alibaba's published open weights stop at Wan 2.2 under Apache 2.0, and the 2.5, 2.6, and 2.7 generations were all commercial API models. For a team with a self-hosting requirement, Wan 2.2 remains the newest downloadable Wan, and that is the detail that decides whether this line fits a self-hosted pipeline at all.
Wan 3.0 and Morphic
Wan 3.0 is not in the Morphic catalog. Morphic runs a broad video lineup today, including Seedance, Kling, and Veo, so you can generate video now and pick the model that fits the shot. The prompt habits above, writing a long take as an arc, labelling every reference, and building pacing into the prompt, carry across all of them, so drafting against them costs nothing.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
1100 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Pro Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
ChatGPT Images 2.5
OpenAI
OpenAI's image model, released September 2026. Sharper detail, precise edits, faster.
Lyria 3.5
Google DeepMind
Google's best-sounding music model. Full songs with structure, vocals, and lyrics.
MiniMax H3 Max Turbo
fal.ai
The fast tier of fal's post-trained H3. Twice the speed, nearly the same look.
Inworld TTS 2
Inworld
Ninety-five voices, 100+ languages. Speech that starts in under a second.