Alibaba Cloud · Wan 3.0 · Live on Molyin

Wan 3.0: 30 Seconds, All-in-One Reference, One Model

Wan 3.0 is Alibaba's new All-in-One video model: text-to-video, image-to-video, and reference-to-video unified in a single model, up to 30 seconds per take, with up to 10 reference images and reference audio in the mix. It's in the generator below right now: same prompt box, same credits — 480p for volume, 1080p for the final cut.

What it is
Alibaba Cloud's next-generation Wan AI video model (released August 2026): an All-in-One architecture that unifies text-to-video, image-to-video, and reference-to-video in one model.
Video
Up to 30 seconds per take (double Wan 2.7), at 480p/720p/1080p — audio on or off, same price either way.
References
Omni-reference: up to 10 reference images + 5 reference videos + 5 reference audio clips mixed in one pass — character, scene, and voice pinned in one go.
On Molyin
Live — generate below right now: text-to-video, image-to-video (first-frame / first-and-last-frame), and reference-to-video, 2–30 seconds, on the same credits as everything else.

Facts on this page verified as of 2026-08-06.

Live on Release Day — Roll Camera

Wan 3.0 is ready below: text-to-video, image-to-video (first-frame / first-and-last-frame), and reference-to-video on a 2-to-30-second slider at 480p/720p/1080p, in the same prompt box on the same credits. Iterate at 480p to save credits, then switch to 1080p for the final cut.

What Is Wan 3.0?

Wan 3.0 is Alibaba Cloud's next-generation Wan video model, and the keyword is All-in-One: last generation, text-to-video, image-to-video, and reference-to-video were three separate models — this generation unifies them into one. You just hand it your assets and the model figures out what to do. Single-take duration doubles from Wan 2.7's 15 seconds to 30, reference images loosen from 5 to 10, and for the first time you can mix in reference audio: a character's face, a scene's mood, a voice's timbre — all locked with one prompt.

Two more changes are quietly practical: output audio goes from always-on to on-or-off (same price) — anyone who wants a clean plate for post no longer has to strip the track afterward. And a new 480p value tier cuts drafting costs in half. The aspect ratio picker gains a clever adaptive default too: leave it unset and the model picks the frame from your assets and intent.

Our approach is the same as always: integrated on launch day. Wan 3.0 is in our AI video generator right now, sharing one credit system with the Wan 2.7 family, Seedance, Kling, and Veo — every model in one model hub. A 30-second film, shot today.

Wan 3.0 at a Glance

One model holds the entire pre-production crew — writing, shooting, voicing, revising — no more switching tools.

30 Seconds in One Pass

Single-take duration doubles from 15 to 30 seconds: longer narrative arcs, more complete action — one prompt shoots the whole take, no stitching segments together.

Omni-Reference, Mixed Feeds

Up to 10 reference images + 5 reference videos + 5 reference audio clips in the same input: images pin the look, videos pin the motion, audio pins the voice — take what you need from each.

Audio On or Off

Turn it on for a finished piece with sound; turn it off for a clean plate you'll score in post — same price, your call. It's the first time the Wan family has offered this switch.

The 480p Value Tier

A new 480p tier: draft, revise, and explore ideas here first at about a quarter of 1080p's cost — then rerun the same prompt at a higher tier for the final cut.

adaptive Smart Framing

The ratio picker gains an adaptive default: leave it unset and the model picks the frame from your reference assets and intent. Need exact control? The five explicit ratios are still there.

First and Last Frames Intact

First frame sets the opening, first-and-last frames set both ends — the signature image-to-video controls carry over in full. You own the start and the finish; the model handles the middle.

Wan 3.0 vs Wan 2.7

Spec by spec — and both columns are live on Molyin. 3.0 pushes long video and omni-reference; 2.7 holds video editing and 1080p value. Pick by need, no agonizing.

CapabilityWan 2.7Wan 3.0
Max single-take duration15 seconds (10s for reference-to-video)30 seconds
Reference inputs≤5 images + ≤5 videos, ≤5 combined≤10 images + ≤5 videos + ≤5 audio
Output audioAlways on, no switchOn or off, same price
Resolution tiers720p / 1080p480p / 720p / 1080p (new value tier)
Aspect ratioFive explicit ratiosadaptive smart default + five explicit
Video editing & extensionYes — Wan 2.7 Video Edit + video extensionNo — use Wan 2.7 Video Edit for edits

Specs from official documentation. Molyin currently offers text-to-video, image-to-video (first/last frame), and reference-to-video at 2–30 seconds across 480p/720p/1080p — credits work across both families.

Wan 3.0 Across the Whole Lineup

Every row in this table is live on Molyin right now — including the just-released Wan 3.0 and the Wan 2.7 family it complements. Pick the tier that fits; credits work everywhere.

ModelLineup roleAvailabilityOn Molyin
Wan 3.0Omni-reference flagship, 30s + three resolution tiersReleased August 2026Live — generate on this page
Wan 2.7Classic workhorse, 1080p valueGenerally availableLive
Wan 2.7 Video EditVideo editing specialist (restyle / swap elements)Generally availableLive
MiniMax H3Native 2K + always-on audioGenerally availableLive
Seedance 2.530-second single takes, directed by the secondGenerally availableLive
Kling 3.0Rival flagship, named-element referencesGenerally availableLive

Availability verified August 6, 2026.

Wan 3.0 FAQ

The questions creators actually ask, answered straight.

Start Creating with Molyin Today

Sign up free and turn your first idea into a cinematic video in minutes.