Wan 3.0: 30 Seconds, All-in-One Reference, One Model
Wan 3.0 is Alibaba's new All-in-One video model: text-to-video, image-to-video, and reference-to-video unified in a single model, up to 30 seconds per take, with up to 10 reference images and reference audio in the mix. It's in the generator below right now: same prompt box, same credits — 480p for volume, 1080p for the final cut.
- What it is
- Alibaba Cloud's next-generation Wan AI video model (released August 2026): an All-in-One architecture that unifies text-to-video, image-to-video, and reference-to-video in one model.
- Video
- Up to 30 seconds per take (double Wan 2.7), at 480p/720p/1080p — audio on or off, same price either way.
- References
- Omni-reference: up to 10 reference images + 5 reference videos + 5 reference audio clips mixed in one pass — character, scene, and voice pinned in one go.
- On Molyin
- Live — generate below right now: text-to-video, image-to-video (first-frame / first-and-last-frame), and reference-to-video, 2–30 seconds, on the same credits as everything else.
Facts on this page verified as of 2026-08-06.
Live on Release Day — Roll Camera
Wan 3.0 is ready below: text-to-video, image-to-video (first-frame / first-and-last-frame), and reference-to-video on a 2-to-30-second slider at 480p/720p/1080p, in the same prompt box on the same credits. Iterate at 480p to save credits, then switch to 1080p for the final cut.
What Is Wan 3.0?
Wan 3.0 is Alibaba Cloud's next-generation Wan video model, and the keyword is All-in-One: last generation, text-to-video, image-to-video, and reference-to-video were three separate models — this generation unifies them into one. You just hand it your assets and the model figures out what to do. Single-take duration doubles from Wan 2.7's 15 seconds to 30, reference images loosen from 5 to 10, and for the first time you can mix in reference audio: a character's face, a scene's mood, a voice's timbre — all locked with one prompt.
Two more changes are quietly practical: output audio goes from always-on to on-or-off (same price) — anyone who wants a clean plate for post no longer has to strip the track afterward. And a new 480p value tier cuts drafting costs in half. The aspect ratio picker gains a clever adaptive default too: leave it unset and the model picks the frame from your assets and intent.
Our approach is the same as always: integrated on launch day. Wan 3.0 is in our AI video generator right now, sharing one credit system with the Wan 2.7 family, Seedance, Kling, and Veo — every model in one model hub. A 30-second film, shot today.
Wan 3.0 at a Glance
One model holds the entire pre-production crew — writing, shooting, voicing, revising — no more switching tools.
30 Seconds in One Pass
Single-take duration doubles from 15 to 30 seconds: longer narrative arcs, more complete action — one prompt shoots the whole take, no stitching segments together.
Omni-Reference, Mixed Feeds
Up to 10 reference images + 5 reference videos + 5 reference audio clips in the same input: images pin the look, videos pin the motion, audio pins the voice — take what you need from each.
Audio On or Off
Turn it on for a finished piece with sound; turn it off for a clean plate you'll score in post — same price, your call. It's the first time the Wan family has offered this switch.
The 480p Value Tier
A new 480p tier: draft, revise, and explore ideas here first at about a quarter of 1080p's cost — then rerun the same prompt at a higher tier for the final cut.
adaptive Smart Framing
The ratio picker gains an adaptive default: leave it unset and the model picks the frame from your reference assets and intent. Need exact control? The five explicit ratios are still there.
First and Last Frames Intact
First frame sets the opening, first-and-last frames set both ends — the signature image-to-video controls carry over in full. You own the start and the finish; the model handles the middle.
Wan 3.0 vs Wan 2.7
Spec by spec — and both columns are live on Molyin. 3.0 pushes long video and omni-reference; 2.7 holds video editing and 1080p value. Pick by need, no agonizing.
| Capability | Wan 2.7 | Wan 3.0 |
|---|---|---|
| Max single-take duration | 15 seconds (10s for reference-to-video) | 30 seconds |
| Reference inputs | ≤5 images + ≤5 videos, ≤5 combined | ≤10 images + ≤5 videos + ≤5 audio |
| Output audio | Always on, no switch | On or off, same price |
| Resolution tiers | 720p / 1080p | 480p / 720p / 1080p (new value tier) |
| Aspect ratio | Five explicit ratios | adaptive smart default + five explicit |
| Video editing & extension | Yes — Wan 2.7 Video Edit + video extension | No — use Wan 2.7 Video Edit for edits |
Specs from official documentation. Molyin currently offers text-to-video, image-to-video (first/last frame), and reference-to-video at 2–30 seconds across 480p/720p/1080p — credits work across both families.
Wan 3.0 Across the Whole Lineup
Every row in this table is live on Molyin right now — including the just-released Wan 3.0 and the Wan 2.7 family it complements. Pick the tier that fits; credits work everywhere.
| Model | Lineup role | Availability | On Molyin |
|---|---|---|---|
| Wan 3.0 | Omni-reference flagship, 30s + three resolution tiers | Released August 2026 | Live — generate on this page |
| Wan 2.7 | Classic workhorse, 1080p value | Generally available | Live |
| Wan 2.7 Video Edit | Video editing specialist (restyle / swap elements) | Generally available | Live |
| MiniMax H3 | Native 2K + always-on audio | Generally available | Live |
| Seedance 2.5 | 30-second single takes, directed by the second | Generally available | Live |
| Kling 3.0 | Rival flagship, named-element references | Generally available | Live |
Availability verified August 6, 2026.
Wan 3.0 FAQ
The questions creators actually ask, answered straight.
Start Creating with Molyin Today
Sign up free and turn your first idea into a cinematic video in minutes.