01 Core Formula
Wan 3.0 understands structured instructions: who, what they do, how the camera moves, where, and in what tone. This chapter builds the prompt skeleton first, then walks through two new 3.0 dials — the audio toggle and the adaptive smart ratio.
1.1 The Basic Prompt Formula
Everything starts with one sentence: subject plus motion are the two required parts; camera, environment and aesthetics stack on demand. Duration is a continuous 2-30s dial — plan exactly as much as the seconds can hold; for 30-second long-form writing, see chapter 4.
[Subject] + [Motion] + optional [Camera] [Environment] [Aesthetics]
Subject
The protagonist — a person, animal or product. The more specific, the better.
Motion
What the subject does, with body parts and pace spelled out. Slow, coherent small movements are the most reliable.
Camera
Shot size and movement: close-up, push-in, pan. For multi-shot narratives, write the storyboard directly, e.g. "Shot 1 [0-3s] wide shot".
Environment
Scene and atmosphere: place, time, light, weather.
Aesthetics
Art direction and quality: color tone, texture, clarity.
The Bookstore on a Rainy Corner
Prompt
On a rainy night at a street corner, a small bookstore still glows with warm yellow light. The owner, wearing a beige sweater, carries a stack of books onto the windowsill, glances up at the rain outside, and smiles as she hangs out a wooden "Open" sign. The camera slowly pushes in from across the street to a close-up of the window, raindrops sliding down the glass, neon signs casting reflections on the wet pavement. Cinematic quality, warm-cool contrast color grading, rich detail.
Two sentences for subject+motion, one each for camera and environment, style words to close — that's a solid Wan 3.0 baseline prompt.
1.2 The Audio Toggle: Finished Cut or Clean Plate
The Wan family's audio finally has a switch (same price either way): turn it on and ambience, sound effects and score are generated alongside the visuals — what you get is a finished cut; turn it off and you get a clean plate with no audio track, ready for your own music and post-production. Want a "finished cut"? Switch on. Want "raw footage"? Switch off — for commercial jobs that need voiceover, subtitles or a swapped BGM, switching off is the cleanest path.
Silent B-Roll Footage
Prompt
A bamboo forest in the early morning, thin mist drifting slowly between the trunks, sunlight slanting through the leaves to form shafts of light, a few bamboo leaves falling in the wind. The camera pans slowly at a low angle, the frame quiet and ethereal, documentary quality, teal-green tones.
The audio toggle is switched off for you — take this as background footage and lay your own music and narration over it, with no model-imposed BGM to wash out.
1.3 The adaptive Smart Ratio
The ratio selector gains a new adaptive default: leave it unset and the model picks the canvas from your reference media and content intent — vertical source material yields vertical output, cinematic copy yields landscape, and most of the time it chooses better than a human guess. Need precise control (say, a locked 9:16 for feed placement)? The five explicit ratios are still there. When in doubt, hand it to adaptive; when there's a hard requirement, set it explicitly.
Scented Candle Teaser
Prompt
A handmade scented candle burns quietly on a natural wood table, the flame flickering gently, the amber wax glowing softly, with blurred linen cloth and dried flowers in the background. The camera slowly orbits the candle halfway, light spots shifting across the tabletop. Minimalist premium feel, warm tones, commercial-ad quality.
Pre-filled with adaptive — the model picks the right canvas from the "commercial-ad quality" intent; switch to an explicit ratio anytime you need to lock it.
02 Image to Video
One first-frame image plus a motion description, and Wan 3.0 brings the picture to life — the prompt's center of gravity shifts from describing the scene to describing the motion, because the scene already exists. A first frame sets the opening; first-last frame sets both ends. This chapter covers both plays.
2.1 Bringing Photos to Life
An image-to-video prompt covers only three things: what moves, how it moves, and whether the camera moves. Don't force things that aren't in the picture; do "wake up" things that are — fluttering fabric, flowing water, flickering lamplight are all high-hit-rate small motions. Add "keep the original texture" and the style won't drift.
Lights of the Fishing Harbor
Prompt
Bring this photo to life: at the fishing harbor under the night sky, the lamps on the fishing boats light up one after another, their reflections swaying across the water, the flags on the masts fluttering gently in the sea breeze, and the distant lighthouse beam sweeping slowly past. The camera stays fixed, keeping the blue-toned film texture of the original photo throughout.
Upload 1 first-frame image; write only the motion without restating the scene, and let "keep the original texture" hold the style.
2.2 First-Last Frame: Own the Start and the End
Upload a first frame and a last frame, and the model generates everything in between. This changes how you prompt: skip the start and the end (they're in the images) and write only "how to get from A to B" — the pace of change and how the camera plays along. Keep both images at the same ratio, and the bigger the visual gap, the more process description you should provide.
One Tree, Four Seasons
Prompt
The same tree slowly transforms from a spring canopy in full bloom: petals drift away on the wind, green leaves grow dense into shade, then turn yellow and red, until only snow-covered branches remain. The camera slowly pulls back, the light and shadow of the four seasons flowing naturally, with the tree's position and framing strictly following the first and last frames.
In image-to-video mode, upload a first frame (spring) + a last frame (winter); 8 seconds gives the four seasons room to turn.
03 Omni References
Wan 3.0's headline capability: mix three channels of reference media — up to 10 images + 5 videos + 5 audio tracks. Images own appearance, videos own motion, audio owns voice. Name them in your prompt as "Image 1", "Video 1", "Audio 1" in upload order to assign their roles (the three media types are numbered independently).
3.1 The Multi-Image Reference Formula
A reference-to-video prompt is three beats: bring the media on stage with "Image N", describe what happens in the new scene, then close with "keep it consistent". Multiple angles of the same subject work best; put only one subject per reference image — collages leave the model unable to tell who's who.
[References: Image 1…] + [Scene description] + keep [Subject] consistent
References
Up to 10 reference images, referenced as "Image N" in upload order.
Scene Description
What happens in the new scene: action, environment, camera — the frame is generated from scratch, all driven by your words.
Consistency
A closing "keep it consistent with the reference images" reinforces the appearance lock.
The Barista's New-Product Day
Prompt
The barista from Image 1, wearing the dark green apron from Image 2, stands behind the minimalist coffee bar from Image 3 and gently pushes a pour-over coffee toward the camera, the latte art on the foam clearly visible, and looks up with a smile. The camera slowly pulls back from a close-up of the cup to a medium shot, warm-toned morning light, a cozy and healing everyday mood. The person, outfit and setting stay consistent with the reference images.
Upload 3 reference images (person / outfit / setting); "Image 1", "Image 2" and "Image 3" each name their own role, and the closing "stay consistent" locks the appearance.
3.2 Reference Videos: Migrate Motion and Camera Work
Reference videos own "how things move": the model learns the character's action rhythm and the camera's movement language from your video and applies them to your new scene. Three hard rules to remember: up to 5 reference videos with a combined length of no more than 15 seconds; the reference videos' total seconds count against the 30-second generation budget (10 seconds of reference leaves roughly 20 seconds of output); and when used as a subject reference, keep a single subject per clip.
Borrow the Camera Work from a Film
Prompt
Follow the camera movement and shot rhythm of Video 1: the camera rises from a low angle at desk level, orbits the subject halfway, then pulls back and settles. Swap in the vintage typewriter from Image 1, struck by an invisible pair of hands on a desk late at night, the paper trembling softly with each keystroke, the desk lamp glowing warm. The camera work strictly follows Video 1, the subject's appearance stays consistent with Image 1, cinematic quality.
Upload 1 reference video + 1 subject image; "follow the camera movement and shot rhythm of Video 1" is the standard phrasing for migrating camera language.
3.3 Reference Audio: Let the Character Speak in Your Voice
Reference audio is the Wan family's new channel: upload a clean voice recording (up to 5 tracks, no more than 15 seconds combined) and the model clones its timbre, letting the on-screen character speak your written lines in that voice. Name "Audio 1" in the prompt and add one line of tone direction; lip movements sync with the pronunciation automatically.
The Founder's Opening Line
Prompt
The brand founder from Image 1 sits in the natural-wood studio from Image 2, a few handmade leather goods arranged in front of her. She looks into the camera and says naturally in the voice of Audio 1: "Eight years of handcraft, and what I still want to say most is — slow down, and you'll get there faster." Keep the warm, unhurried tone and pause rhythm of Audio 1, lip movements strictly in sync with the pronunciation, and her appearance consistent with the reference images.
Upload 1 portrait image + 1 scene image + 1 voice recording; "keep the tone and pause rhythm of Audio 1" is the key phrasing for high-fidelity voice reproduction.
04 30-Second Long Videos & Field Workflows
30 seconds is Wan 3.0's generational leap — one clip can now tell a complete story. This chapter gives you the segmented writing method for long videos, plus the veteran's cost workflow: draft at 480p, switch to a higher tier for the final cut.
4.1 The 30-Second Segmented Script
The way to write long videos is to slice the clip into time segments: each segment states its start-end seconds, picture content and camera work, and a global setup paragraph at the top locks the style and mood. Match the action load to the segment length — cramming 10 actions into 3 seconds just forces the model to fast-forward; if you want a cut, name the transition inside the segment.
[Global Setup] + [Start-End Seconds] + [Picture Content & Camera] + optional [Transition]
Time Slice
Explicit start-end seconds (e.g. 0-5s, 5-12s). Cutting 30 seconds into 3-5 segments is the most stable.
Picture & Camera
What happens in this slice and how the camera moves — action load matched to the slice length.
Global Setup
An opening paragraph that locks style, color tone, mood and hard no-gos for the whole clip.
Transition
How slices join: natural cut, dissolve, wipe — just name it.
A City Wakes Up in 30 Seconds
Prompt
[Global Setup] Documentary aerial style, a color gradient from cold pre-dawn blue to warm gold, true physical lighting, no subtitles, no score (pure ambient sound). 0-6s: The city skyline before dawn, lights scattered like stars, low-hanging clouds, the camera pushing in slowly from high altitude. 6-14s: The horizon pales into dawn; the camera descends through the cloud layer as the city's outline gradually sharpens, and an early light-rail train glides between the buildings, dragging a trail of light. 14-22s: The camera drops to street level and tracks along: a breakfast shop rolls up its shutter, steam rises from the bamboo steamers, a cyclist crosses the morning light — the city begins to wake. 22-30s: The sun leaps over the horizon; the camera pulls up fast back to high altitude, the whole city bathed in golden morning light, and the frame freezes on a magnificent wide shot.
The global setup sets the tone and four slices each own a segment — write the skeleton of a 30-second clip first, then fill in the flesh, and the pacing won't fall apart.
4.2 The 480p Draft Workflow
The veteran's workflow: draft at 480p during the creative phase — at roughly a quarter of the 1080p cost — and once the composition, motion and pacing check out, switch the same prompt to 1080p for the final cut. Ten revisions at draft stage don't hurt; the final passes in one go.
The Skater (Draft Version)
Prompt
At a skatepark at dusk, a teenager in an oversize hoodie drops in from the ramp, jumps, flips the board, lands steady, rides toward the camera and brakes hard in front of it, the kicked-up dust clearly visible against the backlight. The camera follows at a low angle, then orbits halfway and settles on his smiling face. Fast-moving camera with a sense of speed, street documentary style, warm golden tones.
Pre-filled with 480p — validate the action and camera work with this one first; once satisfied, switch the resolution to 1080p, stretch the duration, and run the same prompt for the final cut.