01 Core Formula
Wan 2.7 understands structured instructions: who, what they do, how the camera moves, where, and in what tone. This chapter breaks a great prompt into five reusable parts, then walks through two handy tools — five aspect ratios and the seed. Lock the skeleton first; every later chapter is just an extension of it.
1.1 The Basic Prompt Formula
Everything starts with one sentence: subject plus motion are the two required parts; camera, environment and aesthetics stack on demand. Shoot the clip in your head first, then write it down with the formula. Duration is a continuous 2-15s dial — plan exactly as much as the seconds can hold. One bonus: Wan 2.7 auto-expands short prompts with an intelligent rewrite, but a well-structured original always beats its guesses.
[Subject] + [Motion] + optional [Camera] [Environment] [Aesthetics]
Subject
The protagonist — a person, animal or product. The more specific, the better.
Motion
What the subject does, with body parts and pace spelled out. Slow, coherent small movements are the most reliable.
Camera
Shot size and movement: close-up, push-in, pan. For multi-shot narratives, write the storyboard directly, e.g. "Shot 1 [0-3s] wide shot".
Environment
Scene and atmosphere: place, time, light, weather.
Aesthetics
Art direction and quality: color tone, texture, clarity.
Rider in the Neon Rain
Prompt
On a rainy Tokyo street at night, a female rider in a black leather jacket sits astride a heavy motorcycle, taking off her helmet and shaking out her rain-soaked hair, neon signs casting pink and blue reflections on the wet asphalt. The camera slowly pushes in from a low angle to a half-body close-up, drizzle clearly visible against the backlight. Cinematic quality, cyberpunk color grading, rich detail.
Two sentences for subject+motion, one each for camera and environment, style words to close — that's a solid Wan 2.7 baseline prompt. Suggested: 1080p, 16:9.
1.2 Choosing Among Five Ratios
Wan 2.7 offers five aspect ratios: 16:9, 9:16, 1:1, 4:3 and 3:4. There's only one rule — decide backwards from the publishing platform: 16:9 for landscape video, 9:16 for mobile feeds, 1:1 for square feeds, 4:3 or 3:4 for portraits and still life. Set the ratio before generating; cropping afterwards just throws away quality. Note: image-to-video has no ratio selector — the output follows your first frame.
Trench Coat Around the Corner
Prompt
On an old street bathed in afternoon sun, a model in a beige trench coat walks around the corner at an unhurried pace, the hem of her coat swaying with each step, and she pauses in front of a shop window with a sideways smile. The camera tracks at medium shot then slowly pushes in. Warm vintage film texture, rich street detail.
Pre-filled with 9:16 vertical — the standard canvas for mobile feeds, a full-body composition that fills the frame.
1.3 Reproduce and Iterate with Seed
The seed is Wan 2.7's randomness dial: fix it, and the same prompt regenerates highly similar results — change one word and see exactly what that word does; swap the seed, and the same prompt yields a fresh batch. Two habits worth building: note the seed when you get a result you love, and lock the seed while A/B testing descriptions. Generation is probabilistic — the same seed doesn't guarantee pixel-identical output, but composition and subject stability far exceed random draws.
Paper Crane on the Windowsill
Prompt
An origami paper crane stands on a sun-warmed windowsill, a breeze slipping through the window crack, its wings trembling gently before the wind lifts it off the sill, swirling across the room toward the bookshelf. The camera follows slowly, light and shadow dappling the wooden floor, fine dust floating in the air.
Pre-filled with seed 42 — change the words, keep the seed, and compare every fine-tuned iteration.
1.4 Score Reference: Let the Picture Follow the Music
The audio input slot in text-to-video means "score" — upload a music track and the model treats it as background music, letting the pacing, motion beats and camera moves follow the music's emotion. In the prompt, simply name it: "use Audio 1 as the score", then add one line about how the picture should respond to the music (beat-matching, emotional build-up, mood fit). This is different from "every video has sound": the score slot is your deliberate musical direction — without it the model still scores the video on its own.
Dusk Drive on the Coastal Highway
Prompt
A vintage convertible cruises along a coastal highway at dusk, the girl in the passenger seat humming softly into the wind, her hair flying, palm trees rushing past by the roadside, the sea reflecting the sunset like scattered gold. The camera tracks alongside the car and slowly orbits toward the front, warm golden tones, film texture. Use Audio 1 as the score — the pacing and camera movement follow the music's emotion, and the orbit speeds up at the chorus.
Upload 1 music track to the audio slot (the t2v audio slot means score); pre-filled with 10 seconds — enough room for the music's emotional build and beat-matching.
02 Image to Video
One first-frame image plus a motion description, and Wan 2.7 brings the picture to life — the prompt's center of gravity shifts from describing the scene to describing the motion, because the scene already exists. One hard rule first: the output ratio of image-to-video automatically follows your first frame — there's no ratio selector. Want a 9:16 result? Upload a 9:16 image.
2.1 Bringing Photos to Life
An image-to-video prompt covers only three things: what moves, how it moves, and whether the camera moves. Don't force things that aren't in the picture; do "wake up" things that are — fluttering fabric, flowing water, flickering lamplight are all high-hit-rate small motions. Add "keep the original texture" and the style won't drift.
Dusk at the Fishing Harbor
Prompt
Bring this photo to life: the setting sun dyes the fishing harbor golden-red, shimmering ripples slowly shift across the water, moored fishing boats rock gently on the small waves, flags on the masts flutter in the sea breeze, and a few seagulls skim past the distant breakwater. The camera stays fixed, keeping the warm dusk film texture of the original photo throughout.
Upload 1 first-frame image; write only the motion without restating the scene, and let "keep the original texture" hold the style. The output ratio follows the first frame.
2.2 First-Last Frame: Own the Start and the End
Wan 2.7 is one of four model families on Molyin with first-last-frame support (alongside the Seedance family, Kling 3.0 and Veo 3.1) — upload a first frame and a last frame, and the model generates everything in between. This changes how you prompt: skip the start and the end (they're in the images) and write only "how to get from A to B" — the pace of change and how the camera plays along. Keep both images at the same ratio, and the bigger the visual gap, the more process description you should provide.
Ink Drop Blooms into Mountains
Prompt
A drop of thick ink falls into clear water, slowly blooming and stretching, ink filaments curling and flowing through the water, gradually forming a layered ink-wash landscape of mountains with clouds drifting through the negative space. The camera slowly pushes into the heart of the ink mass, seamless throughout, with rice-paper texture and ink-wash quality from start to finish.
Turn on the "Add last frame" switch in image-to-video mode and upload both frames (ink drop → landscape); the prompt only describes the transformation, and 8 seconds gives the bloom room to breathe.
2.3 Driving Audio: Make the Photo Talk
The audio input slot in image-to-video means "driving audio" — upload a voice recording and the person in the photo speaks along with it, lip movements, expressions and head motion driven by the speech rhythm. Talking-head clips, product pitches and character monologues all rely on it. In the prompt, one line suffices: "lip movements strictly in sync with the speech in Audio 1", then describe the camera and mood as usual.
Podcast Host on the Mic
Prompt
Bring this photo to life: the podcast host speaks naturally into the microphone with a lively expression, slight nods and hand gestures, lip movements strictly in sync with the speech in Audio 1, the emotional shifts in the voice carrying the facial expressions along. The warm desk lamp on the bookshelf behind glows quietly; the camera holds a fixed medium shot, keeping the documentary-photography texture of the original photo throughout.
Upload 1 portrait photo + 1 voice recording (the i2v audio slot means driving audio); clean, noise-free speech gives the highest lip-sync hit rate.
2.4 Video Continuation: Keep Shooting, No Reshoot
Image-to-video also accepts a 2-10s video as a "continuation base" (first_clip) — the model picks up from the end of that clip and keeps shooting, with motion, style and lighting naturally carried over. The prompt only describes the new segment; no need to recap what came before. Three boundaries to remember: the continuation base is mutually exclusive with a first-frame image (pick one); it can be paired with a last frame to lock the ending; but it cannot be combined with driving audio. The duration dial sets the total length after continuation, and billing follows the total.
Continuation: [new content description] + optional [continuity requirement]
New Content
Only describe what happens after the base clip ends: new action, new scene, new camera work.
Continuity
One line — "seamless action, style and lighting carried over from the original video" — prevents jumps at the seam.
The Skater's Second Jump
Prompt
Continue the uploaded video: after landing, the skater rolls forward smoothly, the camera keeps following at a low angle as he accelerates toward the stairs ahead, jumps, flips the board, lands steadily on the flat ground below, waves at the camera and skates out of frame. The action connects seamlessly, with style and lighting carried over from the original video.
Upload 1 clip of 2-10s as the continuation base (mutually exclusive with a first-frame image; cannot be paired with driving audio); pre-filled with 10s total — the dial sets the full length after continuation.
03 Reference to Video
Appearances and signature moves that words can't capture — characters, products, gestures — hand them to reference media. Wan 2.7's reference-to-video accepts a mix of images and videos: up to 5 images, up to 5 videos, and no more than 5 items in total. Name them in your prompt as "Image 1", "Image 2", "Video 1" in upload order (images and videos are numbered independently), and the model locks appearance and motion from your media while the scene and story remain entirely yours. Audio goes through a separate single-audio input slot (voice reference, see 3.3) and doesn't count against the reference quota. Note: the r2v duration cap is always 10 seconds — with or without reference videos.
3.1 The Character Reference Formula
A reference-to-video prompt is three beats: bring the media on stage with "Image N", describe what happens in the new scene, then close with "keep it consistent". Upload order is the numbering order — the first image is Image 1, the second video is Video 2. Multiple angles of the same character work best; put only one character per reference item, or the model can't tell who's who.
[References: Image 1…] + [Scene description] + keep [Subject] consistent
References
Images or videos — up to 5 images, up to 5 videos, 5 total — referenced as "Image N" and "Video N" in their own sequences.
Scene Description
What happens in the new scene: action, environment, camera — the frame is generated from scratch, all driven by your words.
Consistency
A closing "keep it consistent with the references" reinforces the appearance lock.
Chef Panda's Midnight Diner
Prompt
The giant panda from Image 1, wearing a dark blue apron, tosses and stirs food in a wok in a small kitchen with warm yellow lighting, steam rising, and it smugly tastes a spoonful of broth and nods with satisfaction. The camera circles the stove halfway and settles on a close-up of its face. Rich kitchen detail, a cozy, heartwarming everyday mood. The panda's appearance stays consistent with the reference image, 3D animation quality.
Upload 1+ reference images; "Image 1" refers to your uploaded character, and the closing "stay consistent" is the standard appearance-lock phrase.
3.2 Mixing Image and Video References
Wan 2.7's signature move: images lock appearance, videos lock motion, and both can share the stage — "the dancer from Video 1 walks onto the stage from Image 1", and the model takes what it needs from each. Three hard rules: up to 5 images, up to 5 videos, and no more than 5 items in total; the duration cap is always 10 seconds (with or without reference videos); and each reference item should contain a single character.
The Puppeteer Takes the Stage
Prompt
The puppeteer from Image 1 walks onto the vintage theater stage from Image 2 holding his strings, a spotlight catching him as he bows to the audience with the gesture from Video 1, then lifts his puppet and begins the show as the curtain slowly draws closed behind him. The camera slowly pushes in from a mid-shot in the audience, the theater mood solemn and mysterious. Appearances stay consistent with the reference media.
Upload 2 reference images + 1 reference video (3 items total, within the 5 cap); "Image 1", "Image 2" and "Video 1" each name their own role; the r2v duration cap is always 10s — pre-filled with 8.
3.3 Voice Reference: Let the Character Speak in Your Voice
The audio input slot in reference-to-video means "voice reference" — upload a recording of the target voice and the generated character speaks in that voice, faithfully reproducing accent, tone and emotional tension. Branded voiceovers, virtual-presenter clips and multi-character dialogue all rely on it. In the prompt, name it: "say the line in the voice of Audio 1" — write the line itself directly into the prompt, and lip movements sync with the pronunciation automatically.
The Brand Ambassador's New Tagline
Prompt
The brand ambassador from Image 1 stands in the minimalist studio from Image 2, facing the camera, and says the brand tagline in the voice of Audio 1: "Make every departure a view." The tone is warm and resolute, lip movements strictly in sync with the pronunciation, with a slight smile on the word "view". The camera slowly pushes in to a half-body close-up. Appearance stays consistent with the reference images, soft studio lighting, a clean premium look.
Upload a portrait image + a scene image + 1 voice-reference track (the r2v audio slot means voice reference); write the line into the prompt and leave the voice reproduction to the audio.
3.4 First-Frame Anchor: Own the Opening Composition
Reference-to-video supports an extra "first frame" upload to anchor the opening shot — the composition, camera position and lighting of frame one are yours to control precisely, and the model unfolds according to the prompt from frame two. One side effect to know: once a first frame is provided, the output ratio follows it and the ratio selector stops applying (the ratio parameter is ignored in this case).
The Barista's New-Product Opening
Prompt
Use the uploaded first frame as the opening shot: the barista from Image 1 turns from behind the counter to face the camera, raising the new signature drink from Image 2, steam rising from the cup, and smiles as he presents it toward the lens. The opening composition is strictly anchored to the first frame, then the camera slowly pushes in to a close-up of the drink. The person's and product's appearances stay consistent with the reference images, in a warm wood-toned café atmosphere.
In r2v mode upload a first-frame image + 1-2 reference images; once a first frame is provided the ratio follows it — the ratio selector becoming inert is expected behavior.
04 Video Editing
Edit video with plain language — upload a 2-10s source video and write one plain-spoken instruction: change the style, the outfit, the prop, and everything else stays as it was. Output duration automatically follows the source video (no duration selector); the ratio offers six options — auto follows the source by default, or force an explicit ratio like 16:9 or 9:16. You can also attach 1 reference image to show the model "what to change into". Audio control is your second pen: the "keep original audio" switch decides the fate of the source track, and the settings dialog has a "smart rewrite" switch (prompt_extend, on by default) that auto-expands short instructions with detail.
4.1 Style Transfer
Give the whole clip a new art skin: claymation, watercolor, pixel art, cyber neon — one sentence is enough. Name three traits of the target style — stroke, material, color tone — then add "keep the camera movement and pacing as-is", and the model changes the look without touching the story.
Turn Street Footage into Claymation
Prompt
Convert this video into claymation style: hand-sculpted rounded forms with visible fingerprint texture and clay material, soft miniature-set lighting, characters and buildings shaped like molded clay, camera movement and pacing kept exactly as in the original.
Upload a 2-10s source video; it routes to Wan 2.7 Video Edit automatically, duration follows the source and the ratio defaults to auto (follow source). Naming the "form + material + lighting" trio is the most reliable style recipe.
4.2 Element Replacement
Swap outfits, props and products without touching the rest of the shot: state clearly "replace A with B", then add "keep everything else unchanged". Replacement edits are most precise with a reference image — "the leather jacket in Image 1" beats a hundred adjectives.
Swap in a Leather Jacket
Prompt
Replace the character's jacket in the video with the brown leather jacket from Image 1, the fabric creasing naturally with the movement, its texture and lighting direction matching the original scene, everything else unchanged.
Upload a source video, optionally with 1 reference image; "replace A with B + keep everything else" is the most reliable replacement phrasing.