Guide

Wan 2.7 Guide

The complete writing playbook from core formula to first-last frame, mixed references and video editing — every chapter ships with ready-to-submit example prompts you can preload into the Molyin generator with one click. Wan 2.7 is Alibaba Tongyi's 27B MoE video model: cinematic 1080p, 2-15s continuous duration, native audio on every video, and first-last-frame interpolation — this guide walks through each weapon one by one.

Updated 2026-08-04

01 Core Formula

Wan 2.7 understands structured instructions: who, what they do, how the camera moves, where, and in what tone. This chapter breaks a great prompt into five reusable parts, then walks through two handy tools — five aspect ratios and the seed. Lock the skeleton first; every later chapter is just an extension of it.

1.1 The Basic Prompt Formula

Everything starts with one sentence: subject plus motion are the two required parts; camera, environment and aesthetics stack on demand. Shoot the clip in your head first, then write it down with the formula. Duration is a continuous 2-15s dial — plan exactly as much as the seconds can hold. One bonus: Wan 2.7 auto-expands short prompts with an LLM rewrite, but a well-structured original always beats its guesses.

[Subject] + [Motion] + optional [Camera] [Environment] [Aesthetics]

Required

Subject

The protagonist — a person, animal or product. The more specific, the better.

Required

Motion

What the subject does, with body parts and pace spelled out. Slow, coherent small movements are the most reliable.

Optional

Camera

Shot size and movement: close-up, push-in, pan. For multi-shot narratives, write the storyboard directly, e.g. "Shot 1 [0-3s] wide shot".

Optional

Environment

Scene and atmosphere: place, time, light, weather.

Optional

Aesthetics

Art direction and quality: color tone, texture, clarity.

Rider in the Neon Rain

Prompt

On a rainy Tokyo street at night, a female rider in a black leather jacket sits astride a heavy motorcycle, taking off her helmet and shaking out her rain-soaked hair, neon signs casting pink and blue reflections on the wet asphalt. The camera slowly pushes in from a low angle to a half-body close-up, drizzle clearly visible against the backlight. Cinematic quality, cyberpunk color grading, rich detail.

Two sentences for subject+motion, one each for camera and environment, style words to close — that's a solid Wan 2.7 baseline prompt. Suggested: 1080p, 16:9.

1.2 Choosing Among Five Ratios

Wan 2.7 offers five aspect ratios: 16:9, 9:16, 1:1, 4:3 and 3:4. There's only one rule — decide backwards from the publishing platform: 16:9 for landscape video, 9:16 for mobile feeds, 1:1 for square feeds, 4:3 or 3:4 for portraits and still life. Set the ratio before generating; cropping afterwards just throws away quality. Note: image-to-video has no ratio selector — the output follows your first frame.

Trench Coat Around the Corner

Prompt

On an old street bathed in afternoon sun, a model in a beige trench coat walks around the corner at an unhurried pace, the hem of her coat swaying with each step, and she pauses in front of a shop window with a sideways smile. The camera tracks at medium shot then slowly pushes in. Warm vintage film texture, rich street detail.

Pre-filled with 9:16 vertical — the standard canvas for mobile feeds, a full-body composition that fills the frame.

1.3 Reproduce and Iterate with Seed

The seed is Wan 2.7's randomness dial: fix it, and the same prompt regenerates highly similar results — change one word and see exactly what that word does; swap the seed, and the same prompt yields a fresh batch. Two habits worth building: note the seed when you get a result you love, and lock the seed while A/B testing descriptions. Generation is probabilistic — the same seed doesn't guarantee pixel-identical output, but composition and subject stability far exceed random draws.

Paper Crane on the Windowsill

Prompt

An origami paper crane stands on a sun-warmed windowsill, a breeze slipping through the window crack, its wings trembling gently before the wind lifts it off the sill, swirling across the room toward the bookshelf. The camera follows slowly, light and shadow dappling the wooden floor, fine dust floating in the air.

Pre-filled with seed 42 — change the words, keep the seed, and compare every fine-tuned iteration.

02 Image to Video

One first-frame image plus a motion description, and Wan 2.7 brings the picture to life — the prompt's center of gravity shifts from describing the scene to describing the motion, because the scene already exists. One hard rule first: the output ratio of image-to-video automatically follows your first frame — there's no ratio selector. Want a 9:16 result? Upload a 9:16 image.

2.1 Bringing Photos to Life

An image-to-video prompt covers only three things: what moves, how it moves, and whether the camera moves. Don't force things that aren't in the picture; do "wake up" things that are — fluttering fabric, flowing water, flickering lamplight are all high-hit-rate small motions. Add "keep the original texture" and the style won't drift.

Dusk at the Fishing Harbor

Prompt

Bring this photo to life: the setting sun dyes the fishing harbor golden-red, shimmering ripples slowly shift across the water, moored fishing boats rock gently on the small waves, flags on the masts flutter in the sea breeze, and a few seagulls skim past the distant breakwater. The camera stays fixed, keeping the warm dusk film texture of the original photo throughout.

Upload 1 first-frame image; write only the motion without restating the scene, and let "keep the original texture" hold the style. The output ratio follows the first frame.

2.2 First-Last Frame: Own the Start and the End

Wan 2.7 is one of four model families on Molyin with first-last-frame support (alongside the Seedance family, Kling 3.0 and Veo 3.1) — upload a first frame and a last frame, and the model generates everything in between. This changes how you prompt: skip the start and the end (they're in the images) and write only "how to get from A to B" — the pace of change and how the camera plays along. Keep both images at the same ratio, and the bigger the visual gap, the more process description you should provide.

Ink Drop Blooms into Mountains

Prompt

A drop of thick ink falls into clear water, slowly blooming and stretching, ink filaments curling and flowing through the water, gradually forming a layered ink-wash landscape of mountains with clouds drifting through the negative space. The camera slowly pushes into the heart of the ink mass, seamless throughout, with rice-paper texture and ink-wash quality from start to finish.

Turn on the "Add last frame" switch in image-to-video mode and upload both frames (ink drop → landscape); the prompt only describes the transformation, and 8 seconds gives the bloom room to breathe.

03 Reference to Video

Appearances and signature moves that words can't capture — characters, products, gestures — hand them to reference media. Wan 2.7's reference-to-video accepts a mix of images and videos, up to 5 items total. Reference them in your prompt as "Image 1", "Image 2", "Video 1" in upload order (images and videos are numbered independently), and the model locks appearance and motion from your media while the scene and story remain entirely yours.

3.1 The Character Reference Formula

A reference-to-video prompt is three beats: bring the media on stage with "Image N", describe what happens in the new scene, then close with "keep it consistent". Upload order is the numbering order — the first image is Image 1, the second video is Video 2. Multiple angles of the same character work best; put only one character per reference item, or the model can't tell who's who.

[References: Image 1…] + [Scene description] + keep [Subject] consistent

Required

References

Images or videos, up to 5 total, referenced as "Image N" and "Video N" in their own sequences.

Optional

Scene Description

What happens in the new scene: action, environment, camera — the frame is generated from scratch, all driven by your words.

Optional

Consistency

A closing "keep it consistent with the references" reinforces the appearance lock.

Chef Panda's Midnight Diner

Prompt

The giant panda from Image 1, wearing a dark blue apron, tosses and stirs food in a wok in a small kitchen with warm yellow lighting, steam rising, and it smugly tastes a spoonful of broth and nods with satisfaction. The camera circles the stove halfway and settles on a close-up of its face. Rich kitchen detail, a cozy, heartwarming everyday mood. The panda's appearance stays consistent with the reference image, 3D animation quality.

Upload 1+ reference images; "Image 1" refers to your uploaded character, and the closing "stay consistent" is the standard appearance-lock phrase.

3.2 Mixing Image and Video References

Wan 2.7's signature move: images lock appearance, videos lock motion, and both can share the stage — "the dancer from Video 1 walks onto the stage from Image 1", and the model takes what it needs from each. Three hard rules: images plus videos may not exceed 5 items total; with a reference video, duration is capped at 10 seconds (only pure-image references can use the full 15); and each reference item should contain a single character.

The Puppeteer Takes the Stage

Prompt

The puppeteer from Image 1 walks onto the vintage theater stage from Image 2 holding his strings, a spotlight catching him as he bows to the audience with the gesture from Video 1, then lifts his puppet and begins the show as the curtain slowly draws closed behind him. The camera slowly pushes in from a mid-shot in the audience, the theater mood solemn and mysterious. Appearances stay consistent with the reference media.

Upload 2 reference images + 1 reference video (3 items total, within the 5 cap); "Image 1", "Image 2" and "Video 1" each name their own role. With a reference video, duration tightens to 8s (10s cap).

04 Video Editing

Edit video with plain language — upload a 2-10s source video and write one plain-spoken instruction: change the style, the outfit, the prop, and everything else stays as it was. Output duration and ratio follow the source video — there are no duration or ratio selectors. Attach up to 4 reference images to show the model exactly "what to change into". The audio switch is your second pen: "keep original audio" decides the fate of the source track.

4.1 Style Transfer

Give the whole clip a new art skin: claymation, watercolor, pixel art, cyber neon — one sentence is enough. Name three traits of the target style — stroke, material, color tone — then add "keep the camera movement and pacing as-is", and the model changes the look without touching the story.

Turn Street Footage into Claymation

Prompt

Convert this video into claymation style: hand-sculpted rounded forms with visible fingerprint texture and clay material, soft miniature-set lighting, characters and buildings shaped like molded clay, camera movement and pacing kept exactly as in the original.

Upload a 2-10s source video; it routes to Wan 2.7 Video Edit automatically, with duration and ratio following the source. Naming the "form + material + lighting" trio is the most reliable style recipe.

4.2 Element Replacement

Swap outfits, props and products without touching the rest of the shot: state clearly "replace A with B", then add "keep everything else unchanged". Replacement edits are most precise with a reference image — "the leather jacket in Image 1" beats a hundred adjectives.

Swap in a Leather Jacket

Prompt

Replace the character's jacket in the video with the brown leather jacket from Image 1, the fabric creasing naturally with the movement, its texture and lighting direction matching the original scene, everything else unchanged.

Upload a source video plus up to 4 reference images; "replace A with B + keep everything else" is the most reliable replacement phrasing.

Wan 2.7 Parameter Cheat Sheet

Every tier the Molyin generator actually exposes for the Wan 2.7 family.

Wan 2.7Wan 2.7 Video Edit
ModesText to video / Image to video / Reference to videoVideo edit
Resolution720p / 1080p720p / 1080p
Duration2–15sFollows source video
Aspect ratio16:9 / 9:16 / 1:1 / 4:3 / 3:4Follows source video
AudioAudio always on, no controlFollows source video
First-last frameYesNo
Reference images1–5Up to 1
Reference videosUp to 3No
Reference audioNoNo
Return last frameNoNo

FAQ

The six most frequent questions, with answers we verified ourselves.