Guide

Seedance 2.5 Guide

Seedance 2.5 stretches a single generation to 30 seconds and natively supports second-level timestamped scripts — the era of point-and-shoot precision is here. This guide reorganizes the official usage manual around real creative tasks: from the core formula and real-person writing to timestamped long takes, multimodal references, and local editing — every example pre-fills the Molyin generator with one click.

Updated 2026-08-06

01 The Core Formula

Seedance 2.5's grasp of structured direction climbs another level: a complete prompt is built from four layers — asset descriptions, a one-line summary, the detailed plot, and a global wrap-up. This chapter sets up the skeleton first, then hands you the officially validated real-person formula — purpose-built to cure the "AI face" and plastic skin.

1.1 The Complete Prompt Formula

A complete Seedance 2.5 prompt = asset descriptions + one-line summary + detailed plot + global wrap-up. Skip the first layer if you uploaded no assets; break the plot into segments along a timeline or story arc; and put the global wrap-up at the end, reinforcing what runs through the whole film and what must not appear (subtitles, BGM).

[Asset Descriptions] + [One-Line Summary] + [Detailed Plot] + [Global Wrap-Up]

Required

One-Line Summary

Subject + place + event + genre/style + any special camera move — one sentence that sets the film's tone.

Required

Detailed Plot

Write along a timeline or story arc: each shot carries its picture content, camera move, actions, dialogue, and sound effects — plus what you don't want (negative descriptions).

Optional

Asset Descriptions

Number your assets in upload order and state what each one is for: a character's appearance, a voice, a motion, or a scene. With no reference assets, this whole layer can be omitted.

Optional

Global Wrap-Up

At the end, reinforce the details that run through the whole film: camera position, environmental traits, light and atmosphere — or global bans (subtitles, BGM).

A Wetland Red-Crowned Crane Documentary

Prompt

Photorealistic nature documentary style with cinematic, true-to-life lighting. In the early-morning mist, across a vast wetland of reeds (scene references Image 1), an elegant red-crowned crane (appearance references Image 2) spreads its wings in a dance — feathers white as snow, a vivid red crown, a tall and slender frame. The setting is a shallow autumn wetland, the water glittering, surrounded by golden reeds, duckweed, and aquatic plants, with distant morning fog and the rising sun soft-blurred in the background. Low-angle medium shot with a subtle handheld feel, largely locked off, keeping the crane centered at all times. 0s-3s: The crane stands still in the shallow water and slowly spreads its broad black-and-white wings, the motion gentle, the airflow from its wingbeats rippling the surface. A morning breeze blows; sunlight slants in from the horizon, piercing the mist in a Tyndall effect. 3s-8s: The crane hops lightly across the water, its feet skimming the surface in turn and tossing up crystalline spray. Its long neck arches back, beak pointing at the sky; then it lands softly, folds its wings, sweeps its gaze gracefully around, and lets out a clear, ringing call. The low camera gently follows the crane's hops up and down, with natural depth of field and the foreground reeds softly blurred. Ambient nature sound only: flowing water, powerful wingbeats, and the crane's cry — serene, airy, and beautiful.

Requires 2 reference images (scene + subject); open with one sentence that sets the tone, break the plot down along a timeline in the middle, and close with the global mood — the four-layer structure reads at a glance.

1.2 The Real-Person Formula (Curing the AI Face)

2.5 has greatly dialed down the "AI look," but how you write is still the deciding factor for photorealistic people. The officially validated seven-dimension formula breaks a person into age/ethnicity, skin texture, facial details, eyes, hair, wardrobe, and build/presence — the eyes are the soul, and the fidelity suffix "preserve the real fine pores and skin texture" is nearly mandatory.

[Age/Ethnicity] + [Skin Tone/Texture] + [Facial Details] + [Eyes/Soul] + optional [Hair] [Wardrobe] [Build/Presence]

Required

Age/Ethnicity

Exact age + nationality/ethnicity + a style adjective + a face-shape noun, e.g. "a 22-year-old East Asian woman with a cool, classically sculpted face."

Required

Skin Tone/Texture

Warm/cool tone + skin-tone noun + texture adjective + the fidelity suffix "preserve the real fine pores and skin texture" (adding freckles or blemishes reads even more real).

Required

Facial Details

Eye shape + brow bone + nose bridge + lip shape + jawline — write at least 3-4 of these together; they decide how recognizable the person is.

Required

Eyes/Soul

An adjective for the gaze + the message or metaphor the eyes convey + the underlying emotion — this decides the character's emotional range.

Optional

Hair Style/Color

Exact hair color + condition/texture + a hairstyle noun + environmental interaction (lifted by the breeze, etc.).

Optional

Wardrobe/Texture

Cut and tailoring + color + garment noun + fabric material and wear state + how it's worn — wardrobe builds character from the side.

Optional

Build/Presence

Frame and shoulder traits + framing requirement + posture and eyeline + an overall mood word.

Close-Up of a Classical Beauty

Prompt

A 22-year-old East Asian woman with a gentle, classical face made for cinema. Cool-toned fair skin with a fine, supple texture, preserving the real fine pores and skin texture. Slender, peach-blossom eyes (the rims slightly moist), a softly relaxed brow, a small straight nose, full lips with a faint, barely-there tender smile, and a softly drawn jawline. Her gaze is deeply affectionate and full of feeling, a pool of spring water behind her eyes — tender, focused, carrying deep attachment and a trace of reluctance. Jet-black hair is coiled into an effortless, elegant low classical bun, held by a single plain jade hairpin, a few fine strands drifting loose along her cheeks and stirring in the breeze. She wears a minimalist, plain-white cross-collared robe of softly lustrous, understated silk, the collar slightly parted. Her frame is slender, her shoulders narrow, radiating a gentle, wistful, classically romantic air. Close-up framing, eyes looking straight into the lens, soft natural light, shallow depth of field.

Works as pure text-to-video; the same writing applies to animated characters too — just swap the "photoreal skin" language for the target art style.

02 Timestamps and the 30-Second Long Take

Seedance 2.5's headline upgrade: a single generation doubles from 15 to 30 seconds, and second-level timestamps are natively supported — you can specify exactly what happens at which second and where each cut lands, for script-level precision. This chapter teaches the three-layer structure for long videos and the transition formula.

2.1 Timestamped Shot Scripts

The core of the long-video formula is slicing the film into time segments: each slice spells out "start–end seconds + stage theme + physical directions for action/expression + emotional subtext." Subtext explains to the AI why you designed it this way — write it and the acting nuance is night and day. Before the slices, don't forget the "[Global Setup]" block: environment and texture, visual style, camera language, character design, performance core, and prohibitions.

[Start Second–End Second] + [Stage Theme] + [Action/Expression Directions] + optional [Emotional Subtext]

Required

Time Slice

Explicit start and end seconds (e.g. 0-3s, 00:08-00:15) — 2.5 executes down to the second.

Required

Action/Physical Directions

Shot size + composition + the character's detailed movements — everything the frame must strictly follow in this slice.

Optional

Stage Theme

Name the slice (e.g. "Pressing," "Resignation") to help the model lock the paragraph's emotional key.

Optional

Emotional Subtext

Explain to the AI the intent behind the physical directions, e.g. "she isn't venting — she's waiting for an answer."

29-Second One-Take: A Riverside Farewell

Prompt

29-second one-take (a woman in period dress says goodbye by the river) [Global Scene Setup] Base environment: a riverside at dawn, a blurred small boat in the background — a lonely farewell scene. The mood is quiet and restrained. Visual style: cinematic, shallow depth of field (clean, blurred background), soft natural light. Camera language: hold a tight close-up on the woman throughout, as if from the subjective viewpoint of the man standing across from her. A subtle handheld breathing feel, no fast cuts. Character design: a 22-year-old East Asian woman with a gentle, classical face made for cinema. Cool-toned fair skin with a fine, supple texture, preserving the real fine pores and skin texture. Slender, peach-blossom eyes (the rims slightly moist), a small straight nose, full lips with a faint, barely-there tender smile. Her gaze is deeply affectionate and full of feeling, a pool of spring water behind her eyes. Jet-black hair coiled into a low classical bun, held by a single plain jade hairpin. She wears a minimalist, plain-white cross-collared robe of soft, lustrous silk. A slender frame radiating a gentle, classically romantic air. Performance core: restrained and fine-grained. Focus on the shifting gaze, the rise and fall of her breathing, the slight tremble of her lips, the tightening and release of her brow, the swallow in her throat, and the natural slide of tears. Prohibitions: no exaggerated sobbing, no fast cuts, no large body movements, no extra dialogue or BGM, and the tears must not fall early. [Emotion and Action Storyboard (building over 0-29 seconds)] Stage 1, 0-3s [Pressing]: she looks straight into the lens, no tears in her eyes yet, brow faintly knitted, lips parting as she murmurs softly: "Do you really have to go?" Subtext: she isn't venting — she's waiting for him to say out loud an answer she guessed long ago. Stage 2, 3-10s [Resignation]: her gaze slowly drifts away to the empty space beside her, eyelids lowering, the corners of her mouth pulling into a brief bitter smile that falls at once; her nostrils tighten slightly, she draws the smallest restrained breath, her chest rising once. Subtext: with this motion she is forcing the hurt and the bitterness back down. Stage 3, 11-17s [Remembering]: the camera pushes in slightly. She turns her eyes back and slowly, carefully scans his face. Her eye rims are red, but the tears are held back hard — not a drop falls. In the middle, a 0.5-second dead-silent pause: her lips stir once and press shut, her chin tightens, her throat rolls faintly. Subtext: this look is not goodbye — she is carving his face into memory one last time, for life. Stage 4, 18-23s [Lamenting]: she lowers her eyes, and without warning a single tear drops straight onto her collar. She doesn't wipe it; when she looks up again, the pleading in her gaze has turned into profound sorrow. The knitted brow begins to loosen, and she shakes her head — barely — once. Subtext: the relaxing brow signals inner release; the tiny head-shake is one last, helpless refusal of fate. Stage 5, 24-29s [Letting go]: the camera pushes to an extreme close-up. She works a soft smile onto her face (no self-mockery in it), and just as it reaches the corners of her mouth, a second tear slides from the corner of her eye past her nose and stops at her lip. In a near-inaudible voice she struggles to keep steady, she says: "Go." The smile stays frozen on her face as the tears keep sliding down, her gaze never leaving his. The camera finally rests on her face — smiling through tears, her eyes full of nothing but sorrow.

Global setup plus timestamped storyboard is the standard skeleton of a 30-second long take; timing prohibitions like "the tears must not fall early" are exactly the precision only 2.5 can hold.

2.2 Cuts and Transition Formula

For multi-shot narrative, write the transition directions right into the timestamp slices. A cut specification = transition type + base constraints + cut logic. Common transitions: natural cut, fade in/out, dissolve, white flash/black flash, occlusion wipe (an object covers the lens, the frame goes black, then pulls away), match-cut morph (a full moon becomes a wine bowl), action whip (whip pan), dynamic handoff (an outfit change), dolly in/out (pupil-to-scene), ink-wash transition. You can also offer candidates and let the model choose: "from natural cut / occlusion wipe / match-cut morph, pick whichever best fits this film's style."

[Transition Type] + optional [Base Constraints] + [Cut Logic]

Required

Transition Type

Name the transition: "use an occlusion wipe at the cut (no hard cuts, no objects appearing out of thin air)."

Optional

Base Constraints

Keep the footage's breathing feel and durations in balance; state what shot A and shot B each contain.

Optional

Cut Logic

Reinforce that "scenes must switch naturally"; when you want richer camera language, add a shot-size change (e.g. close-up to medium shot).

30-Second Wuxia Anime: Five Transitions in a Row

Prompt

[Global Setup] Base environment and texture: an open-air roadside inn deep in a quiet bamboo forest at night. Push for top-tier 2D animation craft — the drift of falling bamboo leaves and the ripples in the wine must obey real physics, with the floating grace unique to Chinese-style wuxia. Visual style: premium donghua period look (cel-shaded 3D with realistic rendering), ink-wash edges. High-contrast interweaving of cold and warm light (desolate blue-white moonlight against the warm yellow candlefire on the inn's table). Camera language: razor-sharp montage cutting and multiple types of physically seamless transitions, with extreme contrast between stillness and motion. Character design: a 22-year-old swordsman with the handsome, sculpted face of donghua heroes. Sword-like brows and starry eyes, a high ponytail, loose strands over his forehead. He wears a white cross-collared combat outfit trimmed in black; a black-sheathed longsword rests on the table. Performance core: the killing speed of a top-tier wuxia "one-strike finish," the swordsman's unhurried composure, and the poetic negative space of Eastern aesthetics. Prohibitions: no greasy 3D look (strictly keep the pure cel-shaded donghua style). No sloppy fight choreography, no clipping or warped limbs. No subtitles and no extra dialogue. The ink-wash effect may only appear after the 25-second mark, with the "click" of the sheath — never earlier. [Timestamped Storyboard (with explicit cuts and transitions)] 00:00-00:04 Shot 1 [A Quiet Night's Drink] (match-cut morph): the camera holds a static low angle on the night sky, a huge, bright full moon dead center. The moon gradually dims and fades as a round wine bowl — identical in shape and size (top-down view of the table) — slowly materializes where the moon was and fully replaces it, moonlight dotted across the wine's surface with fine ripples. The transition must be seamless. The camera then slowly pulls back and down from the top view to eye level: the swordsman sits at the inn's table and slowly lifts the bowl. 00:05-00:10 Shot 2 [Hidden Killing Intent] (dolly-in close-up transition): the swordsman's motion hitches; his gaze turns suddenly razor-sharp. The camera drives forward at extreme speed, magnifying his eye until the pupil fills the screen — and the frame unfolds seamlessly inside the pupil: high above the bamboo forest, a black-clad assassin gripping a short blade reverse-handed dives out of the sky like a hawk, straight at the lens. 00:11-00:18 Shot 3 [Blades Cross] (action whip-pan transition): the camera yanks back to the swordsman's full-body framing as he leaps from his seat, drawing his sword. The camera follows his sword arm in a violent swing toward frame right (whip pan), smearing the whole screen into horizontal motion blur. Inside the blur the scene jumps to mid-air and freezes: the swordsman's longsword and the assassin's short blade collide hard, bursting into huge orange sparks and a ring-shaped shockwave where the weapons meet. 00:19-00:24 Shot 4 [One Strike Lands] (occlusion wipe): the two pass each other in mid-air. The shockwave of the exchange whips up a storm of bamboo leaves; one huge emerald leaf sweeps sideways across the lens, fully covering it and blacking out the frame. As the leaf drifts off, the picture has already cut to a ground-level eye-height shot: the assassin falls limply to the ground in the background while the swordsman, back to camera in the foreground, slowly sheathes his longsword. 00:25-00:30 Shot 5 [The Lingering Note] (ink-wash transition): with the "click" of the sword seating in its sheath, the entire bamboo forest and the fallen assassin suddenly dissolve like a drop of thick ink hitting water, spreading outward as black-and-white ink wash — and the frame freezes for good inside the splashed-ink landscape.

Each slice names one transition and spells out how it executes; negative timing constraints like "never earlier" are what keep the effect from jumping the gun.

03 Multimodal References

2.5 opens the reference limits wide: up to 30 reference images, and reference video and audio clips up to 30 seconds each (check the parameter cheat sheet for the tiers actually offered on Molyin) — plus pure-audio driving for the first time. This chapter covers the four most useful reference scenarios: multi-person frames, voice cloning, creative transfer, and multi-grid storyboards.

3.1 Multi-Person References (Fixing the Twins Problem)

2.5 focused on fixing "twins" and "face-swap drift" in multi-person scenes, but writing still matters: bind each character to a reference image one by one, refer to them by the same name throughout, and add a global lock at the end. Use solo photos for character references — never multi-view collages.

[Character Binding] + optional [Relationships and Actions] + [Global Lock]

Required

Character Binding

"Character A's appearance references Image 1, Character B's references Image 2" — one solo reference image per character, each named in turn.

Optional

Relationships and Actions

Spell out the spatial relationships and interactions: "A hands B a glass of water," "four people dance in sync at a fixed spacing."

Optional

Global Lock

Close with "each character's facial features and wardrobe stay consistent with the reference images throughout — no face-swapping, no switching, no distortion."

Four-Person Retro Street Dance

Prompt

A joyful short set on a retro American cartoon-animation city street (scene references Image 5). The four young dancers' appearances, outfits, and facial features reference Image 1, Image 2, Image 3, and Image 4 respectively; the four dance in sync on the retro pedestrian street, their body movements tight and continuous, the spacing between them fixed. The cool-warm clash of the street neon stays unified throughout; locked-off frontal camera with a slow lateral tracking move. Every costume detail, facial feature, and body shape stays highly consistent with the reference images throughout — no face-swapping, no trading moves, each identity carries through, motion fluid with no clipping or distortion, in a retro cartoon-animation look.

Requires 5 reference images (4 characters + 1 scene); the official guidance is 1-8 subject images for best results — beyond that, stability drops and you'll be rerolling.

3.2 Voice References

Upload a clean, noise-free voice clip as the reference audio and the model clones its vocal character with high fidelity — even the emotional tension and speech rhythm carry over: a fiery speech, an icy newsreader tone, a sleepy whisper all work. In the prompt, name it directly — "voice references Audio N" — and add one emotional direction for fine control.

A Detective's Whisper on a Rainy Night

Prompt

On a rain-slick rooftop flashing with neon at night, the private detective from Image 1 leans against the railing and slowly raises his head toward the camera, rain dripping off his trench coat. In the voice of Audio 1 he says, low: "I found you." Keep Audio 1's low, suspenseful intonation, breathing rhythm, and gravelly texture, with lip movements strictly synced to the pronunciation. The only ambient sound is rain and distant traffic — no background music, no subtitles.

Requires 1 character reference image + 1 reference audio clip; "keep Audio 1's intonation and breathing rhythm" is the key phrasing for high-fidelity cloning.

3.3 Creative Transfer (Not Just Motion Transfer)

2.5 upgrades reference-video understanding from "motion imitation" to "creative transfer": beyond body dynamics it can grab the source's camera language, color logic, emotional tone, even its viral internet feel. The key is describing the core you want transferred — "the cute vibe," "the comedic timing," "the edit rhythm" — not just the surface motion.

A Black Shiba's Playful Day

Prompt

Reference the Shiba Inu's cute animated expressions, demeanor, movements, and creative pattern in Video 1 to make a short video of the black Shiba Inu from Image 1, precisely capturing the playful, lively state of the dog in the reference video. The setting is a bright living room with soft natural light, the camera gently following the dog's movements. Keep the black Shiba's coat color and physical features from Image 1 unchanged; the mood is warm and healing.

Requires 1 reference video + 1 subject image; "precisely capture the ... state/demeanor" triggers creative-level transfer far better than "do the same moves."

3.4 Multi-Grid Storyboards

Upload a multi-grid storyboard image (a simple line drawing or stick figures are fine) as the plot skeleton, and the model generates a coherent video that strictly follows the shot plan. Four steps: first introduce what each panel represents, then bind the subject designs to their reference images, then fill in each shot's composition, shot size, camera move, and action, and finally the global style and prohibitions.

A Stray Cat Finds a Home (9-Grid Storyboard)

Prompt

Image 1 is a complete film storyboard containing 9 consecutive shots; each panel represents one full shot, together forming a complete, warm, healing cinematic narrative. The middle-aged woman references Image 2: a woman in her 40s-50s, 3D animated-film CG style, fair skin, short dark-brown curly hair, large eyes, a gentle, kind face, always with a natural smile. She wears a pale-yellow knit cardigan over a white inner layer and a blue ankle-length skirt — a homemaker-mother look; her appearance, wardrobe, and proportions stay consistent throughout. The kitten references Image 3: a round, chibi orange-and-white tabby kitten, 3D animated-film CG style, big eyes, a pink nose, fluffy soft fur. It wears a yellow knit hat and carries a green drawstring cloth bag — sweet and endearing; its coat, outfit, and proportions stay consistent throughout. Shot 1 (0-3s): evening in a city residential block; a medium shot at low angle slowly pushes in as the stray kitten hops onto a curbside trash can, flips the lid open with its front paws, and pokes its head in looking for food. Warm sunset light on the street; the mood is a little lonely but real. Shot 2 (3-5s): a close shot pushing slightly into the trash can; the kitten grips the rim with both paws and rummages inside, ears twitching, eyes focused — hunger and helplessness front and center. Shot 3 (5-9s): a locked-off medium shot at the apartment building entrance; the woman stands at the door holding food and spots the kitten. The kitten looks back at her, wary but not running — their first eye contact. Shot 4 (9-12s): a full-body medium shot; the kitten stands in the open yard and slowly approaches the doorway, hope in its eyes. Shot 5 (12-16s): the camera drops to the cat's eye level; the woman crouches and gently sets the food bowl in front of the kitten, which immediately lowers its head to eat, tail swaying lightly, while the woman watches with a smile — the frame full of warmth. Shot 6 (16-20s): a tracking medium shot; the woman walks slowly out of the yard carrying a basket of vegetables, the kitten quietly following beside her, always half a step behind. Shot 7 (20-23s): at the interior doorway, the woman slowly opens the door; warm yellow light spills out, forming a warm-cool contrast with the evening outside. Shot 8 (23-26s): a medium shot from behind and to the side; the woman turns back, smiles, and waves softly, inviting the kitten home. The kitten stands at the threshold looking up at her, ears perked, tail swaying, hesitating before gradually accepting the invitation. Shot 9 (26-30s): the camera slowly pulls back; the woman waits patiently at the door as the kitten finally steps forward and follows her inside, the warm yellow light stretching both their silhouettes long as the door slowly closes — the frame resting on the warm, quiet courtyard. Use a Pixar animated-film CG style, PBR physical materials, high-quality global illumination, soft volumetric light, cinematic depth of field. Camera movement natural and steady, the animals' behavior true to life, and continuity of time, space, character positions, and lighting maintained between shots. No face-swapping, distortion, or proportion errors on the characters; no extra unrelated people, animals, or props.

Requires 3 reference images (storyboard + woman + kitten); per-shot timestamps map one-to-one onto the storyboard grids for maximum fidelity.

04 Voice and Language

2.5 breaks the language barrier: focused optimization for Chinese, English, Spanish, Indonesian, and Malay, plus full coverage of Thai, Arabic, Portuguese, Vietnamese, Japanese, and Korean — prompting in your native language and multi-language dialogue pronunciation are both far more accurate. At the same time, negative instructions like "no subtitles / no BGM" are finally truly obeyed, and a reference video's BGM and vocals can be losslessly separated.

4.1 Multilingual Prompts and Dialogue

Write the prompt in whichever language you know best; when a character needs to speak, write their lines in the target language and the lip-sync will match that language's pronunciation. For ads aimed at a specific market, remember to add a language constraint: all signage, captions, and logos unified in the target language, no mixed-in foreign text.

A Japanese Ad for a Tokyo Ramen Shop

Prompt

The overall style is a real commercial ad: cinematic photography, warm yellow lighting, a Tokyo-at-night atmosphere, natural and fluid visual language. Shot 1 (0-3s): the camera slowly pushes from a Tokyo street at night (streetscape references Image 1) toward the entrance of a ramen shop (storefront references Image 2). The shop sign, lanterns, business-hours board, and menu are all in natural, correct Japanese — no Chinese, English, or garbled text anywhere. Shot 2 (3-7s): the camera moves inside, where the clerk (appearance references Image 3) stands in the open kitchen, smiles at the camera, and says in natural, fluent Japanese: "Hello, would you like a bowl of our homemade tonkotsu ramen?" His mouth shapes stay strictly synced to the Japanese pronunciation, the voice natural, with the characteristics of a real Japanese male speaker. The background keeps the real ambient sounds of boiling noodles, steam, and clinking tableware. Shot 3 (7-11s): cut to a close-up of the ramen (references Image 4), steam rising continuously, the chashu swaying slightly, the oil on the broth flowing naturally. Brand caption text appears on screen, all in natural, correct Japanese: "Rich flavor, a bowl of happiness," in a typeface true to Japanese commercial-ad design. Shot 4 (11-15s): the customer (appearance references Image 5) sits at the table, tastes the ramen, breaks into a satisfied smile, and says naturally to the camera: "It's really delicious — I'll definitely be back!" The voice keeps natural Japanese pronunciation with accurate lip-sync. The brand logo (references Image 6) and a Japanese slogan appear: "A bowl I crave every day." All signage, menus, posters, and logos stay in unified Japanese — no other language mixed in; every line of dialogue in the video is in Japanese, and no Chinese subtitles or dubbing may be generated. Character identities stay consistent — the clerk and the customer never swap faces or outfits.

Requires 6 reference images (street / storefront / clerk / ramen / customer / logo); the constraint "everything in the target language, no mixed-in text" must be spelled out hard.

4.2 Clean Plates and BGM Separation

2.5 greatly improves generation stability, and the negative instruction "no subtitles, no BGM" finally lands with high probability — the clean plate commercial post-production wants is here. In the other direction, for a reference video with sound, one line — "remove the background music, keep the vocals" — losslessly separates the audio tracks while perfectly preserving the original picture and subtitles.

Silent E-Commerce Outfit Showcase

Prompt

A vertical 9:16 e-commerce outfit showcase video against a pure-white, minimalist seamless studio backdrop (layout references Image 1), the fixed split layout holding throughout: on the left side of the frame, three rectangular modules with thin black rounded borders stacked vertically, each topped by a black serif numeral (1, 2, 3), each module containing a flat-lay of the corresponding outfit piece with a small "1x" label beneath it; on the right side, a young female model stands full-body displaying the outfit for real — full-body eye-level locked-off camera, soft even studio lighting, no hard shadows, fabric textures crisp and fine. The model holds a natural, elegant stance with a few slow motions — a hand rising to adjust the clothing, a shift of posture — the overall style clean and premium. Absolutely no text captions, slogans, watermarks, logos, or extra annotations at any point; no background music, no voiceover, no ambient noise — a completely silent, purely visual picture; the camera stays locked with no push, pull, pan, or tilt, and the layout structure stays unchanged throughout.

Requires 1 layout reference image, and the audio toggle is already switched off for you at generation time; the "no subtitles + no BGM + silent" trio finally executes reliably on 2.5.

Strip the BGM, Keep Only the Voice

Prompt

Remove all the background music from Video 1, keeping only the voices of the people speaking and the ambient noise. The vocals must keep their original timbre, tone, breathing, and detail; the picture, people, subtitles, image quality, color grade, and composition all stay completely unchanged — do not alter any visual content.

Requires 1 reference video with vocals; this is the go-to move for fan edits, interview cleanup, and swapping in multi-language dubs.

05 Editing, White Models, and Green Screen

Upload a source video as a reference video and Seedance 2.5 performs precision surgery without touching the rest of the footage: local removal and replacement, white-model rendering into a finished film, green-screen background swaps, and seamless transitions between two videos. All four scenarios in this chapter run in reference-to-video mode.

5.1 Local Removal and Replacement

The core of the editing formula is spelling out "from A to B": the exact edit target + the action and change + when it takes effect. Workhorse verbs: add, remove, modify/replace/change to. Pair with timestamps to lock the time range of the edit — leave it out and the edit applies to the whole clip.

[Exact Edit Target] + [Action and Change] + optional [Effective Timing]

Required

Edit Target

Lock the target: "the woman on the left in the video," "the glass on the table," "the green English title text" — the more specific, the less collateral damage.

Required

Action and Change

Issue the command: from A to B — "remove it and naturally fill in the background," "replace with black dress trousers."

Optional

Effective Timing

"Apply the change consistently across the whole clip" or "only during seconds 5-8" — locking the timeline prevents continuity breaks.

Remove Extra People

Prompt

Remove the woman on the left and the man on the right from Video 1, keeping only the woman in the center, her movements unchanged. Fill the removed areas naturally to match the original footage's lighting and perspective, leaving no ghosting or blurry patches; every other part of the picture, the composition, and the camera movement stay exactly the same.

Requires 1 source video; name who stays and who goes by position, add one line about clean fill-in — that's the highest-success recipe for invisible removal.

Recolor the Outfit + Remove Text

Prompt

Completely remove the green English title text that appears in Video 1, and edit the blue sportswear worn by the person doing back stretches into red sportswear — change only the clothing color. Also fully remove the extra fitness props in the background: the blue kettlebell on the right and the blue balance pad on the left. The person themself, their skin, head, hands and feet, body posture, exercise movements, the yoga mat, the carpet, the wall, and the camera frame all stay unchanged. After the text and props are removed, the areas they covered must be filled in naturally — no text ghosting, no blurry patches, no prop remnants. The red sportswear must stay stable as the person rolls over, lifts a leg, and stretches, keeping the original folds and fabric texture of the clothing.

Requires 1 source video; multiple edit targets can go in one prompt, but each target needs its own clear "change into what."

5.2 White-Model Rendering

Upload a white-model (gray-box) animation exported from Maya/Blender as the reference video, and the AI renders the final film from your prompt — camera moves, blocking, and movement paths strictly follow the white model, while materials, lighting, and characters come from your words. Coarse white models (simple geometric primitives) work best today: spell out "reference statement + which model maps to which character + detailed plot + scene treatment + global wrap-up." For white models with limbs or wings, write out the full motion sequence or they'll come out stiff.

White Model into an Elf Queen

Prompt

Reference the camera movement and physical lighting changes of Video 1, and precisely replace the white model in the video with the elf queen from Image 1. Her appearance strictly follows the image: long silver-white hair, a magnificent white-and-gold embroidered robe, and a magic staff topped with a glowing crystal — keep the character consistent, no clipping, no face changes. The setting is an elven great hall, keeping the scene from Video 1. The camera positions match the original video, and the elf queen gracefully raises her magic staff following the model's movements in the reference video. As she moves, her silver-white hair and the wide skirt of her white robe flow naturally with the air currents, silky with real drape. Along with the motion, the lighting effects from the reference video transform into a sacred pillar of light rising from the ground, echoing the dazzling golden glow of the staff in her hand. The final video requires cinematic CG photoreal quality, sacred and epic lighting, a high-definition picture, finely detailed skin, crisp metallic textures on the costume, rich detail, and fluid, natural motion.

Requires 1 white-model video + 1 character reference image; "precisely replace the white model with the character from Image N" is the core phrasing of white-model rendering.

5.3 Green-Screen Generation and Replacement

Both directions are supported: turn an ordinary video's background into a green screen (so you can composite it yourself in post), or seamlessly composite a person from green-screen footage into a brand-new scene — foreground separation fine to individual hairs, with the person's perspective and lighting automatically fused into the new background, no flicker, no popping.

One-Click Green-Screen Plate

Prompt

Turn the entire white background of the video into a pure green-screen background. The person's movements, the foreground elements, and the frame composition stay unchanged; the edges of the person are clean with no leftover white fringe, and hair detail stays intact. Remove the audio and keep it silent.

Requires 1 source video, and the audio toggle is already switched off for you at generation time; once you have the green-screen plate, swap the background in any editing software.

5.4 Seamless Two-Video Transitions

Upload the before and after clips, and the model automatically generates the bridging frames that fill the gap — without modifying the source footage. Advanced techniques like action-triggered transitions, eyeline pulls, and occlusion wipes are all supported. Your prompt only needs to describe what happens at the seam.

A Carp Leaps the Dragon Gate

Prompt

Seamlessly join Video 1 and Video 2 without modifying either one. The carp in Video 1's final frame starts swimming rapidly forward, then leaps quickly out of the water into the air, where a heavy, ancient stone gate appears; the carp crosses over the gate into the clouds and mist, finally transforms into a dragon, and the shot connects to Video 2.

Requires 2 videos (the opening and closing clips); the line "without modifying the source videos" is mandatory, and the transition description covers only the missing middle stretch.

Seedance 2.5 Parameter Cheat Sheet (vs 2.0)

Below are all the tiers actually offered for Seedance 2.5 in the Molyin generator, side by side with 2.0 Standard.

Seedance 2.5Seedance 2.0
ModesText to Video / Image to Video / Reference to VideoText to Video / Image to Video / Reference to Video
Resolutions480p / 720p480p / 720p / 1080p / 4k
Duration4–30 s4–15 s
Aspect Ratios16:9 / 9:16 / 4:3 / 3:4 / 21:9 / 1:116:9 / 9:16 / 4:3 / 3:4 / 21:9 / 1:1
AudioSupportedSupported
First/last frameSupportedSupported
Reference images1–91–9
Reference videosUp to 3Up to 3
Reference audioUp to 3Up to 3
Return last frameSupportedSupported

FAQ

The eight most common ways clips go wrong — and the fixes we've verified.