Most people who try Seedance 2.0 notice the same thing: type "a man running through the streets, very cinematic" and you're gambling — regenerate five times, hope one sticks. Meanwhile, some prompts land a near-final shot on the first try. The difference isn't luck. It's understanding what the model actually reads.
Seedance 2.0 ingests your text, images, videos, and audio together, and internally splits everything into two dimensions: a spatial layer (what's in the frame) and a temporal layer (how things change over time). It isn't a copywriting assistant that responds to mood words — it's a multimodal director waiting for a work order. A good prompt reads like production notes, not poetry: who, where, doing what, how the camera moves, in what order.
This post condenses what we've learned across hundreds of generations into five rules and a troubleshooting table. Read it once and you'll waste far fewer generations.
Step one: pick the right task verb
If you're generating from reference material (images, videos, audio), decide which kind of task you're doing first — your sentence structure tells the model how to interpret everything else:
| Task | What you want | Recommended phrasing |
|---|---|---|
| Multimodal reference | Extract elements (a subject, a motion, a style, a voice) and generate a brand-new video | "Referencing the subject in <Image 1>, generate…" |
| Video editing | Add, remove, or change elements in an existing video; everything else stays | "Strictly edit <Video 1>, changing A to B" |
| Video extension | Continue an existing video forward or backward in time | "Extend <Video 1> forward, generating…" |
Here's the trap almost everyone falls into: for editing and extension tasks, say "Video 1" — never "referencing Video 1." That one extra word can make the model misread your edit as a reference task, and you'll get a new video that only vaguely resembles the original.
The three task types also combine: "Referencing the camera work in <Video 2>, strictly edit <Video 1>…"
The core formula
For text-to-video or any complex task, organize your prompt around this skeleton:
precise subject + action detail + scene + lighting & tone + camera work + visual style + quality + constraints
The order has logic behind it: first lock down who is doing what (the temporal layer's protagonist), then establish where and in what atmosphere (the spatial layer), then tell the model how to shoot it, and finally tighten the output with style, quality, and constraint words. The five rules below unpack this formula.
Rule 1: Name your cast
A reference image usually contains more than one thing, and the model doesn't know who "he" is. So the first move is subject definition:
Define the woman in Image 1 wearing a red dress and a straw hat as Rosa.
Three things matter:
- Define with 2–3 stable, static features (clothing, hairstyle, category) so the subject is uniquely identifiable;
- Once defined, use the same label consistently — don't alternate between "Rosa" and "the woman";
- In simple scenes without formal definitions, bind subject to source every time you mention it, using
subject@Image N— e.g., "the barista@Image 1".
Multiple characters get separate definitions and separate labels: "Define the tall man in Video 1 as the cop, and the short man as the thief." From then on, the chase scene always says cop and thief, and references never tangle.
One advanced move: if a character's face matters, prepare a dedicated head-only closeup and spell out the division of labor — "Leo's facial features follow Image 1 (headshot); his outfit and styling follow Image 2 (full-body shot)." This is the single most effective defense against mid-video face swaps.
Rule 2: Write shots, not vibes
The model's internal representation decouples space and time, so the ideal prompt shape for a complex video is a timeline of shots: break the video into Shot 1, Shot 2, Shot 3, described in the order events happen.
Bad: "A man runs nervously through the streets, very cinematic." The model can only guess.
Good:
Shot 1: Side view of an alley; the man breaks into a slow run, breathing hard.
Shot 2: He crashes into a fruit stand; the camera whips over to a closeup of his panicked face.
Shot 3: He vaults a low wall and disappears; the camera slowly pulls back and holds on the empty street.
Within each shot, cover four things in a fixed order:
- Camera move or transition ("medium shot, slow push-in", "cut to…")
- Subject's action and expression
- Position or spatial change
- Audio (sound effects, dialogue, music)
One caveat: don't write exact timestamps like "0–3 seconds." The model's support for precise timing is unstable, and forcing durations tends to produce artifacts. Let it pace the story naturally.
Rule 3: Choreograph, don't label
The principle for actions: specific body parts + quantified intensity. Get concrete about hands, legs, head — and add amplitude, speed, and force: "slowly raises a hand," "whips her head around," "pushes off hard," instead of a vague "he moves."
Two hard-won lessons:
- Prefer slow, gentle, continuous small movements. Walking slowly, lightly raising a hand, easing down into a chair — these succeed far more often than sprinting, leaping, or violent tumbling;
- Write the transitions between actions. "Carried by the momentum of his turn, he raises a hand" flows far better than two disconnected actions.
The same goes for emotion — don't write abstract labels; externalize feelings into physical detail:
| Abstract | Externalized |
|---|---|
| She is sad | Head bowed, shoulders trembling slightly, fingers unconsciously clutching the hem of her jacket, tears welling but not falling |
| He is nervous | Checking his watch again and again, fingers drumming the table, breathing shallow, eyes darting away |
| She is relieved | A long exhale, tense shoulders finally dropping, a faint smile as she lifts her gaze to the distance |
Rule 4: One camera move per shot
Seedance 2.0 understands standard cinematography vocabulary very well — just use it: medium shot, closeup, wide shot, slow push-in, smooth lateral tracking, locked-off camera, handheld follow…
The one iron rule: specify only one camera move per shot. Asking for push, pull, pan, and track simultaneously visibly destabilizes the frame. For complex camera choreography, split it into multiple shots — or better, hand the model a camera-reference video.
Rule 5: Close the frame with constraints
The final paragraph of your prompt exists to cage the randomness. Three word groups, each with a job:
- Quality: high definition, rich detail, cinematic texture, natural color, soft lighting;
- Style: name the target look explicitly — "2D anime style," "vintage film," "cyberpunk cold blue-violet palette." This matters most when your reference image is photorealistic but you want an animated result; leave it out and the output drifts toward live-action;
- Constraints: keep this trio on permanent duty — "no subtitles," "no watermark," "no logo." For multi-character scenes, add a global line: "Throughout the video, never show duplicate characters with identical appearance and clothing."
Building your reference kit: 4–5 assets, each with a job
More reference material is not better. Assign each asset one of four functional roles:
- Character anchor: locks appearance (headshot + full-body, 1–2 images)
- Scene setter: locks environment and art direction (1 image)
- Camera reference: locks shot language and motion rhythm (1 video)
- Mood control: locks emotion and voice timbre (1 audio clip)
Four to five assets total is the sweet spot. Maxing out the upload limit actively hurts — with too many inputs the model can't rank feature priority, and you get style conflicts and subject confusion. And one counterintuitive rule: never use multi-view character sheets (front/side/back in one image). The model tends to read different angles as different people, which makes face drift and "twin" glitches worse.
Ordering matters too: the more precisely an asset must be reproduced, the earlier it should appear in your prompt.
Punctuation the model understands
Seedance 2.0 has a symbol convention for distinguishing information types. Using it correctly measurably improves comprehension:
| Type | Symbol | Example |
|---|---|---|
| Music | () | (Fast-paced rock music plays in the background) |
| Sound effect | <> | |
| Dialogue | {} | {Hello, world}; for languages other than Chinese/English, name the language: says in Japanese {こんにちは} |
| Caption / title card | 【】 | 【Chapter One: Departure】 |
Keep dialogue in one language throughout — don't mix languages mid-script (proper nouns excepted).
Troubleshooting table
| Symptom | Fix |
|---|---|
| Character's face changes mid-video | Add a dedicated head-only closeup; write "face from Image 1, styling from Image 2"; put critical references first; drop multi-view sheets |
| Identical "twins" appear in frame | Bind every character to its image; add the global no-duplicates constraint; use single-person photos; don't paste an entire script as the prompt |
| Style drifts toward photorealism | Name the style explicitly ("3D stylized animation"); if needed, convert the reference image to the target style first |
| Unwanted subtitles appear | Add "no subtitles"; strip text from reference material; landscape output triggers subtitles far less often than portrait |
| Visible jump at extension seams | End each segment on a cut; in editing, trim ~6 frames from the end of the first clip and 1 frame from the start of the next |
| Quality degrades over chained extensions | Limit how many times you extend; use HD sources; advanced trick — convert the source into a white-model (untextured 3D) video first, then extend |
| More than 4 characters gets chaotic | Group characters into images first (max 4 per image), then go image-to-video |
| Special effect doesn't match intent | Don't describe the effect in words — supply an effect reference video |
One decision worth its own paragraph: when to extend vs. when to stitch? For dialogue-driven scenes in a single location — long conversations, emotional build — use video extension for an immersive long-take feel. For plot turns or fast action — chases, fights, montage — generate segments separately and cut them together to protect rhythm and impact. Real projects usually combine both.
A complete example: night-market noodle stall
Here's every rule applied in one prompt.
Assets: Image 1 = vendor headshot; Image 2 = vendor full-body shot; Image 3 = night-market stall scene; Audio 1 = upbeat street-market background music.
Define the man in Image 1 and Image 2 as Leo; his facial features follow Image 1 (headshot), his outfit and styling follow Image 2 (full-body shot). The scene follows the night-market noodle stall in Image 3; background music blends with Audio 1.
Shot 1: Locked-off medium shot. Leo stands behind the wok station tossing noodles; flames leap up as his wrist flicks in quick, controlled bursts, a sheen of sweat on his forehead. (Upbeat street-market music plays in the background)
Shot 2: The camera slowly pushes in to a close shot. Leo plates the noodles and hands them over, the corner of his mouth lifting, and says to an off-screen customer {Eat it while it's hot}.
Shot 3: The camera slowly pulls back to a wide shot. The stall's sign lights up against the night, crowds drift past, and Leo turns back to the wok.
Overall: high definition, rich detail, warm tones, cinematic texture, soft lighting. No subtitles, no watermark, no logo. Only one Leo appears throughout; never show duplicate characters with identical appearance.
Map it back: task verb (reference), subject definition with divided duties (headshot owns the face, full-body owns the styling), three shots each covering camera/action/space/audio, one camera move per shot, correct symbols, and the quality + style + constraint close.
Image-to-video: two more ready-to-use prompts
The noodle-stall example shows the full-asset-kit approach. But image-to-video usually starts far more modestly: you have one or two pictures and want them to move. The two prompts below — one simple, one full-scale — demonstrate a judgment call that's easy to miss: not every prompt needs a shot list. A single continuous action in one location fits in one paragraph; only a multi-event story across spaces earns the full "setup + shots + constraints" structure.
Example 1 — single image, micro-motion: bringing an old photo to life
Assets: Image 1 = an old photo of an elderly woman sitting in a rattan chair in a courtyard.
Referencing the elderly woman in Image 1 — silver-haired, in a deep-blue cotton blouse, seated in a rattan chair — generate a scene of her in the courtyard on a quiet afternoon: she gazes calmly ahead, then her eyes blink slowly, the corners of her mouth lift into a faint smile, she tilts her head slightly toward the right of the frame and, carried by that motion, raises her right hand to gently tuck a stray strand of hair behind her ear; behind her, drying laundry and tree shadows sway softly in the breeze. Locked-off medium-close shot. Vintage film texture, warm tones, soft lighting; face stable and undistorted, movements slow and continuous; no subtitles, no watermark, no logo.
Check it against the rules: the subject is locked with three static features (rattan chair, deep-blue blouse, silver hair); every action is a slow, continuous micro-movement with a written transition (the head tilt carries into the hair tuck); exactly one camera setup throughout; and one closing sentence folds in quality, stability, and the constraint trio. A simple request should be this short — forcing a shot list onto it just gives the model extra ways to improvise.
Example 2 — multi-image storyboard: a rain-lane short
Assets: Image 1 = woman's headshot; Image 2 = her full-body shot in a qipao; Image 3 = a rain-soaked stone lane in an old southern town. Mind the upload order: the face — the asset that must be reproduced most precisely — goes first.
Overall setting: a misty, rain-veiled stone lane in an old southern town, cinematic realism. Define the woman in Image 1 and Image 2 as Mei: her facial features follow Image 1 (headshot), her outfit and styling follow Image 2 (full-body shot); the scene follows the stone lane in Image 3, with Image 3 as the opening frame.
Shot 1: Locked-off wide shot. Fine rain falls on the stone pavement as Mei walks slowly in from the mouth of the lane under an oil-paper umbrella, steps unhurried, the hem of her qipao swaying gently with each stride.
Shot 2: The camera slowly pushes in to a medium-close shot. Mei stops before a wooden door, tilting the umbrella slightly as she turns her head, her gaze settling on the brass knocker, tiny raindrops clinging to her lashes.
Shot 3: Smooth lateral tracking. Carried by the momentum of raising her arm, Mei folds the umbrella in one motion, turns, and steps through the doorway; the frame holds on the hazy lantern light deep in the lane. (Slow guzheng music rises in the background)
Overall: high definition, rich detail, film-grain cinematic texture, cool low-saturation palette, layered lighting; face stable and undistorted, movements slow, continuous, and natural, no clipping or stutter; no subtitles, no watermark, no logo; only one Mei appears throughout — never show duplicate characters with identical appearance.
This prompt threads all five rules together and adds one technique specific to image-to-video: declaring the opening frame. "Image 3 as the opening frame" lets the video unfold from the scene image's composition — empty lane first, character enters after — far more stable than letting the model guess the opening. Note the emotion writing too: the word "wistful" never appears; what's written is "gaze settling on the brass knocker, raindrops clinging to her lashes."
Text-to-video: writing with no images at all
Every example so far leaned on reference assets. But what if you have nothing — pure text-to-video? The formula is still the skeleton from the top of this post, with one key difference: there's no image to bind to, so the subject and style live or die on your words — pack the features denser. Same pairing: one simple, one full-scale.
Example 1 — a cat outside a convenience store on a snowy night (one-paragraph assembly)
Late at night outside a convenience store, an orange-and-white shorthair cat sits on the doorstep as fine snow drifts down. Its ears flick now and then, shaking snowflakes off their tips; the end of its tail sweeps unhurriedly side to side; each breath condenses into a small puff of white in the cold air. Behind it, warm yellow light spills through the glass door, and snowflakes sink slowly through the glow. Locked-off close shot. High definition, rich detail, cinematic texture, cold-warm night contrast, soft lighting; fur detail stable, movements slow and continuous; no subtitles, no watermark, no logo.
With no reference image, "orange-and-white, shorthair, sitting on the doorstep" is this cat's entire identity — feature density directly determines subject stability. And the actions are still all slow micro-movements (ear flicks, tail sweeps, visible breath), which matters even more with animal subjects.
Example 2 — a coffee shop ad (shot list + on-screen text)
Overall setting: an independent coffee shop in early morning, warm cinematic tones, natural light slanting in through large glass windows.
Shot 1: Locked-off closeup. Deep-brown coffee beans pour slowly from above into the metal hopper of a grinder, bouncing as they land, fine dust floating in the backlight.
Shot 2: The camera slowly pushes in. In a white ceramic cup, latte art blooms from the center into a leaf as steam curls upward, catching the window light.
Shot 3: The camera slowly pulls back to a wide shot. The morning-lit shop sits clean and quiet, the latte steaming on the counter; in the lower third of the frame, white handwritten text fades in: "First a sip, then the day." (Gentle acoustic guitar in the background)
Overall: high definition, rich detail, warm tones, cinematic texture, soft lighting; liquid and steam behave naturally, no distortion; no text other than the specified line, no watermark, no logo.
This one hides a rule that's easy to get backwards: when you want text on screen, the constraints flip. Drop "no subtitles" and instead spell out the text's four essentials — content, timing, position, entrance — then close with "no text other than the specified line" as the backstop. Slogans are the most common commercial use of text-to-video; this reversal is worth memorizing.
Closing thoughts
Seedance 2.0 moved the bottleneck of video generation from model capability to your ability to communicate. You don't need to know how to shoot a film — you need to write prompts the way a director writes production notes: name the cast, tell the story shot by shot, choreograph the movement, one camera move per shot, and close with constraints.
The rest is practice. That noodle-stall prompt above? Take it, tweak it, make it yours.