01 The Core Formula
Seedance 2.0 is essentially a multimodal AI director: it reads your text, images, video, and audio at once, and internally splits your prompt into a "space layer" (what's in the frame) and a "time layer" (how things change over time). A good prompt is therefore not prose — it's an engineering instruction: who, where, doing what, how the camera moves, in what order. This chapter breaks the official advanced formula into reusable parts, then teaches you to organize multi-shot narratives with numbered shots.
1.1 The Basic Prompt Formula
The official advanced formula has eight parts: precise subject + action detail + scene + light and color + camera + visual style + image quality + constraints. You don't need all eight on day one — subject plus motion are the two required parts, the rest stack on as needed. But make the closing constraints a habit: negative instructions like "keep it free of subtitles" and "do not generate watermarks" measurably cut the failure rate.
[Subject] + [Motion] + optional [Environment] [Aesthetics] [Camera] [Audio] [Constraints]
Subject
The star of the frame — a person, animal, or product. The more specific, the better.
Motion
What the subject does, with body parts and pace spelled out.
Environment
Setting and mood: place, time of day, lighting, weather.
Aesthetics
Art direction and image quality: palette, texture, clarity.
Camera
Framing and camera movement: close-up, push-in, dolly — one move per shot.
Audio
Voice, sound effects, and background music.
Constraints
Closing negative instructions: "keep it free of subtitles", "do not generate watermarks", "do not generate logos".
An Orange Cat on the Windowsill
Prompt
An orange cat dozes on a windowsill bathed in afternoon sunlight, the tip of its tail twitching now and then as a breeze lifts a corner of the white gauze curtain. The camera slowly pushes in to a close-up of the cat's face, the afternoon sun tracing a golden rim along its fur. Warm, cozy mood, cinematic quality, soft natural light, rich detail. Keep the video free of subtitles; do not generate any watermark.
Two sentences for subject + motion, one each for environment, camera, and image quality, then constraints to close — the official advanced formula in one full pass.
1.2 Defining Subject and Motion
The ceiling of the model's understanding is your action description. Three official rules. One: break motion into body parts with intensity — write to the hands, shoulders, head, and qualify with "slowly", "firmly", "slightly". Two: prefer slow, continuous small motions; explosive moves like sprinting and leaping break easily. Three: write the transition between consecutive actions, e.g. "riding the momentum of the turn, she raises her hand." And externalize emotion through physical detail — "her fingers unconsciously grip the corner of her shirt" always beats "very nervous".
The Barista's Latte Art
Prompt
A barista in a dark green apron stands behind a wooden counter, holding a milk pitcher with both hands, tilted slightly, his wrist swaying at an even pace as he slowly draws a leaf pattern into the foam. Finishing, he straightens up and, riding that motion, glances down at the cup — the corners of his mouth rise despite himself. Static medium shot, warm café interior, softly blurred background, gentle lighting. His face stays stable and undistorted, motion natural and fluid, no stutter, no flicker; keep the video free of subtitles.
Motion lands on the wrist with pace quantified ("swaying at an even pace"); "riding that motion" is the official action-transition phrasing; emotion is externalized through "the corners of his mouth rise"; constraints close the prompt.
1.3 Shot Sequencing and Camera Work
For anything beyond one shot, number your storyboard: "Shot 1 / Shot 2 / Shot 3" — each shot covers, in order, the camera move, the subject's action, the position change, and the audio. Note the official warning: the model's support for exact seconds is unstable, and forcing durations can produce abnormal results — use shot numbers and let the model pace the narrative (this is the opposite of Seedance 2.5's second-level timestamps; don't mix the two). And one camera move per shot — stacking push, pan, and track only buys you jitter.
A Dorm-Room Mini Drama (Official Field Case)
Prompt
The girl's appearance references Image 1, the dorm-room scene style references Image 2, the camera work references Video 1, and the ambient sound references Audio 1. Shot 1: in the evening, the girl walks briskly to the dorm door; the camera follows smoothly in a medium shot, warm yellow daylight spilling into the hallway through the windows; she pauses at the door, takes a deep breath, faintly nervous. Shot 2: she pushes the door open and walks in; cut to an indoor medium shot — her roommates look up from their books, and one of them asks with a grin {How did the exam go? Did you pass?}; the camera slowly alternates half-body close-ups among them. Shot 3: the girl first lowers her head with a crestfallen look — the camera moves to a close shot — then she looks up, unable to hold back her smile, and bursts out laughing {Got you!}; the roommates chase and roughhouse; the camera slowly pulls back and settles on a wide shot of the dorm full of laughter. Throughout: high-definition cinematic documentary style, warm palette, soft light; faces stable and undistorted, motion natural and fluid, no stutter, no flicker; the ambient sound blends naturally with Audio 1; keep the video free of subtitles.
Requires 2 reference images + 1 camera-reference video + 1 ambient audio clip. The full skeleton of an official field case: asset binding up front, numbered shots, dialogue in {} symbols, image-quality and constraint words to close.
1.4 Quality, Style, and Constraint Words
Three kinds of closing words each do one job: style words set the art direction, quality words hold the clarity floor, and constraint words close the loop — "keep it free of subtitles", "do not generate watermarks", "do not generate logos" are the official constraint template trio, and they measurably cut the failure rate. The example below stacks all three.
A Cyberpunk Rainy Alley
Prompt
Cyberpunk style with a cold blue-violet palette: in a narrow rainy alley, neon signs and holographic ads flicker on both sides as a figure in a glowing jacket walks slowly toward the far end, the wet ground mirroring the lights. The camera pans smoothly to follow. Cinematic quality, rich detail, saturated color, layered lighting. Avoid generating any text or subtitles, no logos, no watermark.
Style words open to set the tone, the official constraint trio closes — the fixed move for cutting subtitle and watermark failures.
02 Text in Video
Slogans, subtitles, speech bubbles — Seedance 2.0 renders readable text right into the frame. Learn the official symbol protocol first: dialogue goes in {}, subtitles in 【】, music in (), sound effects in <> — the symbols tell the model what kind of information each span is. Use common characters for on-screen text (rare characters and special symbols break easily), and keep dialogue in one language — no mixing Chinese and English except proper nouns.
2.1 Slogans and Title Text
The official slogan template is a five-piece set: text content + when it appears + where it appears + how it appears + text features (color, style). The model matches a fitting typeface on its own; when you're strict, name the color and style directly. Close with "no other text beyond the specified wording" to keep stray signs and captions out.
[Text Content] + [Timing] + [Position] + [Entrance] + [Style]
Text Content
The exact wording to appear in frame, spelled out letter by letter.
Timing
When the text appears, e.g. "in the second half of the shot".
Position
Where the text appears, e.g. "center of frame", "bottom".
Style
Color, typeface personality, and motion, e.g. "bright-yellow handwritten letters popping in".
A Lemon-Tea Ad Slogan
Prompt
Hand-drawn comic style: three friends sit around a camping table, drinking the lemon tea from Image 1 together, the mood friendly and relaxed. In the second half the frame gradually blurs, and the text "Summer Only, Always Refreshing" emerges at the center of the frame, in a rounded handwritten typeface, bright yellow, with a slight bouncing motion. No text or subtitles other than the specified wording appear anywhere in the frame.
Requires 1 product reference image. The official five-piece set in one pass, closed with the key line that keeps stray text out.
2.2 Subtitles
Subtitles and voiceover are natural partners: write the line in {} letter by letter, then mark the subtitles with 【】 or use the official template sentence — "subtitles appear at the bottom of the frame, fully in sync with the audio rhythm." Quoting the template verbatim is the steadiest approach.
A Voiceover for the Dawn of the Universe
Prompt
Generate a video with voiceover. A deep, calm male voice says {In the grand universe, our world is but a fleeting moment. Yet within it, life flourishes against all odds.} The scene slowly transitions from a star-dense night to dawn — the stars fade one by one, the sun rises from behind the mountain ridge, and the sea of clouds is dyed gold. Subtitles appear at the bottom of the frame, fully in sync with the audio rhythm.
The closing sentence is the official subtitle template verbatim; with the line written in {}, the subtitles align to the voice automatically.
2.3 Speech Bubbles
Comic-style bubbles grow a character's line right into the frame. The official phrasing: "the character says {the line}, and a bubble appears around the speaking character with the corresponding line inside." It suits light narrative — short dramas, marketing skits — and the shorter the line, the better it lands.
Three Lines on the Running Track (Official Case)
Prompt
The two people from Image 1, dressed in sportswear, are running on the school track. The girl turns to the boy and says with a confident smile {We can definitely do it!}; the camera cuts to a close shot of the boy, who answers hesitantly {Are you sure?}; the camera cuts back to a medium close-up of the girl, who says brightly {Yes!}, the mood bright and firm. A bubble appears around each speaking character with the corresponding line inside. No text other than the bubble lines appears anywhere in the frame.
Requires 1 two-person reference image. The camera switches between the two speakers with each line — one short line per person is all a few seconds can hold.
03 Image References
One reference image beats a hundred adjectives. Upload product shots, character images, or a logo, then call them out in the prompt as "Image 1", "Image 2" — the model locks appearance and style from your assets. Upload order is the numbering order. For multi-subject scenes, use the official definition phrasing: "define the woman in the red dress and straw hat from Image 1 as Subject 1", then keep using the same label — never alternate between a name and "that woman". And more assets is not better: the official recommendation is a 4–5 asset setup (1–2 character images + 1 scene image + 1 camera-reference video + 1 audio clip); maxing out the slots only confuses the model's feature priorities.
3.1 Multi-Angle Subject References
Orbit shots for e-commerce and cross-scene character reuse both rely on multi-angle reference images to lock appearance: upload 2–3 images of the same subject from different angles and the model builds a complete identity file that survives any scene change. For people, a clean facial close-up plus a full-body shot is enough — avoid multi-view collages, which the model tends to read as different people.
Reference [subject] from [Image N] + generate [scene description] + keep [subject] consistent
Reference Subject
The subject whose appearance to lock, pointed at with "Image N".
Scene Description
What happens in the new scene: action, camera, environment.
Consistency
One line of "stay consistent with the reference images" reinforces the appearance lock.
A Camera in the Round (Official Case)
Prompt
Extract the camera from Image 1, Image 2, and Image 3, replace the background with pure white, and place the camera on a white table. The shot opens with a close-up on the body's material, then slowly orbits the camera for a full turn, clearly showing the front, sides, and back. The product appearance stays consistent with the reference images. Studio-grade lighting, clean frame; do not generate any watermark or text.
Requires 2–3 multi-angle product images. "Extract … stays consistent" is the official fixed phrasing for multi-angle appearance locking, closed with constraint words.
3.2 Multi-Image References
Each image gets one job: one for the character, one for the logo, one for the scene. Call out each role in upload order — "Image 1", "Image 2" — and the model won't cross them. The more important the asset, the earlier it should appear in the prompt.
A Logo Entrance in a Neon City (Official Case)
Prompt
The backdrop is a sky corridor in a neon-lit future city, aircraft and holographic ads interweaving. Referencing the girl from Image 2, open with a medium shot of her releasing a silver levitating lantern with a holographic glow; the camera then slowly pulls back to reveal a sky full of floating lanterns; the frame gradually blurs, and the logo from Image 1 emerges, centered, and holds. Overall in a 3D cyberpunk sci-fi animation style with a cold blue-violet palette; the girl's appearance stays consistent with the reference image.
Requires 2 images (logo + character); the logo lands right after the blur transition for maximum ceremony.
Five Assets, Five Jobs in a Café (Official Case)
Prompt
The scene is set inside the restaurant from Image 4, with customers coming and going. The girl from Image 1, wearing the outfit from Image 2, is tidying items on the counter. The boy from Image 3 is a customer who walks up, wanting to ask for her contact. The sign from Image 5 stays visible in the bottom-right corner throughout. Faces, hairstyles, and outfit details stay consistent with the reference images; the picture is natural and fluid; keep the video free of subtitles.
Requires 5 reference images (girl / outfit / boy / scene / sign). One job per image, each called out by number — the official standard for multi-element references.
04 Video References
The things words can't pin down — fight choreography, dive-bomb camera paths, particle effects — hand them to a reference video. The model replicates the motion, camera work, or effect morphology; you swap in your own subjects and scenes. Upload order is the numbering order for "Video 1", "Video 2". Remember the asset-budget rule: 4–5 assets is the sweet spot; maxing out the slots invites style conflicts and blurry subject recognition.
4.1 Motion References
For complex motion — fights, dance, sprints — pure text always loses something. Upload a motion-reference video and the model replicates its action rhythm and camera language; you bring your own characters and setting.
A Two-Person Fight Scene (Official Case)
Prompt
Reference the character movements and camera language of Video 1, and generate a fight between the two characters from Image 1 and Image 2: Image 1 is the fighter on the right, Image 2 the one on the left. The exchanges are tight and fast-paced, the camera rides the rhythm of the action, (intense percussion playing in the background). Faces stay consistent with the reference images; motion stays fluid, never stiff, no clipping, no stutter; keep the video free of subtitles.
Requires 1 motion-reference video + 2 character images. Motion and camera go to the video, faces to the images, and the music is marked with the () symbol — the official division of labor.
4.2 Camera-Movement References
For camera trajectories like FPV dives and orbiting ascents, one reference video is more precise than any camera term. The model copies the path; you only name the new visual anchor and the scene style.
An FPV Dive Through a Tech Campus (Official Case)
Prompt
Reference the camera movement of Video 1 to make a concept video for a tech campus: with the high-rise from Image 1 as the visual anchor, the same first-person dive — plunging from above the clouds, threading the sky bridges between the towers, finally skimming low across the water — conveying the campus's high-tech, futuristic feel. Footage smooth and steady, cinematic quality; keep the video free of subtitles, do not generate any watermark.
Requires 1 camera-reference video + 1 scene image. "The same first-person dive" is the official phrasing — the trajectory is copied from the reference video, far more precise than words.
4.3 Effects References
Particles, light effects, transformation trajectories — an effect's shape and motion logic resist verbal description. The official field conclusion is unanimous: when an effect misses, feed the target effect video in as a reference so the model grasps its exact morphology — the hit rate jumps immediately.
Golden Particles Around a Flute (Official Case)
Prompt
Reference the golden particle effect from Video 1: as the costumed character from Image 1 plays a bamboo flute, the same golden particles swirl around them, gathering and scattering in rhythm with the melody. The face stays consistent with the reference image, the motion natural, the overall feel that of a wuxia film; keep the video free of subtitles.
Requires 1 effects-reference video + 1 character image; pointing at an effect with a video beats piling on adjectives.
05 Video Editing
Upload an existing video and add, remove, or change elements, extend it forward or backward, or stitch clips — all with one instruction. Pick the model by the job: for add/remove/change use Seedance 2.0 Video Edit — upload a 3–15 second source clip in video edit mode and the output duration follows the source; for extensions use Seedance 2.0 Extend — each pass adds 4–15 seconds. Two-clip stitching still runs in reference-to-video mode with both clips attached as references. One writing rule holds throughout: editing tasks address the clip as "Video 1" directly — never "reference Video 1", which gets misread as a reference task instead of an edit.
5.1 Adding, Removing, and Changing Elements
These jobs go to the Seedance 2.0 Video Edit model. The official phrasing differs by operation: for additions, spell out "element features + where they appear"; for changes, use "replace A in Video N with B from Image N"; for removals, name what goes — and also emphasize what must stay, which the official guide found performs measurably better. For replacements, attach a reference image of the new element for a more accurate look.
In [Video N] at [time/space position] + add/remove/replace [element description]
Edit Operation
Add, remove, or replace — name the operation type in one phrase.
Time Position
When in the video it happens, e.g. "in the second half".
Space Position
Where in the frame, e.g. "on the left side of the table".
Adding Snacks to the Table (Official Case)
Prompt
Add fried chicken, pizza, and other snacks onto the table surface in Video 1. The food sits naturally, its lighting matches the original footage, and everything else stays unchanged.
Requires 1 source video. The official addition phrasing = element features + position, closed with "everything else stays unchanged".
Clearing the Table (Official Case)
Prompt
Remove the other parts and tools from the table in Video 1, keeping the tabletop neat and clean — only the items in the two people's hands remain. Their movements, the lighting, and the camera movement stay completely unchanged.
Requires 1 source video. The official removal rule: name what goes, and emphasize what must stay — it performs measurably better.
Swapping the Perfume for a Cream (Official Case)
Prompt
Replace the perfume in Video 1 with the face cream from Image 1. The actions and camera movement stay unchanged; the replacement's perspective and reflections blend naturally into the original footage, and everything else stays the same.
Requires a source video + 1 replacement reference image. The official change phrasing = "replace A in Video N with B from Image N" + actions and camera unchanged.
5.2 Video Extension
Switch to the Seedance 2.0 Extend model and pick 4–15 seconds of new footage on the slider. Extend backward for a prologue or forward for an outcome; audio-visual style and subjects carry over automatically. Describe only the new segment — the model auto-trims the join and the source footage is never regenerated.
A Preceding Over-the-Shoulder Reassurance (Official Case)
Prompt
Extend Video 1 backward: give the man in white an over-the-shoulder shot, and he says gently {It's not that bad. You're just stressed. Everyone goes through this, you just need to keep going.} The indoor lighting stays consistent with the original footage.
Requires 1 source video. Write "extend Video 1 forward/backward" as-is — never add the word "reference", which gets misread as a reference task.
Continuing to the Friends' Reunion (Official Case)
Prompt
Generate what happens after Video 1: the two latecomers run up to them, the five friends finally meet, and chat warmly. Appearances, setting, and lighting carry over from the original footage; the motion is natural and continuous.
Requires 1 source video. A forward extension describes only the new segment — the model auto-trims the join and never regenerates the source footage.
5.3 Bridging Two Clips
Give the model an opening clip and a closing clip, and it fills in the missing transition. Your prompt only needs to describe what happens at the seam — a gust of wind, a particle transition, a match cut all work as glue. Up to 3 input videos are supported, with a combined duration of no more than 15 seconds.
A Falling-Leaf Transition (Official Case)
Prompt
Video 1. The instant the falling leaf hits the ground, a ring of golden particles bursts outward and a gust of wind sweeps across the frame — then cut to Video 2.
Requires 2 videos (opening and closing); describe only the missing middle stretch — the shorter, the more focused.