01 The Core Formula
Seedance 2.0 understands structured direction: who, what they do, where, and how it's shot. This chapter breaks a good prompt into six reusable building blocks, then shows you how to script multi-shot stories with shot numbers — once you have this skeleton, every later chapter is just an extension of it.
1.1 The Basic Prompt Formula
Everything starts with one sentence: subject plus motion are the two required blocks, the other four stack on as needed. Play the clip in your head first, then write it down following the formula.
[Subject] + [Motion] + optional [Environment] [Aesthetics] [Camera] [Audio]
Subject
The star of the frame — a person, animal, or product. The more specific, the better.
Motion
What the subject does, with body parts and pace spelled out.
Environment
Setting and mood: place, time of day, lighting, weather.
Aesthetics
Art style and image quality: color grade, texture, sharpness.
Camera
Framing and camera movement: close-up, push-in, pan — one move per shot.
Audio
Voice, sound effects, and background music.
Cat on a Sunlit Windowsill
Prompt
An orange cat dozes on a windowsill bathed in afternoon sunlight, the tip of its tail twitching now and then as a breeze lifts the white gauze curtain. The camera slowly pushes in to a close-up of the cat's face. Warm, cozy mood, soft natural light, rich detail.
Two sentences for subject + motion, one each for environment and camera — that's a solid basic prompt.
1.2 Defining Subject and Motion
The ceiling of the model's understanding is your motion description: name the hands, shoulders, and head; add pace and weight; and externalize emotion through physical detail — "her fingers absently grip the hem of her shirt" always beats "very nervous." Favor slow, connected micro-motions; explosive, large-scale movement is where clips break.
The Barista's Latte Art
Prompt
A barista in a dark green apron stands behind a wooden counter, tilting a milk pitcher with both hands and gently swaying his wrist to draw a leaf pattern into the foam. When he finishes, he glances down and can't hold back a smile. Static medium shot, warm café interior, softly blurred background, gentle lighting.
Pin motion to specific body parts like the wrist and fingertips, and externalize emotion with details like "can't hold back a smile" instead of abstract adjectives.
1.3 Shot Sequencing and Camera Moves
Anything longer than one shot deserves a numbered storyboard: Shot 1, Shot 2, Shot 3 — each naming its camera move, subject action, and change of position. Don't pin exact seconds; the model paces the story itself. But give each shot only one camera move — stacking push, pull, pan, and track in one shot just makes the footage unstable.
A Rainy-Night Walk Home in Three Shots
Prompt
Shot 1: In a rainy alley at night, a girl with a transparent umbrella walks briskly across the wet pavement. The camera tracks her smoothly from the side as neon signs stretch long reflections on the ground. Shot 2: She stops in front of a coffee shop glowing with warm light, folds her umbrella, and shakes off the raindrops. Cut to a close-up of her face as her expression shifts from tired to hopeful. Shot 3: She pushes the door open, a wind chime rings softly, and the camera slowly pulls back to settle on the lit storefront in the rainy street. Cinematic look throughout, cool blue night against warm yellow light, no subtitles, no watermark.
Number the shots, give each one a single camera move, then close with image-quality and constraint words.
1.4 Quality, Style, and Constraint Words
Three kinds of closing words each do one job: style words set the art direction, quality words hold the sharpness floor, and constraint words seal the edges — negative instructions like "no text or subtitles" and "no watermark" measurably cut your reject rate. The example below stacks all three.
Cyberpunk Rain Alley
Prompt
Cyberpunk style with a cold blue-violet palette: in a narrow rainy alley, neon signs and holographic ads flicker on both sides as a figure in a glowing jacket walks slowly toward the far end, the wet ground mirroring the lights. The camera pans smoothly to follow. Cinematic quality, rich detail, saturated color, layered lighting. Avoid generating any text or subtitles, no logos, no watermark.
Open with style words to set the tone, close with constraint words — a fixed habit that keeps subtitles and watermarks out of your clips.
02 Text in Video
Slogans, subtitles, speech bubbles — Seedance 2.0 renders readable text right inside the frame. The trick is specifying four things: what it says, when it appears, where it sits, and how it looks. Stick to common words; rare characters and special symbols are where things go wrong.
2.1 Slogans and Title Text
You don't need design skills for on-screen titles: state the text content, then its timing, position, and style. The model matches a fitting typeface on its own — or you can name the color and style directly.
[Text content] + [Timing] + [Position] + [Entrance] + [Style]
Text Content
The exact words that appear on screen, written out in quotes.
Timing
When the text appears, e.g. "in the second half of the shot."
Position
Where the text sits, e.g. "center of the frame" or "bottom."
Style
Color, typeface personality, and animation, e.g. "bright yellow handwritten letters fading in."
Iced Tea Ad Slogan
Prompt
Hand-drawn comic style: three friends sit around a camping table, clinking glasses of the iced lemon tea from Image 1, the mood relaxed and cheerful. In the second half the frame gradually blurs, and the text "Summer Only. Stay Fresh." appears at the center in a rounded handwritten typeface, bright yellow, with a gentle bounce-in animation.
Requires 1 product reference image; "Image 1" in the prompt points to your uploaded product photo, and the slogan text works best spelled out letter for letter.
2.2 Subtitles
Subtitles and voiceover are a natural pair: write the spoken lines first, then ask for "subtitles at the bottom of the frame, synced to the narration." Put the lines in quotes and the subtitles will align with the audio on their own.
Dawn Over the Universe, Narrated
Prompt
Generate a video with voiceover. A deep, calm male voice says: "In the vastness of the universe, our world is but a fleeting instant of light. Yet within that instant, life blooms against all odds." The scene transitions slowly from a star-filled night to dawn: the stars fade one by one, the sun rises behind a mountain ridge, and the sea of clouds turns gold. Subtitles appear at the bottom of the frame, synced to the narration, in a thin white typeface.
Quote the narration, ask for subtitles "synced to the lines" at the bottom — picture and audio lock together.
2.3 Speech Bubbles
Comic-style bubbles grow the dialogue right onto the frame: as a character speaks, a bubble follows with the matching line inside. Great for short dramas and playful marketing skits — the shorter the line, the better it lands.
A Sunset Run by the Sea
Prompt
The girl from Image 1, in sportswear, jogs along a seaside path at dusk. She turns toward the camera and says with a confident smile: "One more kilometer and we're done!" Cut to a medium close-up as she picks up the pace and says brightly: "The finish line is right there!" A comic speech bubble appears beside whoever is speaking, with the matching line inside. Warm sunset tones, natural fluid motion.
Requires 1 character reference image; quote each line exactly and the bubble follows the speaker automatically.
03 Image References
One reference image beats a hundred adjectives. Upload a product shot, a character, or a logo, then call it out in the prompt as "Image 1" or "Image 2" — the model locks appearance and style to your assets. Upload order is the numbering order; this chapter covers the two core patterns.
3.1 Multi-Angle Subject References
Orbiting product showcases and reusing a character across scenes both rely on multi-angle references: upload 2-3 shots of the same subject from different angles, and the model gets a complete "appearance file" — the face or product survives any change of scene.
Reference the [subject] in [Image N] + generate [scene description] + keep the [subject] consistent
Reference Subject
The subject whose appearance gets locked, pointed to with "Image N."
Scene Description
What happens in the new scene: action, camera, environment.
Consistency
One line of "consistent with the reference images" reinforces the lock.
360° Headphone Showcase
Prompt
Extract the silver over-ear headphones from Image 1, Image 2, and Image 3, place them on a white pedestal against a pure white background. The camera opens with a close-up on the ear-cup material, then orbits the headphones slowly for a full rotation, revealing the front, sides, and back. Keep the product appearance consistent with the reference images. Studio lighting, clean frame.
Requires 2-3 product photos from different angles; "extract ... keep consistent" is the go-to phrasing for multi-angle identity lock.
3.2 Multi-Image References
Each image owns one job: one for the character, one for the logo, one for the scene. Call out each job in upload order — "Image 1," "Image 2" — and the model won't mix up their roles.
Logo Reveal over a Neon City
Prompt
Set against the neon-lit night sky of a futuristic city: referencing the girl from Image 2, she releases a silver levitating lantern with a holographic glow. The camera pulls back from a medium shot to reveal a sky full of floating lanterns. The frame gradually blurs, and the logo from Image 1 emerges, centered, and holds. 3D cyberpunk sci-fi animation style, cold blue-violet palette.
Requires 2 images — a logo and a character; timing the logo's entrance after the blur transition gives it the most ceremony.
04 Video References
What words can't fully say — fight choreography, dive-bomb camera paths, particle effects — hand over to a reference video instead. The model replicates the motion, camera work, or effect behavior; you just swap in your own subject and scene. Upload order sets the "Video 1," "Video 2" numbering.
4.1 Motion References
Fights, dance, sprints — complex motion always loses something in translation to text. Upload a motion reference video and the model copies its choreography and camera language; you supply the characters and the setting.
Two-Character Fight Scene
Prompt
Reference the character movements and camera language of Video 1, and generate a fight between the two characters from Image 1 and Image 2: Image 1 is the fighter on the left, Image 2 the one on the right. Tight, fast-paced exchanges, the camera riding the rhythm of the action, driving percussion in the soundtrack. Faces stay consistent with the reference images; motion stays fluid, never stiff.
Requires 1 motion reference video + 2 character images; the video owns the choreography, the images own the faces — clean division of labor.
4.2 Camera-Movement References
FPV dives, rising orbits — a reference video nails a camera path more precisely than any camera term can. The model traces the trajectory; you only need to name a new visual anchor and a new scene.
FPV Dive Through a Tech Campus
Prompt
Reference the camera movement of Video 1 to make a concept video for a futuristic tech campus: with the twin towers from Image 1 as the visual anchor, dive first-person from above the clouds, thread the sky bridges between the buildings, then skim low across the water. Emphasize the campus's high-tech, futuristic feel — glass facades catching the morning light, footage smooth and steady.
Requires 1 camera reference video + 1 scene image; letting the video dictate the trajectory is far more precise than describing it.
4.3 Effects References
Particles, glows, transformation trails — the shape and motion logic of an effect is nearly impossible to write down. Hands-on experience points to one conclusion: when a text-described effect misses, point to a reference video instead and your hit rate jumps immediately.
Golden Particles Around a Flute Melody
Prompt
Reference the golden particle effect from Video 1: as the costumed character from Image 1 plays a bamboo flute, the same golden particles swirl around them, gathering, flowing, and scattering in rhythm with the melody. Keep the face consistent with the reference image, motion natural, overall wuxia-film quality.
Requires 1 effects reference video + 1 character image; pointing at the effect in a video beats stacking descriptive text.
05 Video Editing
Upload an existing video and rewrite it with one sentence: add, remove, or swap elements, extend it forward or backward, or bridge two clips. This chapter has exactly one rule: for editing tasks write "Video 1" directly — never "reference Video 1" — or the model treats it as a reference task instead of an edit.
5.1 Adding, Removing, and Swapping Elements
Change the frame without touching the rest of the footage: name the operation, the position, and the new element's traits, then close with "keep everything else unchanged." For swaps, upload a reference image of the replacement and the appearance lands much more accurately.
In [Video N] at [time/position] + add/remove/replace [element description]
Edit Operation
Add, remove, or replace — state the operation in one word.
Time Position
When in the video it happens, e.g. "in the second half."
Space Position
Where in the frame, e.g. "on the left side of the table."
Add Dessert to the Table
Prompt
In Video 1, add a plate of strawberry cake and two hot lattes onto the dining table, gentle steam rising from the cups. The food sits naturally, its lighting matches the original footage, and everything else stays unchanged.
Requires 1 source video; name where the new elements appear and close with "everything else stays unchanged."
Swap the Can for a Sparkling Water Bottle
Prompt
In Video 1, replace the canned drink on the table with the glass sparkling-water bottle from Image 1. Keep the tabletop lighting and camera movement unchanged; the replacement's perspective and reflections should blend naturally into the original footage.
Requires the source video + 1 reference image of the replacement; "replace A with B" is the most reliable swap phrasing.
5.2 Video Extension
Extend backward to add a prologue, forward to continue the story — audio-visual style and subjects carry over automatically. Only describe the newly added stretch; the original footage is never regenerated.
Add a Prologue Before the Scene
Prompt
Extend Video 1 backward: open with an over-the-shoulder shot toward the window, raindrops sliding down the glass, then cut back inside where a girl in a grey sweater slides a cup of hot tea across the table to her friend and says softly: "Warm your hands first. Then tell me everything." Gentle tone, warm interior light, quiet mood.
Requires 1 source video; write "extend Video 1 forward/backward" — and leave the word "reference" out of it.
5.3 Track Completion
Hand over a first clip and a last clip, and the model fills in the missing transition between them. Your prompt only needs to describe what happens at the seam — a gust of wind, a particle wipe, a match cut all work as glue.
Falling-Leaf Transition Into a New Scene
Prompt
Video 1. The instant the falling leaf hits the ground, a ring of golden particles bursts outward and a gust of wind sweeps across the frame — then cut to Video 2.
Requires 2 videos (the opening and closing clips); describe only the missing middle stretch — shorter and sharper works best.