01 The Core Formula
Veo 3.1 understands structured direction: who, what they do, how the camera moves, where, and what it sounds like. This chapter breaks a good prompt into five reusable parts — and note the fifth one is sound, because every Veo clip carries audio. Writing sound isn't optional homework; it's part of the craft. Two official ground rules on top: prompts should be clear and specific with ambiguity eliminated; and an 8-second short focuses on a single moment — don't chain "A then B then C" into one prompt.
1.1 The Basic Prompt Formula
Everything starts with one sentence: subject plus motion are the two required parts; camera, environment, and audio stack on as needed. Veo 3.1's duration is fixed at 8 seconds — treat it as a narrative unit, not a limitation: 8 seconds is exactly one complete beat (one subject, one core action, one camera move). Before writing, ask yourself: if only one thing happens in these 8 seconds, what is it? A second action crammed in usually means both play badly.
[Subject] + [Motion] + optional [Camera] [Environment] [Audio]
Subject
The star of the frame — a person, animal, or product. The more specific, the better.
Motion
What the subject does, with body parts and pace spelled out; one core action per 8 seconds.
Camera
Framing and camera movement: close-up, push-in, dolly — one move per shot.
Environment
Setting and mood: place, time of day, lighting, weather.
Audio
Sound effects, dialogue, music — every Veo clip carries audio; leave it blank and the model improvises.
Sunrise at the Fishing Harbor
Prompt
A fishing harbor at dawn. A few wooden boats rise and fall on the swell while an old fisherman in a conical hat stands at the bow, bending to haul in a dripping net, coil by coil. The camera pushes in slowly from a low angle on the pier as golden morning light spreads across the water. Audio: waves lapping against hulls, gulls crying as they skim past, the patter of water shaken from the net, the low rumble of a diesel engine in the distance.
Two sentences for subject + motion, one each for camera and environment, then name the sounds to close — that's a solid basic Veo prompt. Suggested: 1080p, 16:9.
1.2 Camera Language
Veo 3.1 executes the terms in Google's official taxonomy reliably, and three families can be combined. Camera angles: eye-level, low-angle, high-angle, bird's-eye, over-the-shoulder, POV. Camera movements: static, pan, tilt, dolly, truck, crane, aerial, arc, whip pan. Optical effects: shallow depth of field, rack focus, telephoto compression. One rule of discipline — a single movement per 8-second shot; stacking push, orbit, and pan together just buys you drifting footage.
A Dolly Through the Morning Market
Prompt
An open-air market in the early morning. The camera trucks slowly along a row of stalls: bamboo baskets piled with dew-covered greens, a vendor lifting the lid off a bamboo steamer as white vapor billows out, an elderly woman picking through tomatoes and dropping her choices into a cloth bag. Morning light slants through gaps in the canvas awnings, dust motes drifting in the beams. Audio: vendors calling out over one another, the rustle of plastic bags, the hiss of the steamer lid lifting — lively and full of everyday life.
"The camera trucks slowly along…" is the official term and the steadiest phrasing for a lateral move; queue the sights in the order the camera passes them. One move — the truck — for the whole 8 seconds.
1.3 Style Control
Style words set the art direction of the whole clip — position is flexible, but they can't be missing: "cinematic", "3D animation", "watercolor illustration", "stop-motion", and "documentary handheld" are all high-hit terms from the official lists. Keep the style self-consistent: an origami world should sound like rustling paper. When picture and sound are written in the same style, the result feels like one complete world.
A Deer in an Origami Forest
Prompt
A world made of origami: layer upon layer of folded-paper trees. A paper deer steps out from behind a tree, lowers its head to sniff an origami flower, then looks up, its paper ears giving a gentle twitch. The camera slowly pushes in to a side view of the deer. Stop-motion feel, soft studio light, visible paper texture. Audio: the delicate rustle of paper, a light and playful xylophone score — like opening a pop-up storybook.
The style words "stop-motion feel + paper texture" set the tone, and the audio matches with paper rustle and xylophone — one world, one style. Suggested: 1080p.
02 Native Audio
Always-on audio is Veo 3.1's headline feature: every clip comes with sound, there is no off switch, and no post-production dubbing is needed — effects, dialogue, and music are generated with the picture, automatically in sync. Google recommends describing audio in separate sentences. The more specific you are, the more precise the result; write nothing and the model improvises — not always the mood you wanted. This chapter splits sound into effects, dialogue, and music, and covers each in turn.
2.1 Sound Design
Sound effects are written as "name it + layer it": list each sound you want, then state which is foreground and which is distant — the near main sound first, details in the middle, the far backdrop last. Pair onomatopoeia with its source for best results: "thunder booms" beats "there is thunder", and "rain drumming on a tin roof" beats "rain sounds".
A Storm Breaks Over the Forest
Prompt
A coniferous forest in a storm. Gale-force wind bends the treetops to one side, lightning tears across the dark sky, fat raindrops hammer the fern leaves, and a mountain stream rushes between boulders. Static low-angle shot among the trees, the frame flaring bright with each lightning strike. Audio: dense rain drumming on leaves up close, wind howling through the forest, the stream rushing, then a thunderclap booming directly overhead and rolling away into the distance.
Layer the audio — near rain, wind and water, thunder moving from near to far — with the thunder given a full arc. 8 seconds is exactly one thunderclap's complete journey.
2.2 Character Dialogue
The official dialogue notation is "who + how + colon + the line": identify the speaker by appearance, attach a delivery cue like "with a smile" or "softly", and let the line follow a colon. One official best practice deserves special attention: do not wrap the line in quotation marks — quotes risk being rendered as on-screen text in the video, so let a colon introduce the words. The model lip-syncs automatically; 8 seconds holds one or two short lines, and the shorter the line, the steadier the sync. For longer conversations, splitting into two clips works better.
The Head Chef's Taste Test
Prompt
In a restaurant kitchen, a middle-aged head chef in a dark apron lifts a spoonful of soup, blows on it, tastes it, and closes his eyes for a second. He turns to the young commis beside him and says with a smile: One thing is missing — wasabi, just a touch. The commis freezes for a beat, then nods and dashes to the spice rack. Static medium shot, the stove's firelight playing on both faces. Audio: the soup pot bubbling, the low hum of the extractor hood, the chef's steady voice, the bright clatter of spatulas in the background.
The line follows a colon with no quotation marks (official rule: quotes get rendered as on-screen text), the speaker is identified as "the head chef in a dark apron", the delivery cue "with a smile" is attached — then the kitchen sounds are layered last. Suggested: 16:9.
2.3 Music and Mood
Music is written as a trio of "genre + mood + instruments": "upbeat jazz" is vague, while "a laid-back jazz trio, brushes circling on the snare" is precise. Match the score to the picture's rhythm — quick-cut montages want driving tempo, long takes want an ambient bed. In visual-only passages with no dialogue or effects, music is the sole voice in the mix and deserves the extra sentence.
A Waterfront City at Dusk
Prompt
A waterfront city at dusk: a ferry carves through golden water as it docks, office towers catch the setting sun in their glass, pedestrians on a footbridge trail long shadows, and a corner record store switches on its warm lamps. The camera drifts between these four vignettes, the mood warm and nostalgic. Music: a laid-back jazz trio — plucked double bass as the foundation, brushes circling softly on the snare, a saxophone drifting in with an occasional phrase. Relaxed, wistful.
The music is spelled out as "jazz trio + laid-back nostalgia + bass/brushes/saxophone" — the full trio, matched to the pace of a dusk montage. Suggested: 1080p.
03 Image to Video
One first-frame image plus one motion description — Veo 3.1 brings the frame to life. Two official best practices: prompt for motion only (the image already provides the subject, scene, and style — re-describing them only confuses the model), and refer to people in the image with general terms ("the woman", "he", "she"). One practical rule: crop your image to the target ratio (16:9 or 9:16) before uploading, so the framing stays in your hands. Write sound as usual: a still image carries no audio information, so effects and music come entirely from your prompt.
3.1 Bring the Frame to Life
An image-to-video prompt writes only three things: what moves, how it moves, and whether the camera moves — the official taxonomy calls these camera motion, subject animation, and environmental animation, alone or in combination. Don't force in things that aren't in the image — but anything already there can be awakened: a fluttering coat hem, drifting clouds, a flickering flame. Small motions hit most reliably. Add one line — "carry over the image's original texture" — and the style won't drift.
A Hand-Drawn Sketch Comes Alive
Prompt
Bring this hand-drawn sketch to life: the girl standing on the bridge, her hair and scarf lifting gently in the wind. She raises a hand to tuck a wind-blown strand behind her ear, the inked lines of the river below the bridge ripple with flowing light, and the penciled clouds drift slowly overhead. The camera stays fixed; the motion is subtle and natural, preserving the original pencil-sketch texture and paper grain throughout. Audio: a soft breeze, a faint paper-like rustle, a quiet solo piano piece.
Requires 1 first-frame image; write only the motion, never re-describe the scene (official rule), and let "preserving the original pencil-sketch texture" hold the style.
3.2 Product Showcase
For e-commerce product clips, the prompt covers two things: how the product moves (rotate, levitate, tilt) and how the camera plays along (orbit, push in). Add one line — "keep the product identical to the image" — to lock the details, then name the lighting and backdrop. Audio is the detail most product clips forget: the click of a box lid, a premium-feeling score — small things that move conversion.
A Levitating Running Shoe
Prompt
The running shoe from the image rises slowly against a dark gray backdrop, levitating as it rotates one full, even turn — the woven texture of the upper and its reflective strips catching the light. The camera pushes in slowly to a close-up of the midsole, revealing the cushioning structure. Keep the product identical to the image. Studio lighting, clean premium look. Audio: a soft whoosh as the shoe lifts, backed by a minimal, tech-flavored electronic score with a steady beat.
Requires 1 product image; "one full rotation + push-in close-up" is the classic product-showcase combo, with a tech-styled score to match.
04 Reference-to-Video and Fast
Veo 3.1 Fast is the budget-friendly volume tier: same picture specs as the standard model (720p/1080p/4K, 8 seconds, always-on audio) at a friendlier price — ideal for drafts and batch production. It also exclusively owns one capability: reference-to-video (r2v). Upload 1–3 reference images to lock a character's appearance and send that character into any new scene. Hard rule: reference-to-video output is locked to 16:9 landscape; for vertical video, go back to text-to-video or image-to-video.
4.1 Two Paths to Character Consistency
Path one (reference images): the prompt is a three-part structure — bring the assets on stage with "Image N", describe what happens in the new scene, then close with a consistency line. Upload order is the numbering order, up to 3 images, and multiple angles of the same character work best. Path two (pure text): the official best-practices approach — give the character a name, write one detailed, word-for-word-fixed description of appearance and voice (age, hair, features, wardrobe, accent), reuse the whole paragraph verbatim in every new scene's prompt, and pair it with the same seed for cross-scene consistency of both face and voice. The clara-archive example below is the official documentation's own demonstration.
[References: Image 1…Image 3] + [Scene description] + keep [subject] consistent
Reference Images
Up to 3 images, cited in the prompt as "Image N"; upload order is the numbering order.
Scene Description
What happens in the new scene: action, environment, camera, audio.
Consistency
A closing line like "keep it identical to the reference images" reinforces the lock.
A Robot Mascot Dances in the Square
Prompt
The white, round-headed little robot mascot from Image 1 and Image 2 appears in a city square at night, a giant ground screen flowing with colorful light. It breaks into street dance to the music: side-to-side glides, robotic freezes, and a final one-arm-raised finishing pose. The camera orbits half a circle around it at medium distance. Keep the mascot identical to the reference images; 3D animation quality. Audio: an energetic funk track, faint mechanical whirs from the robot's joints, and one crisp hit of sound on the final freeze.
Requires 1–3 reference images; this example automatically switches to Veo 3.1 Fast with 16:9 landscape output. Multiple angles of the same character give the best consistency.
Clara in the Archive (Official Consistency Example)
Prompt
A medium shot, with the camera slowly dollying forward in a dimly lit, grand Parisian archive. Dust motes dance in a single beam of light from a high window. Clara, a historian in her early 30s, with observant, dark brown eyes that hold a quiet intensity. She has chin-length, black hair styled in a classic bob. She is dressed in a sophisticated, dark navy-blue wool coat, with a silk scarf patterned with subtle gold and cream designs tied around her neck. She stands before a large, ancient wooden table, carefully turning the fragile, yellowed page of a massive, leather-bound book. Her expression is one of deep concentration. In a voice that is crisp and clear, with a thoughtful, analytical tone and a standard American accent, Clara says: It has to be here.
The verbatim example from Google's official best-practices doc: appearance and voice are written as one word-for-word-fixed paragraph, reused wholesale in every new scene (only the action and setting change), paired with the same seed — the standard pure-text approach to cross-scene consistency. Note the colon introduces the line, no quotation marks.
4.2 When to Use Fast
In one line: draft with Fast, finish with Standard. Prompt iteration, storyboard animatics, and batch social content multiply Fast's cost advantage; for the final hero shots where every detail matters, switch back to the standard model for maximum fidelity. And whenever you use reference-to-video, you're already on Fast — the one case where the choice is made for you.
Bringing an Oil Painting to Life
Prompt
Image 1 is an oil painting: dusk in a seaside town, fishing boats resting on the sand, a church clock tower lit in the distance. Bring the painted world to life: cooking smoke curls up from the chimneys, the sea shimmers in thick impasto strokes, a fisherman crosses the beach with an oar on his shoulder, and the clock tower's lamplight flickers gently. The camera pans extremely slowly, preserving the heavy brushwork and paint texture throughout. Keep the style identical to the reference image. Audio: distant waves, the church bell tolling three evening chimes, gulls calling — quiet and far-reaching.
Requires 1–3 reference images; this example automatically switches to Veo 3.1 Fast. For style references, double-lock with "preserve the brushwork texture" plus "keep the style identical".