01 The Core Formula
HappyHorse 1.1 understands structured direction: who, what they do, how the camera moves, where, in what tone. This chapter breaks a good prompt into five reusable parts, then demonstrates the full skeleton of a high-complexity scene with a 15-second flagship sample. Get the skeleton right and every later chapter is just an extension of it. One premise to remember: HappyHorse outputs sound-on video by default — leave the sound unwritten and the model improvises, so every prompt should carry an audio layer.
1.1 The Basic Prompt Formula
Everything starts with one sentence: subject plus motion are the two required parts; camera, environment, and style stack on as needed. Shoot the clip in your head first, then write it down following the formula. Duration is a continuous 3–15 second dial — plan exactly as much action as the seconds you pick. A second action crammed in usually means both play badly. Audio is always on: give the ambience and effects a line each, and the clip comes back complete.
[Subject] + [Motion] + optional [Camera] [Environment] [Style]
Subject
The star of the frame — a person, animal, or product. The more specific, the better.
Motion
What the subject does, with body parts and pace spelled out; slow, continuous small motions are the safest.
Camera
Framing and camera movement: close-up, push-in, dolly — one move per shot.
Environment
Setting and mood: place, time of day, lighting, weather.
Style
Art direction and image quality: color palette, texture, clarity.
A Caravan on the Dunes
Prompt
A desert at dusk. A camel caravan moves slowly along the ridge of a dune, the packs on the camels' backs swaying with each step, the leader in a russet cloak walking at the front. The camera dollies slowly from a low side angle as the setting sun stretches the shadows of people and camels across the sand, fine grains streaming along the ground. Cinematic quality, warm golden palette, rich detail. Audio: camel bells ringing softly with each step, a low hum of wind-blown sand in the distance, the occasional heavy breath of a camel.
Two sentences for subject + motion, one each for camera and environment, then style and audio layers to close — that's a solid basic HappyHorse prompt. Suggested: 1080p, 21:9.
A 15-Second F1 Onboard Lap (Flagship Sample)
Prompt
[Style and Camera Setup] Ultra-photorealistic F1 onboard camera POV footage, mounted directly behind the driver's helmet on the roll hoop, identical to the official F1 onboard camera angle. The lower foreground of the frame captures the driver's helmet, the full nose and front wing of the car stretching ahead, with rival F1 cars clearly visible on track. Hyper-realistic physics simulation including authentic car vibration, chassis flex, suspension movement and aerodynamic buffeting. 8K cinematic quality, real broadcast color grading, motion blur on high-speed moving elements, tire smoke and rubber debris on track surface. [0-4 seconds] The car launches out of a high-speed corner at approximately 280km/h, the chassis shaking violently from kerb vibration as it exits the corner. Two rival F1 cars appear directly ahead, one slightly to the left and one slightly to the right, their rear diffusers and rear wings filling the frame. The driver's gloved hands make micro-corrections on the steering wheel, the car weaving slightly under braking instability. Engine roar is deafening, the entire frame vibrates with mechanical intensity. Tire squeal and downforce whistle fill the audio. Track surface details, paint marks and rubber racing lines are hyper-realistic. [4-9 seconds] The car approaches a tight chicane at full racing speed, heavy braking causes the nose to dip dramatically, the entire car shuddering under 5G deceleration force. The driver's helmet tilts forward under braking load. The car turns sharply left then immediately right through the chicane, the chassis rolling and yawing with authentic weight transfer, tires visibly deforming under lateral load. A rival car is overtaken aggressively on the inside of the corner, the two cars passing within centimeters of each other, slipstream turbulence causing the camera to shake violently. Tire smoke wisps from the front wheels under trail braking. [9-15 seconds] The car accelerates hard out of the final corner onto a long straight, the rear tires momentarily breaking traction causing a visible rear-end snap and oversteer, the driver catching the slide with a sharp opposite lock steering input. The car straightens and rockets down the straight at increasing speed, the two cars ahead growing smaller in the distance. Grandstand crowds on both sides blur past at extreme speed. The engine note rises through the rev range hitting the limiter, the driver upshifting via paddle shifts causing sharp torque interruptions. The camera shakes continuously with road surface imperfections and aerodynamic turbulence. The final frame shows the car at full speed with heat haze rising from the asphalt ahead. [Quality Enhancement] Authentic F1 onboard footage realism, real physics-based car movement, genuine suspension travel and chassis flex, accurate tire behavior under load, realistic engine and aerodynamic audio, cinematic motion blur on high-speed elements, broadcast-accurate color grading, hyper-realistic track surface with rubber marbles and tire marks, continuous realistic camera vibration throughout the entire clip, no artificial smoothing applied.
A validated flagship sample, verbatim: [Style and Camera Setup] sets the tone → second-level beats schedule the action (0-4 / 4-9 / 9-15) → [Quality Enhancement] closes, with audio (engine, tire squeal, airflow) written into every beat — the complete skeleton for a complex 15-second scene.
1.2 Choosing Among 9 Aspect Ratios
HappyHorse 1.1 offers nine aspect ratios — the most on Molyin: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9, and 9:21. There's only one rule for choosing — work backwards from where you'll publish: 16:9 for widescreen video, 9:16 for vertical mobile feeds, 4:5 or 1:1 for feed posts, 21:9 for a cinematic widescreen look, 9:21 for ultra-tall visual experiments. Decide the ratio before generating; cropping afterwards just throws away quality.
A Solo in the Rehearsal Studio
Prompt
A rehearsal studio bathed in morning light. A contemporary dancer in black practice clothes stretches her arms across the wooden floor, then rises onto pointe and spins once, her skirt flaring with the turn before she lands in a still, controlled pose. Static medium shot; morning light slants in through tall windows, fine dust drifting in the beams. Audio: the soft scuff of dance shoes on wood, a swish of fabric at the turn, one grounded thud on landing, and the quiet breathing room tone of the studio.
Pre-filled with 9:16 vertical — the standard frame for mobile feeds, and a full-body figure fills it perfectly; the audio layer gives the silent studio its breath.
1.3 Reproduce and Iterate with Seed
The seed is HappyHorse 1.1's random seed — and one of the abilities that sets it apart. Lock the seed and the same prompt reproduces a highly similar result: change one word and you can see exactly what that word does. Swap the seed and the same prompt becomes a fresh batch of variations. Two habits worth building: write down the seed when a result clicks, and freeze the seed while A/B-testing wording. Note that generation is probabilistic — the same seed doesn't guarantee pixel-identical output, but composition and subject stay far more stable than a random roll.
A Red Fox in the Snowy Woods
Prompt
A red fox pads through a snow-covered birch forest. It stops suddenly, ears pricked, listening — then dives headfirst into a snowdrift, sending powder flying. The camera tracks it smoothly in a medium shot; winter sunlight slants through the trunks and its breath condenses into white puffs in the cold air. Audio: the creak of paws packing snow, the muffled thump of the dive, wind through the birches, and a distant birdcall or two.
Pre-filled with seed 42 — change the wording, keep the seed, and compare exactly what each tweak does; the audio layer is A/B-testable the same way.
02 Lip-Sync and Spoken Lines
Lip-sync is HappyHorse's flagship scenario on the official playground: write the line letter-by-letter in quotes, attach the phrase "precise lip-sync", and the character delivers it to camera accurately. A short line needs one sentence covering tone and framing; longer delivery is split into second-level beats, one line per beat. Both examples in this chapter are validated official samples.
2.1 News-Anchor Delivery
Multi-line delivery is written as "camera setup + second-level beats + one line per beat": establish the scene and framing first, then split at 0-5s and 5-10s with one line and one small action per beat (turning to the second camera), and close by naming "precise lip-sync" plus the room tone.
A News Anchor's Opening Delivery (Official Sample)
Prompt
Medium shot of a professional news anchor at a sleek desk in a modern broadcast studio, cool blue lighting, softly glowing screens behind. 0-5s: He looks into the camera and says in a clear measured voice, "Good evening. Tonight, a breakthrough that could change how millions of us work." 5-10s: He turns slightly toward a second camera, "We'll have the full story, and what it means for you, right after this." Precise lip-sync, subtle studio room tone, crisp broadcast quality, shallow depth of field.
The official playground lip-sync sample, verbatim: two lines in 10 seconds split at 0-5s / 5-10s, the phrase "precise lip-sync" is mandatory, and the studio room tone gives the voice its space.
2.2 Character Syntax and Vertical Skits
The official notation for vertical skits names the character directly with character1: "character1 does something in some setting, looks up and says in English: the line" — one sentence packs the scene, the action, the language, and the line; pair it with 9:16 vertical and you have a feed-ready skit.
A Guitarist at Three in the Morning (Official Sample)
Prompt
character1 in a cozy dim room strums once, looks up and says in English: "This next one I wrote at three in the morning." Warm practical light, intimate, cinematic, precise lip-sync.
The official character-syntax sample, verbatim: character1 opens by name, "says in English" pins the language, pre-filled with 9:16 vertical — the shorter the line, the steadier the sync.
03 Image to Video
One first-frame image plus one sentence about motion, and HappyHorse 1.1 brings the picture to life — the focus of your prompt shifts from describing the scene to describing movement and sound, because the scene already exists. One hard rule to remember: image-to-video output automatically follows the first frame's aspect ratio, with no ratio selector. Want a 9:16 result? Upload a 9:16 image. Compositional control stays in your hands.
3.1 Bringing a Photo to Life
An image-to-video prompt covers only three things: what moves, how it moves, and whether the camera moves. Don't force in things that aren't in the image — instead, "wake up" what's already there: a fluttering hem, flowing water, a swaying lantern are all high-hit-rate small motions. Add one line like "preserve the original texture" and the art style won't drift. Write the audio layer as usual — a still image carries no audio information of its own.
A Canoe on a Morning Lake
Prompt
Bring this photo to life: thin mist drifts slowly across the lake, the canoe rises and falls gently on small ripples, the lantern hanging at its bow sways slightly, and a waterbird lifts off from the distant reeds and skims across the water toward the right of the frame. Keep the camera static, preserving the photo's cool-toned film look throughout. Audio: small waves lapping the hull, the faint metallic clink of the lantern, the waterbird's wingbeats skimming past, the rustle of distant reeds.
Requires 1 first-frame image. Describe only the motion, not the scene — "preserving the original look" holds the style, and the audio layer supplies what a still frame lacks; the output ratio automatically follows the first frame.
3.2 Product Showcase
For e-commerce product clips, write two things clearly: how the product moves (rotate, levitate, push in) and how the camera plays along. Add "product appearance stays consistent with the image" to lock the details, then specify lighting and backdrop — "studio lighting, clean frame" is the all-purpose closing line for product videos. Audio is the detail product clips forget most: the click of a box lid, a premium-feeling score — small things that move conversion.
A Celadon Vase in the Round
Prompt
The celadon vase in the image sits before a light-gray seamless backdrop and rotates slowly at a constant speed through a full turn, its crackled glaze shifting under soft studio light. The camera pushes in slowly from a medium shot to a close-up of the shoulder, revealing the fine gloss of the glaze. Product appearance stays consistent with the image. Studio-grade lighting, clean and premium. Audio: the quiet room tone of a gallery, backed by a minimal, premium-feeling ambient score.
Requires 1 product image. "Full rotation + push-in close-up" is the standard product-showcase combo — a premium score completes the piece.
04 Reference to Video
When words can't pin down a face — a character, a product, a mascot — hand it to reference images. HappyHorse 1.1's reference-to-video accepts images only, 1–9 of them, cited in upload order as "Image 1", "Image 2". The model locks the appearance from your images while scene and action come entirely from your text. The division of labor with image-to-video is simple: a reference image isn't the first frame — it's the character sheet.
4.1 The Character Reference Formula
A reference-to-video prompt is a three-part structure: bring the material on stage with "Image N", describe what happens in the new scene, then close with one line of "stay consistent". Upload order is the numbering order — the first image is Image 1, the second is Image 2. Multi-angle shots of the same character work best: the model builds a complete "character sheet" and the face survives any scene change.
[References: Image 1…] + [Scene description] + keep [subject] consistent
References
Images only, 1–9, cited as "Image N" in upload order.
Scene description
What happens in the new scene: action, environment, camera — the picture is generated from scratch, all driven by your words.
Consistency
One line of "stay consistent with the reference" reinforces the appearance lock.
An Astronaut's Daily Life in Orbit
Prompt
The tuxedo cat from Image 1, wearing a white spacesuit, floats slowly inside a space station module. Curious, it reaches out a paw to touch a drifting water droplet, while the blue Earth turns slowly outside the porthole. The camera orbits halfway around it. The cat's and the suit's appearance stay consistent with the reference image. 3D animation quality, rich detail. Audio: the low hum of the station's equipment, the faint wobble of the disturbed droplet, the particular quiet of zero gravity.
Requires 1 or more reference images. "Image 1" refers to your uploaded character image; the closing "stay consistent" is the standard appearance-locking phrase, and the audio layer fills in the capsule's sense of space.
4.2 Combining Multiple Images
Each image gets one job: one for the character, one for the prop, one for the outfit. Assign each role in upload order — "Image 1", "Image 2" — and the model won't mix them up. The more important the material, the earlier it should appear in the prompt: the lead always comes before the props.
The Bouquet Toss
Prompt
The bride from Image 1 stands on ivy-covered stone steps and, laughing, tosses the white peony bouquet from Image 2 over her shoulder; several bridesmaids reach for it with smiles as petals scatter in the air. Low-angle shot looking up, afternoon sunlight filtering through leaves, warm and natural. The appearances of the people and the bouquet stay consistent with the reference images. Audio: cheers and laughter from the guests, the rustle of scattering petals, birdsong faint in the distance.
Requires 2 or more reference images. Name each role explicitly — people as Image 1, the prop as Image 2 — so nothing gets crossed; the cheers are what truly brings the moment alive.
05 Video Editing
Rewrite a video with plain language — that's HappyHorse Video Edit's turf. Upload a 3–15 second source clip and state one plain instruction: change the style, change the outfit, change the prop; everything else stays untouched. Output duration and aspect ratio automatically follow the source video — there are no duration or ratio selectors. You can also attach up to 5 reference images to show the model what to change things into.
5.1 Style Transfer
Give the whole clip a new artistic skin. The official sample's phrasing is minimal: one line — "recolor the entire scene using the reference image's style" — paired with a target-style reference image. The image answers "change it into what" far more accurately than a hundred adjectives. Without a reference image, name three traits of the target style in words — brushwork, texture, palette.
Recolor by Reference Image (Official Sample)
Prompt
Recolor the scene using the reference image's style. Keep the camera movement, composition, and pacing of the original video unchanged — only the color grade and artistic texture change to match the reference.
Requires 1 source video of 3–15 seconds + 1 target-style reference image; this example switches to Video Edit automatically, and duration and ratio follow the source. The core phrasing of the official recolor sample — style goes to the reference image, the storytelling stays untouched.
5.2 Element Replacement
Swap clothes, props, or products without touching the rest of the footage: state clearly what to replace with what, then close with "everything else stays unchanged". For replacements, one reference image nails the new look — "the beige cardigan in Image 1" beats a hundred adjectives.
Swap the Jacket for a Cardigan
Prompt
Replace the person's jacket in the video with the beige knit cardigan from Image 1. The fabric should crease naturally with the movement, and the texture and lighting direction should match the original footage. Everything else stays unchanged.
Requires a source video, plus up to 5 reference images. "Replace A with B + everything else stays unchanged" is the most reliable replacement phrasing.