Guide

MiniMax H3 Guide

MiniMax H3 generates picture and stereo sound in a single pass — 4 to 15 seconds at 768p or native 2K, with multi-shot editing, lip-synced dialogue and a unified reference context that reads up to 9 images, 3 video clips and 3 audio tracks together. This guide is organized around real creative tasks: from the three-track formula and shot timelines, to audio direction, keyframes and reference casting — every example preloads into the Molyin generator with one click.

Updated 2026-08-09

01 The Three-Track Formula

Most video models generate a picture; H3 generates a film — image, ambient sound and score come out of the same pass. That changes how you write: a good H3 prompt covers what we see, what the scene sounds like, and (optionally) what music plays under it. This chapter builds that skeleton, then covers the two dials every job needs: resolution and duration.

1.1 Write for Eyes and Ears at Once

Start from one sentence of subject plus action, then layer on camera, environment and — the H3-specific habit — sound. Because audio is generated jointly with the picture, anything you don't specify gets invented for you. One line of ambience and one line of music are usually enough to take control. Plain natural language works; H3 was also trained on a stricter structured layout we'll meet in chapter 3.

[Subject] + [Action] + optional [Style] [Camera] [Ambient sound] [Score]

Required

Subject

Who or what the shot is about — the more concrete the description, the more stable the result.

Required

Action

What the subject does, one clear beat at a time. Vague verbs produce vague motion.

Optional

Style

The overall look, stated up front: live-action cinematic, 2D animation, claymation, vintage film.

Optional

Camera

Shot size and movement. H3 follows explicit camera language well — chapter 2 gives the full vocabulary.

Optional

Ambient sound

What the scene itself sounds like: rain, traffic, sizzling oil. Characters can hear this.

Optional

Score

Music only the audience hears. Name instruments and tempo, not moods.

The Night Noodle Stall

Prompt

Live-action, cinematic. A late-night street noodle stall glows under a single hanging bulb; the cook lifts a strainer of noodles from the rolling boil, steam flooding the frame, and tosses them into a waiting bowl. The camera holds a medium shot, then pushes in slowly toward the bowl. Boiling water rumbles, the strainer clangs on the pot rim, distant scooters pass. A lone erhu line plays softly under the scene.

One subject, one clean action, one camera move — and both sound layers named. This is the baseline H3 prompt shape.

Lighthouse in the Storm

Prompt

Live-action, cinematic, cool blue-grey palette. A lighthouse keeper in a yellow raincoat leans into the wind on the gallery deck, gripping the rail as spray bursts over him; the beam sweeps past twice. The camera trucks right with large amplitude at slow speed, keeping him framed against the churning sea. Wind roars, waves slam the rocks below, the raincoat snaps and flutters. Low sustained strings build slowly and hold without resolving.

All six elements in play — style first, then subject, action, an explicit camera move, and one line each for ambience and score.

1.2 Resolution & Duration: 768p to Explore, 2K to Deliver

Two resolutions — 768p and native 2K — and a continuous 4–15 second dial that accepts any whole number. The working rhythm: iterate at 768p, where each second costs a fraction of 2K, then rerun the winning prompt at 2K for the final cut. Match action load to seconds: one beat per 4–5 seconds is reliable; 15 seconds carries a three-beat miniature scene. H3 renders at 24fps with stereo audio at every tier.

The Koi Pond in 2K

Prompt

Live-action, documentary macro style. A koi pond after rain: orange and white koi glide beneath a surface stippled with droplets, lily pads trembling as fish pass under them, a maple leaf spinning slowly in the current. The camera starts on a tilt down at slow speed, then holds a static shot as three koi converge on a drifting crumb of bread. Soft rain patter fading out, water lapping stone, a single splash as a koi breaks the surface. No music.

Preloaded at 15 seconds in 2K — texture-dense subjects like water, scales and foliage are where native 2K visibly pays off. Note "No music" is itself a direction: it keeps the score track empty.

02 Multi-Shot Timelines

H3 models cuts natively — you can direct an edited sequence, not just a single take. The convention: open with [Shot 1], then start every later shot with a tag and a timestamp, like [Shot 2] At 00:04.000. Timestamps must strictly increase and stay inside the video duration. This chapter covers the shot list and the camera vocabulary that moves within each shot.

2.1 The Shot List: Tags + Timestamps

Write each shot as its own sentence block: the tag, the cut time, then what the new shot shows. Use a cut to introduce genuinely new information — a new subject, space, or viewpoint; if only the distance changes, move the camera instead of cutting. Budget seconds honestly: a 12-second video holds three shots comfortably, five is pushing it.

[Shot 1] opening shot … [Shot N] At MM:SS.mmm, the camera cuts to …

Required

Shot tag

[Shot 1] opens the video with no timestamp; every later shot gets its own numbered tag.

Required

Cut time

MM:SS.mmm at the start of each later shot — strictly increasing, always inside the total duration.

Optional

Cut phrase

"the camera cuts to" is the workhorse; "the shot transitions to" and friends also work. Dissolves and fades only when you actually want them.

Optional

New information

What this cut reveals that the last shot couldn't — new subject, new space, new viewpoint. A cut without new information should be a camera move.

Morning Market in Three Shots

Prompt

Live-action, cinematic, warm morning light. [Shot 1] A wide shot of a covered market waking up: vendors rolling up shutters, crates of oranges stacked into pyramids, sunlight cutting through gaps in the roof. [Shot 2] At 00:04.000, the camera cuts to a close-up of a fishmonger's hands laying glistening mackerel on crushed ice, water droplets scattering. [Shot 3] At 00:08.000, the camera cuts to an elderly customer lifting an orange to her nose, closing her eyes, and smiling. Shutters rattle, ice crunches, the crowd murmur swells gradually. A light acoustic guitar pattern at a walking tempo.

Three shots in 12 seconds — wide establishes, close-up details, face reacts. The soundscape and score lines sit after the shot list and cover the whole video.

2.2 Camera Vocabulary: Type × Amplitude × Speed

H3 understands a precise camera vocabulary — pan, truck, tilt, pedestal, push in, pull out, zoom, arc, tracking, static, POV, handheld shake, roll — and each move takes two optional modifiers: amplitude (small or large) and speed (slow or fast). Write the move as a natural action inside the shot, not as a tag pile at the end. Medium amplitude and normal speed are the defaults; only say them when they matter.

Around the Sculptor

Prompt

Live-action, cinematic, a single continuous shot. In a dusty studio flooded with window light, a sculptor chisels the face of a marble figure, pausing to blow powder off the cheekbone. The camera arcs around him with large amplitude at slow speed, revealing the figure's face only as the orbit completes. Chisel taps ring against stone, dust whispers off the surface, a pigeon coos on the skylight. No music.

One shot, one orbit — the arc's amplitude and speed are spelled out, and the reveal is tied to the camera path. Motion this deliberate is exactly what the three-dimension vocabulary is for.

03 Directing Native Stereo

H3 outputs stereo audio on every generation — there is no audio switch, so the real question is whether you direct the soundtrack or leave it to chance. Speech is lip-synced to the speaker and generated with the picture. This chapter covers the three audio moves: on-screen dialogue, off-screen voiceover, and the two-layer split between scene sound and score.

3.1 Dialogue: Who Speaks, What They Say, How

The reliable recipe has three parts: identify the speaker (age, voice character), give the exact line, and direct the delivery. H3's native notation tags each speaker with a stable ID like (S1) and wraps the line as <d>[English] …</d> — the tag names the spoken language, and the line inside is preserved verbatim. Quotes work too, but the tagged form is the most deterministic. Keep lines short: one or two sentences per 10 seconds. Any text that should appear on screen — signs, labels — goes in double quotes instead.

The Bookshop Recommendation

Prompt

Live-action, cinematic, warm tungsten light. In a cramped secondhand bookshop, the owner — a woman in her sixties with a low, unhurried voice (S1) — pulls a clothbound novel from a high shelf, taps its cover twice, and hands it across the counter, saying with quiet certainty: <d>[English] Everyone asks for the famous one. This is the one they keep.</d> The camera holds a medium close-up over the customer's shoulder. Pages rustle, a floorboard creaks, rain ticks on the shop window. No music.

Speaker established with age and voice character, tagged (S1), line wrapped in <d>[English] …</d>, delivery directed with "quiet certainty" — the full dialogue recipe in one prompt.

3.2 Voiceover: Narration Over the Picture

For narration, H3 has a fixed phrase worth using verbatim: the speaker "says in an off-screen voiceover", followed by the line — and right after it, state that the on-screen character's lips stay closed. That last clause is not decoration; it is what stops the model from lip-syncing your narration onto whoever is in frame.

The Letter from the Mountains

Prompt

Live-action, cinematic, muted colors. A young man hikes a switchback trail above a fog-filled valley, breath visible in the cold air. An older man's weathered voice (S1) says in an off-screen voiceover: <d>[English] Your grandfather walked this path every autumn. Now you know why.</d> while the hiker's lips remain completely closed. The camera tracks him from behind with small amplitude at slow speed, then tilts up to the ridgeline. Wind moves through dry grass, boots crunch on gravel, a distant hawk cries. Sparse piano notes at a slow tempo, fading before the video ends.

The voiceover phrase plus the "lips remain completely closed" clause is the pair that separates narration from on-screen speech — drop the second half and the hiker may start mouthing the words.

3.3 Two Layers: Scene Sound vs. Score

H3 thinks of audio in two layers. Scene sound is everything the characters could hear — ambience, action sounds, breaths. Score is music only the audience hears — describe it by instrumentation, tempo and dynamics, never by mood words. In H3's native structured format these are literal labeled fields: the timeline goes under integrated_multimodal_description, scene sound under overall_soundscape, and score under non_diegetic_music (write N/A for silence). Molyin passes your prompt through untouched, so you can paste the full three-field layout into the prompt box — the example below is one, ready to submit.

integrated_multimodal_description: [Shot 1] … ⏎ overall_soundscape: [ambience + action sounds] ⏎ non_diegetic_music: [instrumentation + tempo + dynamics, or N/A]

Required

Ambience

The scene's continuous bed: rain, room tone, traffic. 1–4 sentences covering the whole video.

Optional

Action sounds

Physical events that make noise — footsteps, impacts, fabric. Dialogue stays in the timeline field, not here.

Optional

Score spec

Instruments, tempo, dynamic arc — "sparse piano, slow, swelling strings". Not "emotional music". N/A means silence.

Rooftop Rain, Full Native Format

Prompt

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium shot frames a woman standing under a corrugated awning on an apartment rooftop as evening rain falls, holding a chipped mug of tea. She watches the city lights blur below, then closes her eyes and breathes out slowly. [Shot 2] At 00:05.000, the camera cuts to a close-up of the mug: steam curls up through the rain-light, a drop from the awning strikes the rim and splashes. overall_soundscape: Steady rain drums on the corrugated awning and hisses on the streets below. The mug clinks once against her ring, and a long exhale is audible under the rain. non_diegetic_music: A muted felt-piano motif at a slow tempo, repeating with a slight swell in the second half and no resolution.

This is the complete native three-field layout, pasted as-is into the prompt box. Use it when you want frame-accurate control; for quick jobs, the flowing style from chapter 1 is fine.

04 Image to Video

Give H3 a first frame and it develops the scene forward from exactly that image — identity, clothing, layout and style all anchor to what you upload. The output aspect ratio follows the source image, so compose your still in the frame you want to publish. Add a last frame and H3 interpolates the journey between the two.

4.1 First Frame: Anchor, Then Develop

The proven structure: anchor first, then move. Open by confirming what the image establishes — subject, position, mood — then describe the next action, the camera's behavior, and the sound. Don't re-describe every pixel and don't contradict the image; the frame has already won that argument. Small, physical motions have the highest hit rate.

The Old Bicycle Wakes Up

Prompt

Starting from this frame, the scene comes alive: the bicycle's front wheel begins to turn slowly, the playing card clipped in its spokes ticking faster and faster, and the handlebar streamers lift in a rising breeze. The camera pushes in with small amplitude at slow speed toward the spinning spokes. The card's tick builds into a soft clatter, wind moves through the alley, a dog barks once far away. No music.

Upload one first-frame image. The prompt spends no words describing the bicycle — the image owns the look; the words own the motion and the sound.

4.2 First + Last Frame: Describe the Journey

With both frames uploaded, your prompt's job flips: skip the endpoints — they're already fixed — and write the path between them. How does the subject move, what changes hands, how does the composition evolve, and how do the differences between the two frames progressively close? Keep it to a single continuous shot so the model can interpolate cleanly, and give bigger visual gaps more seconds.

The Bowl Takes Shape

Prompt

On the spinning wheel, the lump of wet clay rises under the potter's cupped hands: her thumbs press a hollow into the center, the walls climb and flare outward, and slip runs down her wrists as the form settles into the finished bowl's silhouette. The camera holds a static close shot on the wheel head the whole way. The wheel hums steadily, wet clay squelches under her palms, water drips from her fingers into the tray.

Upload the lump as the first frame and the finished bowl as the last — the prompt narrates only the transformation. Eight seconds gives the interpolation room; a bigger gap between frames deserves more.

05 Reference Casting

H3's signature capability: a unified context that reads up to 9 images, 3 video clips (each 2–15s, 15 seconds combined) and 3 audio tracks in one request — and a prompt that tells it what each asset contributes. Name assets by kind and upload order — Image 1, Video 1, Audio 1 — and assign each a role. This is casting, not decoration: appearance from one asset, motion from another, voice from a third.

5.1 Name Every Asset, Assign Every Role

The rule that makes multi-reference work: every uploaded asset gets named and given a job. "Use the camera movement from Video 1, put the character from Image 1 on screen" — kind plus upload order, role stated outright. Unassigned references invite the model to guess. Images carry identity, scenes and style; videos carry camera movement, rhythm and structure; audio carries voice timbre and musical feel. The first 5 reference images are free; from the 6th, a small per-image surcharge applies — the estimator shows it before you submit.

[Image/Video/Audio N] provides [role] + target description + optional [consistency lock]

Required

Named assets

Image 1, Video 1, Audio 1 — kind plus upload order. Numbering restarts within each kind.

Required

Roles

What each asset contributes: appearance, scene, camera movement, pacing, voice timbre, music style.

Optional

Consistency lock

The traits that must survive: hair, wardrobe, palette, lighting. Spell them out in words even though the image shows them.

The Trench-Coat Walk

Prompt

The woman from Image 1 walks toward the camera through the rainy neon street, following the camera movement and pacing of Video 1 — a slow backward tracking shot that keeps her centered as she advances. Keep her identity consistent: short silver hair, dark green trench coat with the collar up, calm unhurried stride. Puddle splashes under her boots, rain hisses on neon signs, a synth bassline pulses at a slow tempo.

Upload one character image and one camera-reference clip. Each asset is named, each has exactly one job, and the consistency lock restates in words what must not drift.

5.2 Consistent Characters, Borrowed Voices

For a recurring character, reference does the heavy lifting and language does the locking: name the identity traits explicitly next to the image reference, and they hold across shots. Audio references extend this to the voice — point a speaker at an uploaded track and H3 matches its timbre and delivery without copying the recording. That's how you give a generated character a consistent voice across a series. Audio references are free of surcharge.

The Street Singer

Prompt

The busker from Image 1 — young man, loose curly hair, denim jacket over a white tee, worn acoustic guitar — performs under a stone archway at dusk. He strums twice, closes his eyes, and sings (S1) with the voice timbre and relaxed delivery of Audio 1: <d>[English] Another day is folding, and I'm still on this side of town.</d> The camera pushes in slowly from a wide shot to a medium close-up as passersby slow down to listen. Guitar strings ring under the archway's natural reverb, coins clink into the open case, evening traffic murmurs beyond.

Upload one character image and one vocal track. The (S1) speaker maps to the audio's timbre — the recording steers how the voice sounds, the <d> line decides what it sings. Preloaded at 12 seconds in 2K for the performance detail.

MiniMax H3 Parameter Cheat Sheet

Every tier the Molyin generator actually exposes for MiniMax H3.

ModesText to video / Image to video / Reference to video
Resolution768p / 2k
Duration4–15s
Aspect ratio16:9 / 9:16 / 4:3 / 3:4 / 21:9 / 1:1
AudioStereo audio on every generation, no switch
First-last frameYes
Reference images1–9
Reference videosUp to 3
Reference audioUp to 3
Return last frameNo

FAQ

The eight most frequent questions, with answers we verified ourselves.