MiniMax H3 (Hailuo 3)
An omni-modal video model that reads text, images, video and audio in one context — and answers with a native 2K film, stereo sound included, up to 15 seconds in a single take. Live on Molyin — generate below.
- Failed generations auto-refunded
- Credit packs never expire
- Every model's price is public
- Cancel anytime in one click
- 2K
- Native resolution
- 24 fps
- Frame rate
- 4–15 s
- Any clip length
- Stereo
- Sound in every take
Official MiniMax H3 samples
Straight from the MiniMax release — hover to preview, tap the speaker icon to hear the native stereo mix.
- MiniMax H3
- Reference to video
- 15s
- 2K
Image 1 serves as a reference for the overall texture and mood, while Image 2 provides reference for the character's appearance. Create a 15-second, 16:9 widescreen trendy fashion short film. Maintain consistency in character design: platinum blonde long hair, narrow black vintage sunglasses, black glossy patent leather trench coat, a cold and confident fashion expression, and orange reflections of flames cast on the surface of the leather coat. The overall style is a fast-cut film-like fashion advertisement, blending elements such as nighttime fire scenes, black smoke, orange-red flames, VHS glitches, CCTV broadcast interruptions, 90s analog film grain, scan-lines, chromatic aberration, light leaks, flash-white transitions, and slight camera shake throughout.
- MiniMax H3
- Reference to video
- 15s
- 2K
Create a 16:9 widescreen, high-end fashion brand video. The overall mood, scenes, and film-like texture should reference Image 1; bag assets should reference Image 3; character assets should reference Image 2; and the brand ending logo should reference Image 4. This is a fashion campaign for selling clothes and bags. The overall vibe should be high-end, cool, and restrained, but the editing cannot feel dull—it needs to be more dynamic and infused with a strong sense of fashion rhythm. Avoid the feel of an ordinary narrative film or an e-commerce ad. The core story remains simple: on a desert highway beside a vintage car, a woman walks to the car's trunk, opens it, and retrieves a black bag. There's a brief, quiet connection between her and a man standing beside the car. Then, she walks away with the bag. The bag and the clothing should naturally integrate into the characters' movements, becoming an intrinsic part of their personas.
- MiniMax H3
- Text to video
- 15s
- 2K
A faster pace, grand yet not drawn out. Quick hard cuts, bridge shaking, intense light flashes, short blackouts, and warp jumps. The text design resembles the wide-spaced lettering of a movie trailer, avoiding plain white but featuring a restrained glow and textured finish with subtle edge highlights. The text animation includes emerging from the deep void of space, illuminated by starlight sweeps, expanding letter spacing, afterimages, subtle glowing, and flashes of blackouts.
- MiniMax H3
- Reference to video
- 15s
- 2K
Style: A dark-pop / cyber-grunge / rap music video with a high-fashion yet realistic texture, evoking a film magazine aesthetic. The visuals feature high contrast but avoid appearing cheap. The overall reference draws from late-90s to early-2000s independent magazines, photocopied zines, film scans, underground music posters, and collage aesthetics. Visuals include coarse grain, subtle film jitter, halftone patterns, print edge roughness, and scanning misalignments. The editing rhythm is fast, relying solely on hard cuts with no fades or soft transitions. The text design style and texture should reference the attached imagery.
- MiniMax H3
- Text to video
- 15s
- 2K
A 15-second, 16:9 horizontal short video. A live-action nighttime laundromat scene blended with hand-drawn glowing animations. The setting is a small self-service laundromat with faintly flickering fluorescent lights. Inside are running washing machines, plastic laundry baskets, an old bench, and a single sock discarded on the floor. The overall space feels quiet, carrying a subtle sense of nostalgia. Filmed with a handheld phone in single-hand grip, the footage has noticeable shakiness; the white fluorescent lighting causes fluctuating exposure; surface reflections are visible on the glass panels; and the focus lags slightly when the camera gets close to objects. The visuals should not look polished or meticulous like a commercial ad. The overall texture should resemble a candid, spontaneous recording—like stumbling into the laundromat late at night and instinctively capturing surreal visions with a raw documentary feel.
- MiniMax H3
- Text to video
- 15s
- 2K
Generate a 15-second, 16:9 widescreen light suspense crime film opening sequence. The overall style should reference the visual language of these concepts: retro Japanese anime openings, hard-edged silhouettes, comic-style collage, asymmetrical split screens, strong geometric color blocks, English title credits, minimal Japanese katakana for decoration, and a jazz-crime vibe. The atmosphere should be 60% suspense and 40% jazz: mysterious, cool, agile, with an urban crime feel—neither scary nor heavy, and definitely not a joyful jazz music video. The motion effects of the opening should feel like animated graphic collages: a black-lined frame appears first, split-screen borders are quickly drawn out, color blocks and frames are pasted in piece by piece; silhouettes of characters, close-ups of props, and English title credits slide in, bounce out, or are revealed through masking along with the drum beats. Avoid realistic narrative animations—this should have the polished feel of an opening sequence. The English credits should be legible, with animated effects allowed: thin line frames can first be drawn out, names slide into the frame, with letters appearing one at a time and revealed by moving color blocks, before pausing briefly at the end. Do not add Chinese text, avoid garbled text, and ensure no misspelled English names. Rules for the sequence: Every English credit and title should only appear once. Do not repeat the same title, do not repeat the same name, and do not assign multiple roles to the same individual. Transitions must be varied: circular vinyl record masks, vertical cuts through car doors, character shadows acting as wipes, red-line cuts, large English letter masking, split-frame border reorganizations, hard geometric color-block cuts, and individual frame-by-frame collage reveals. All transitions should sync with the drum beats—precise, suspenseful, agile, and with a comic-collage vibe. Avoid soft dissolves and fluid transitions. BGM: Create an original BGM.
One context in. One film out.
This is MiniMax's own flagship example. Three assets — a camera-movement reference, a character image and a vocal recording — plus one sentence describing how they fit together. MiniMax H3 reads them as a single context and delivers the finished, singing shot.

The prompt
Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3.
What is MiniMax H3?
MiniMax H3 is the third generation of MiniMax's H series, released on July 31, 2026 — and it is not just a video model. MiniMax H3 is a general-purpose multimodal generation model: it understands text, images, video clips and audio inside one unified context, and you describe how they relate in plain language. Ask it to borrow the camera movement from one video, put the character from an image on screen, and have them sing the vocals from an audio file — in a single prompt.
The output side is just as ambitious. MiniMax H3 renders in native 2K by default — not an upscaled image, but in-context regeneration by the base model, which is why fine details like small text survive. Every generation ships with native stereo sound: music, dialogue and effects are produced jointly with the picture, not bolted on afterwards. Add native multi-shot modeling, accurate text and brand rendering, and motion transfer from reference footage, and MiniMax H3 reads like a model built for commercial work — ads, e-commerce, product and UI films. MiniMax has also announced plans to open the model weights.
While MiniMax H3 rolls out, every capability it competes with is already live on Molyin: try our AI video generator, check pricing, or browse the full model library.
MiniMax H3 vs Hailuo 02
MiniMax H3 replaces the Hailuo 02 generation with a new architecture. Here is what actually changed.
| Capability | Hailuo 02 | MiniMax H3 |
|---|---|---|
| Model design | Video-first architecture | Omni-modal unified context, built new |
| Resolution | 768P / 1080P | Native 2K default, 768P economy tier |
| Clip length | Fixed 6s or 10s | Any length from 4 to 15 seconds |
| Audio | Silent output | Native stereo sound in every take |
| References | First-frame image | 9 images + 3 videos + 3 audio clips |
| Billing | Per fixed clip | Per second — pay for exactly what you generate |
Specs verified 2026-08-02.
MiniMax H3 vs models live on Molyin
MiniMax H3 is live on Molyin — and so is every model it competes with. Long takes, audio, reference control: pick your tool, one credit balance.
| Model | Positioning | Release Date | On Molyin |
|---|---|---|---|
| MiniMax H3 | Omni-modal flagship, native 2K + stereo | Jul 31, 2026 · open weights announced | Live — generate above |
| Seedance 2.5 | ByteDance flagship, up to 30s takes | Aug 7, 2026 (API) | Live since API day one — generate above |
| Wan 3.0 | Alibaba's All-in-One model, up to 30s takes | Aug 6, 2026 | Live since launch day — generate above |
| HappyHorse | Expressive generation + instruction video editing | Jun 2026 (1.1) | Live — generate above |
| Seedance 2.0 family | Standard / Fast / Mini workhorses, up to 4K | Apr 2026 (API) | Live — generate above |
| Kling 3.0 | Element references + native audio, up to 4K | Feb 4, 2026 | Live — generate above |
| Veo 3.1 | Google's cinematic model with audio | Oct 2025 | Live — generate above |
One subscription, one credit balance — every live model above is included.
MiniMax H3 full specifications
Every value below is derived live from the same model and billing configuration that powers the generator above — nothing is hand-written, so this table can never drift from what you actually get.
| Parameter | Value | Notes |
|---|---|---|
| Generation modes | Text / Image / Reference to video | Three modes, one unified multimodal context — switch freely in the generator above. |
| Resolution tiers | 768P / 2K | Native 2K — 2560×1440 for 16:9 — regenerated in context by the base model, not upscaled; 768P is the economy tier. |
| Frame rate | 24 fps | Cinema-standard output frame rate. |
| Clip length | 4–15 s | Any integer length in a single take, billed per second — no fixed 6s/10s tiers. |
| Audio | Native stereo, always on | Music, dialogue and effects are generated jointly with the picture; vocals can follow an audio reference. |
| Reference images | Up to 9 | The first few images per request are free; extras carry a small per-image surcharge — exact numbers in the FAQ below. |
| Reference videos | Up to 3 | Each clip 2–15 s, combined reference footage up to 15 s; reference seconds are billed at the output tier rate. |
| Reference audio tracks | Up to 3 | Audio references cannot be the only input — pair them with an image or a video. |
| Credits per second | 18 (768P) · 29 (2K) | Billed per second from your credit balance; failed generations are refunded automatically. |
Everything listed here is already supported on the site, and stays in sync with the model's latest official capabilities.
MiniMax H3 pricing on Molyin
Billed per second from the same credit balance as every other live model. The numbers below are computed straight from our billing engine — and failed generations are refunded automatically.
| Model | Resolution | Credits/sec | 5s Video | 5s Video + 3s Reference |
|---|---|---|---|---|
| 768p | 18 | 90 credits | 144 credits | |
| 2k | 29 | 145 credits | 232 credits |
- Exception: reference video seconds for MiniMax H3 are billed at the full per-second rate.
- MiniMax H3: the first 5 reference images are free, each additional image adds 9 credits.
How to generate with MiniMax H3
Three steps from idea to a native 2K film with stereo sound
Pick MiniMax H3 and a mode
Choose text-to-video, image-to-video or reference-to-video.
Write the prompt, add references
Describe the shot like a director. Optionally attach images, video clips and audio, and say how they relate.
Generate and export
MiniMax H3 delivers picture and stereo sound in one pass. Preview in place, then download the original file.
The three-track prompt formula
MiniMax H3 generates picture, ambient sound and score in one pass — so a good prompt writes for eyes and ears at once. Straight from our full prompt guide, with 12 ready-to-run examples.
[Subject] + [Action] + optional [Style] [Camera] [Ambient sound] [Score]
Subject
Who or what the shot is about — the more concrete the description, the more stable the result.
Action
What the subject does, one clear beat at a time. Vague verbs produce vague motion.
Style
The overall look, stated up front: live-action cinematic, 2D animation, claymation, vintage film.
Camera
Shot size and movement. H3 follows explicit camera language well — chapter 2 gives the full vocabulary.
Ambient sound
What the scene itself sounds like: rain, traffic, sizzling oil. Characters can hear this.
Score
Music only the audience hears. Name instruments and tempo, not moods.
MiniMax H3 — FAQ
Specs, release status and how to try it
Start Creating with Molyin Today
Sign up free and turn your first idea into a cinematic video in minutes.