Create
Support
MiniMax H3

MiniMax H3 (Hailuo 3)

An omni-modal video model that reads text, images, video and audio in one context — and answers with a native 2K film, stereo sound included, up to 15 seconds in a single take. Live on Molyin — generate below.

  • Failed generations auto-refunded
  • Credit packs never expire
  • Every model's price is public
  • Cancel anytime in one click
2K
Native resolution
24 fps
Frame rate
4–15 s
Any clip length
Stereo
Sound in every take

Official MiniMax H3 samples

Straight from the MiniMax release — hover to preview, tap the speaker icon to hear the native stereo mix.

  • MiniMax H3
  • Reference to video
  • 15s
  • 2K

Image 1 serves as a reference for the overall texture and mood, while Image 2 provides reference for the character's appearance. Create a 15-second, 16:9 widescreen trendy fashion short film. Maintain consistency in character design: platinum blonde long hair, narrow black vintage sunglasses, black glossy patent leather trench coat, a cold and confident fashion expression, and orange reflections of flames cast on the surface of the leather coat. The overall style is a fast-cut film-like fashion advertisement, blending elements such as nighttime fire scenes, black smoke, orange-red flames, VHS glitches, CCTV broadcast interruptions, 90s analog film grain, scan-lines, chromatic aberration, light leaks, flash-white transitions, and slight camera shake throughout.

  • MiniMax H3
  • Reference to video
  • 15s
  • 2K

Create a 16:9 widescreen, high-end fashion brand video. The overall mood, scenes, and film-like texture should reference Image 1; bag assets should reference Image 3; character assets should reference Image 2; and the brand ending logo should reference Image 4. This is a fashion campaign for selling clothes and bags. The overall vibe should be high-end, cool, and restrained, but the editing cannot feel dull—it needs to be more dynamic and infused with a strong sense of fashion rhythm. Avoid the feel of an ordinary narrative film or an e-commerce ad. The core story remains simple: on a desert highway beside a vintage car, a woman walks to the car's trunk, opens it, and retrieves a black bag. There's a brief, quiet connection between her and a man standing beside the car. Then, she walks away with the bag. The bag and the clothing should naturally integrate into the characters' movements, becoming an intrinsic part of their personas.

  • MiniMax H3
  • Text to video
  • 15s
  • 2K

A faster pace, grand yet not drawn out. Quick hard cuts, bridge shaking, intense light flashes, short blackouts, and warp jumps. The text design resembles the wide-spaced lettering of a movie trailer, avoiding plain white but featuring a restrained glow and textured finish with subtle edge highlights. The text animation includes emerging from the deep void of space, illuminated by starlight sweeps, expanding letter spacing, afterimages, subtle glowing, and flashes of blackouts.

  • MiniMax H3
  • Reference to video
  • 15s
  • 2K

Style: A dark-pop / cyber-grunge / rap music video with a high-fashion yet realistic texture, evoking a film magazine aesthetic. The visuals feature high contrast but avoid appearing cheap. The overall reference draws from late-90s to early-2000s independent magazines, photocopied zines, film scans, underground music posters, and collage aesthetics. Visuals include coarse grain, subtle film jitter, halftone patterns, print edge roughness, and scanning misalignments. The editing rhythm is fast, relying solely on hard cuts with no fades or soft transitions. The text design style and texture should reference the attached imagery.

  • MiniMax H3
  • Text to video
  • 15s
  • 2K

A 15-second, 16:9 horizontal short video. A live-action nighttime laundromat scene blended with hand-drawn glowing animations. The setting is a small self-service laundromat with faintly flickering fluorescent lights. Inside are running washing machines, plastic laundry baskets, an old bench, and a single sock discarded on the floor. The overall space feels quiet, carrying a subtle sense of nostalgia. Filmed with a handheld phone in single-hand grip, the footage has noticeable shakiness; the white fluorescent lighting causes fluctuating exposure; surface reflections are visible on the glass panels; and the focus lags slightly when the camera gets close to objects. The visuals should not look polished or meticulous like a commercial ad. The overall texture should resemble a candid, spontaneous recording—like stumbling into the laundromat late at night and instinctively capturing surreal visions with a raw documentary feel.

  • MiniMax H3
  • Text to video
  • 15s
  • 2K

Generate a 15-second, 16:9 widescreen light suspense crime film opening sequence. The overall style should reference the visual language of these concepts: retro Japanese anime openings, hard-edged silhouettes, comic-style collage, asymmetrical split screens, strong geometric color blocks, English title credits, minimal Japanese katakana for decoration, and a jazz-crime vibe. The atmosphere should be 60% suspense and 40% jazz: mysterious, cool, agile, with an urban crime feel—neither scary nor heavy, and definitely not a joyful jazz music video. The motion effects of the opening should feel like animated graphic collages: a black-lined frame appears first, split-screen borders are quickly drawn out, color blocks and frames are pasted in piece by piece; silhouettes of characters, close-ups of props, and English title credits slide in, bounce out, or are revealed through masking along with the drum beats. Avoid realistic narrative animations—this should have the polished feel of an opening sequence. The English credits should be legible, with animated effects allowed: thin line frames can first be drawn out, names slide into the frame, with letters appearing one at a time and revealed by moving color blocks, before pausing briefly at the end. Do not add Chinese text, avoid garbled text, and ensure no misspelled English names. Rules for the sequence: Every English credit and title should only appear once. Do not repeat the same title, do not repeat the same name, and do not assign multiple roles to the same individual. Transitions must be varied: circular vinyl record masks, vertical cuts through car doors, character shadows acting as wipes, red-line cuts, large English letter masking, split-frame border reorganizations, hard geometric color-block cuts, and individual frame-by-frame collage reveals. All transitions should sync with the drum beats—precise, suspenseful, agile, and with a comic-collage vibe. Avoid soft dissolves and fluid transitions. BGM: Create an original BGM.

One context in. One film out.

This is MiniMax's own flagship example. Three assets — a camera-movement reference, a character image and a vocal recording — plus one sentence describing how they fit together. MiniMax H3 reads them as a single context and delivers the finished, singing shot.

Video 1 — camera movement reference
Image 2 — character reference
Image 2 — character reference
Audio 3 — vocal reference

The prompt

Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3.

The generated film — press play to hear the character sing the referenced vocals. Official MiniMax sample.

What is MiniMax H3?

MiniMax H3 is the third generation of MiniMax's H series, released on July 31, 2026 — and it is not just a video model. MiniMax H3 is a general-purpose multimodal generation model: it understands text, images, video clips and audio inside one unified context, and you describe how they relate in plain language. Ask it to borrow the camera movement from one video, put the character from an image on screen, and have them sing the vocals from an audio file — in a single prompt.

The output side is just as ambitious. MiniMax H3 renders in native 2K by default — not an upscaled image, but in-context regeneration by the base model, which is why fine details like small text survive. Every generation ships with native stereo sound: music, dialogue and effects are produced jointly with the picture, not bolted on afterwards. Add native multi-shot modeling, accurate text and brand rendering, and motion transfer from reference footage, and MiniMax H3 reads like a model built for commercial work — ads, e-commerce, product and UI films. MiniMax has also announced plans to open the model weights.

While MiniMax H3 rolls out, every capability it competes with is already live on Molyin: try our AI video generator, check pricing, or browse the full model library.

MiniMax H3 vs Hailuo 02

MiniMax H3 replaces the Hailuo 02 generation with a new architecture. Here is what actually changed.

CapabilityHailuo 02MiniMax H3
Model designVideo-first architectureOmni-modal unified context, built new
Resolution768P / 1080PNative 2K default, 768P economy tier
Clip lengthFixed 6s or 10sAny length from 4 to 15 seconds
AudioSilent outputNative stereo sound in every take
ReferencesFirst-frame image9 images + 3 videos + 3 audio clips
BillingPer fixed clipPer second — pay for exactly what you generate

Specs verified 2026-08-02.

MiniMax H3 vs models live on Molyin

MiniMax H3 is live on Molyin — and so is every model it competes with. Long takes, audio, reference control: pick your tool, one credit balance.

ModelPositioningRelease DateOn Molyin
MiniMax H3Omni-modal flagship, native 2K + stereoJul 31, 2026 · open weights announcedLive — generate above
Seedance 2.5ByteDance flagship, up to 30s takesAug 7, 2026 (API)Live since API day one — generate above
Wan 3.0Alibaba's All-in-One model, up to 30s takesAug 6, 2026Live since launch day — generate above
HappyHorseExpressive generation + instruction video editingJun 2026 (1.1)Live — generate above
Seedance 2.0 familyStandard / Fast / Mini workhorses, up to 4KApr 2026 (API)Live — generate above
Kling 3.0Element references + native audio, up to 4KFeb 4, 2026Live — generate above
Veo 3.1Google's cinematic model with audioOct 2025Live — generate above

One subscription, one credit balance — every live model above is included.

MiniMax H3 full specifications

Every value below is derived live from the same model and billing configuration that powers the generator above — nothing is hand-written, so this table can never drift from what you actually get.

MiniMax H3 full specifications
ParameterValueNotes
Generation modesText / Image / Reference to videoThree modes, one unified multimodal context — switch freely in the generator above.
Resolution tiers768P / 2KNative 2K — 2560×1440 for 16:9 — regenerated in context by the base model, not upscaled; 768P is the economy tier.
Frame rate24 fpsCinema-standard output frame rate.
Clip length4–15 sAny integer length in a single take, billed per second — no fixed 6s/10s tiers.
AudioNative stereo, always onMusic, dialogue and effects are generated jointly with the picture; vocals can follow an audio reference.
Reference imagesUp to 9The first few images per request are free; extras carry a small per-image surcharge — exact numbers in the FAQ below.
Reference videosUp to 3Each clip 2–15 s, combined reference footage up to 15 s; reference seconds are billed at the output tier rate.
Reference audio tracksUp to 3Audio references cannot be the only input — pair them with an image or a video.
Credits per second18 (768P) · 29 (2K)Billed per second from your credit balance; failed generations are refunded automatically.

Everything listed here is already supported on the site, and stays in sync with the model's latest official capabilities.

MiniMax H3 pricing on Molyin

Billed per second from the same credit balance as every other live model. The numbers below are computed straight from our billing engine — and failed generations are refunded automatically.

AI video model credit rates by resolution — credits per second and per-clip examples
ModelResolutionCredits/sec5s Video5s Video + 3s Reference
768p1890 credits144 credits
2k29145 credits232 credits
  • Exception: reference video seconds for MiniMax H3 are billed at the full per-second rate.
  • MiniMax H3: the first 5 reference images are free, each additional image adds 9 credits.

See credit pricing for every model on the full price list →

How to generate with MiniMax H3

Three steps from idea to a native 2K film with stereo sound

Pick MiniMax H3 and a mode

Choose text-to-video, image-to-video or reference-to-video.

Write the prompt, add references

Describe the shot like a director. Optionally attach images, video clips and audio, and say how they relate.

Generate and export

MiniMax H3 delivers picture and stereo sound in one pass. Preview in place, then download the original file.

The three-track prompt formula

MiniMax H3 generates picture, ambient sound and score in one pass — so a good prompt writes for eyes and ears at once. Straight from our full prompt guide, with 12 ready-to-run examples.

[Subject] + [Action] + optional [Style] [Camera] [Ambient sound] [Score]

Required

Subject

Who or what the shot is about — the more concrete the description, the more stable the result.

Required

Action

What the subject does, one clear beat at a time. Vague verbs produce vague motion.

Optional

Style

The overall look, stated up front: live-action cinematic, 2D animation, claymation, vintage film.

Optional

Camera

Shot size and movement. H3 follows explicit camera language well — chapter 2 gives the full vocabulary.

Optional

Ambient sound

What the scene itself sounds like: rain, traffic, sizzling oil. Characters can hear this.

Optional

Score

Music only the audience hears. Name instruments and tempo, not moods.

Read the full MiniMax H3 prompt guide

MiniMax H3 — FAQ

Specs, release status and how to try it

Start Creating with Molyin Today

Sign up free and turn your first idea into a cinematic video in minutes.