MiniMax · H series, third generation

MiniMax H3

An omni-modal video model that reads text, images, video and audio in one context — and answers with a native 2K film, stereo sound included, up to 15 seconds in a single take.

Facts verified 2026-08-02

Official H3 sample — native 2K, shown as a muted preview
2K
Native resolution
24 fps
Frame rate
4–15 s
Any clip length
Stereo
Sound in every take

Try the studio H3 will live in

MiniMax H3 is coming to Molyin and is preselected below. Switch to any live model — Seedance 2.5, Kling 3.0, Veo 3.1, HappyHorse — and start creating with one credit balance.

MiniMax H3 isn't generating yet — it activates the day access opens.

What is MiniMax H3?

MiniMax H3 is the third generation of MiniMax's H series, released on July 31, 2026 — and it is not just a video model. H3 is a general-purpose multimodal generation model: it understands text, images, video clips and audio inside one unified context, and you describe how they relate in plain language. Ask it to borrow the camera movement from one video, put the character from an image on screen, and have them sing the vocals from an audio file — in a single prompt.

The output side is just as ambitious. H3 renders in native 2K by default — not an upscaled image, but in-context regeneration by the base model, which is why fine details like small text survive. Every generation ships with native stereo sound: music, dialogue and effects are produced jointly with the picture, not bolted on afterwards. Add native multi-shot modeling, accurate text and brand rendering, and motion transfer from reference footage, and H3 reads like a model built for commercial work — ads, e-commerce, product and UI films. MiniMax has also announced plans to open the model weights.

While H3 rolls out, every capability it competes with is already live on Molyin: try our AI video generator, check pricing, or browse the full model library.

One context in. One film out.

This is MiniMax's own flagship example. Three assets — a camera-movement reference, a character image and a vocal recording — plus one sentence describing how they fit together. H3 reads them as a single context and delivers the finished, singing shot.

Video 1 — camera movement reference
Image 2 — character reference
Image 2 — character reference
Audio 3 — vocal reference

The prompt

Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3.

The generated film — press play to hear the character sing the referenced vocals. Official MiniMax sample.

What makes H3 different

Six capabilities that separate H3 from single-task video models.

One unified context

Text, images, video and audio live in the same context window. You describe relationships in natural language instead of wiring up fixed task modes.

Native 2K, no upscaler

2K output comes from in-context regeneration by the base model, not a super-resolution pass — small text and fine detail stay legible.

Stereo sound in one pass

Music, dialogue, foley and ambience are generated together with the picture. Every take arrives with a native stereo mix.

Native multi-shot

Cuts and shot changes are modeled inside the clip itself, keeping characters and style coherent across shots.

9 + 3 + 3 references

Up to 9 reference images, 3 video clips and 3 audio tracks in one request — characters, motion and voices stay consistent.

Built for brand work

Strong instruction following, accurate text and brand rendering, and V2V motion transfer — aimed squarely at ads, e-commerce, product and UI films.

Official H3 samples

Straight from the MiniMax release — hover to preview, tap the speaker icon to hear the native stereo mix.

  • MiniMax H3
  • Stereo audio
  • MiniMax H3
  • Stereo audio
  • MiniMax H3
  • Stereo audio
  • MiniMax H3
  • Stereo audio
  • MiniMax H3
  • Stereo audio

MiniMax H3 vs Hailuo 02

H3 replaces the Hailuo 02 generation with a new architecture. Here is what actually changed.

CapabilityHailuo 02MiniMax H3
Model designVideo-first architectureOmni-modal unified context, built new
Resolution768P / 1080PNative 2K default, 768P economy tier
Clip lengthFixed 6s or 10sAny length from 4 to 15 seconds
AudioSilent outputNative stereo sound in every take
ReferencesFirst-frame image9 images + 3 videos + 3 audio clips
BillingPer fixed clipPer second — pay for exactly what you generate

Specs from MiniMax's official H3 release and API documentation, verified 2026-08-02.

MiniMax H3 vs models live on Molyin

H3 is on its way. Until it lands, the capabilities it competes on — long takes, audio, references — are already live here.

ModelPositioningAvailabilityOn Molyin
MiniMax H3Omni-modal flagship, native 2K + stereoReleased Jul 31, 2026 · open weights announcedComing soon — preselected above
Seedance 2.5ByteDance flagship, up to 30s takesLiveGenerate now
Seedance 2.0 familyStandard / Fast / Mini workhorses, up to 4KLiveGenerate now
Kling 3.0Element references + native audio, up to 4KLiveGenerate now
Veo 3.1Google's cinematic model with audioLiveGenerate now
HappyHorseExpressive generation + instruction video editingLiveGenerate now

One subscription, one credit balance — every live model above is included.

MiniMax H3 — FAQ

Specs, release status and how to try it

Start Creating with Molyin Today

Sign up free and turn your first idea into a cinematic video in minutes.