Create
Support
FLUX 3

FLUX 3: One Multimodal Model for Video, Image & Audio

Black Forest Labs' world model generates picture and sound in a single pass — write a prompt below and FLUX 3 renders 5–20 seconds of footage with dialogue and effects built in.

  • Failed generations auto-refunded
  • Credit packs never expire
  • Every model's price is public
  • Cancel anytime in one click
5–20 s
Clip length
720P / 1080P
Resolution
Native, on/off
Audio
3–10
Storyboard keyframes

Real Prompts, Real FLUX 3 Footage

Every clip below was generated by FLUX 3 from the exact prompt shown — official samples and community work, sound and all. Hit Remix to load any prompt into the generator above and make it yours.

  • FLUX 3
  • Text to video
  • 10s
  • 1080P

A cozy ramen shop on a rainy Tokyo night: steam rising from the broth, neon reflections in the window puddles, the cook working calmly. The camera drifts slowly past the counter. Rain patter and quiet kitchen sounds.

  • FLUX 3
  • Text to video
  • 20s
  • 720P

ranting about ai

  • FLUX 3
  • Text to video
  • 20s
  • 720P

How pyramids are made? Visualize it.

  • FLUX 3
  • Text to video
  • 15s
  • 720P

claymation story about a man consistently getting smushed by a large giant hand that enters the scene so he cant live his life

  • FLUX 3
  • Text to video
  • 20s
  • 720P

Daily life of the Aztecs

One Photo In, a Fireworks Finale Out

FLUX 3's image-to-video mode animates a single still into motion. Here is the official example end to end: one fireworks photo plus a one-line prompt becomes a 10-second 1080P finale — crowd murmur and explosions generated natively with the picture.

Start frame — fireworks photo
Start frame — fireworks photo

Prompt

A fourth of July fireworks display concludes with a dazzling finale of red, white and blue fireworks.

10s · 1080P · native audio — generated by FLUX 3 image-to-video

What Is FLUX 3?

FLUX 3 is the multimodal foundation model announced by Black Forest Labs — the lab behind the FLUX image family — on July 23, 2026. Where FLUX 1 and FLUX.2 generated images, FLUX 3 learns from images, videos, and audio jointly inside one architecture. The bet is simple: a model that must render how objects move, collide, and sound is forced to learn how the world actually behaves — and every modality it generates benefits from that shared understanding.

That world-model foundation shows up in the footage. FLUX 3 generates clips up to 20 seconds long with native audio in a single pass — dialogue in multiple languages synced to lip movement, sound effects pinned to the physical events that cause them, styles ranging from handheld camcorder footage to claymation and cinematic shots. Give it three words and it stages a coherent scene; give it storyboard keyframes and it directs the transitions between your exact frames. On the image side it handles complex prompts markedly better than earlier FLUX generations and renders accurate multilingual text.

FLUX 3 video is live on Molyin today — the video side of the AI video generator above runs it directly, with storyboard keyframes and clip extension included. The image side is still in early access, so FLUX 3 stays marked coming soon in the AI image generator — the slot is ready, and both sides run on one credit system that covers every model.

FLUX 3 vs. Seedance 2.5

The two newest flagship video models on Molyin, side by side. They complement rather than replace each other: FLUX 3 leads on native dialogue and keyframe direction, Seedance 2.5 on 30-second takes and reference capacity — one credit balance covers both.

CapabilityFLUX 3Seedance 2.5
Native audioDialogue + sound effects in the same pass, toggle on/off at the same priceJoint audio-visual generation, always on
Clip length5–20 s, any length in range4–30 s, any length in range
Resolution720P / 1080P480P / 720P
Reference inputs3–10 storyboard keyframes drive the shot sequenceUp to 50 reference assets — images, clips, and audio
Clip extensionFLUX 3 Extend model — continues an existing clip, 5–20 sSeedance 2.5 Extend model — continues from the last frame, 4–30 s
Multilingual dialogueCharacters speak multiple languages, lip-syncedMultilingual lip-sync — eight languages demonstrated

Both models are live in the generator above — switch between them without leaving the page.

FLUX 3 vs models live on Molyin

Every row in this table is live on Molyin right now. Native dialogue, storyboard keyframes, 30-second takes, 4K — pick the tool per shot, one credit balance for all of them.

ModelPositioningRelease DateOn Molyin
FLUX 3Black Forest Labs' world model — native audio + storyboard keyframesAug 7, 2026 (video API) · open weights announcedLive — generate above
Seedance 2.5ByteDance flagship — 30s takes, directed second by secondAug 7, 2026 (API)Live — generate above
Wan 3.0Alibaba's All-in-One model, up to 30s takesAug 6, 2026Live — generate above
MiniMax H3Omni-modal flagship, native 2K + stereoJul 31, 2026 · open weights announcedLive — generate above
HappyHorseExpressive generation + instruction video editingJun 2026 (1.1)Live — generate above
Seedance 2.0 familyStandard / Fast / Mini workhorses, up to 4KApr 2026 (API)Live — generate above
Kling 3.0Element references + native audio, up to 4KFeb 4, 2026Live — generate above
Veo 3.1Google's cinematic model with audioOct 2025Live — generate above

One subscription, one credit balance — every live model above is included.

FLUX 3 Full Specifications

Every FLUX 3 parameter on Molyin, straight from the model registry and billing engine — what this table says is what the generator enforces.

FLUX 3 Full Specifications
ParameterValueNotes
Generation modesText-to-video · Image-to-video (first/last frame) · Storyboard keyframes · Clip extensionStoryboard keyframes run in reference mode; clip extension is its own model — pick FLUX 3 Extend in video-extend mode.
Resolution720P / 1080PBoth tiers generate with native audio; each tier has its own per-second rate.
Clip length5–20 secondsAny length in the range — billing is per second of output.
Aspect ratiosauto · 16:9 · 9:16 · 4:3 · 3:4 · 1:1 · 21:9 · 2:1auto lets the model pick from the prompt and inputs; 2:1 widescreen is a FLUX 3 exclusive on Molyin.
AudioNative — dialogue and sound effects, toggle on/offGenerated in the same pass as the picture; same price with audio on or off.
LanguagesMultilingual dialogue and on-screen textCharacters speak multiple languages lip-synced, and rendered signs or titles stay legible.
Storyboard keyframes3–10 imagesKeyframes are spread across the timeline; the model generates the motion between your exact frames.
Clip extension1 source clipThe FLUX 3 Extend model continues an existing clip; new footage bills at the extension rate below.
SeedNot supportedVary the wording to explore alternatives — identical prompts still differ between runs.
Credits per second38 (720P) · 65 (1080P)Text-to-video, image-to-video, and storyboard keyframes all bill at this rate.
Extension credits per second91 (720P) · 118 (1080P)Applies only to newly generated footage when extending a clip — the source clip itself is free.

Everything listed here is already supported on the site, and stays in sync with the model's latest official capabilities.

What FLUX 3 Costs in Credits

Transparent per-second pricing from the same billing engine that charges the generator — audio adds nothing, and failed generations are refunded automatically.

AI video model credit rates by resolution — credits per second and per-clip examples
ModelResolutionCredits/sec5s Video5s Video + 3s Reference
720p38190 creditsReference images only
1080p65325 creditsReference images only
720p91455 credits455 credits (source video free)
1080p118590 credits590 credits (source video free)
  • Video-extend models (FLUX 3 Extend) never bill the source video — only the newly generated seconds are billed, at the extend rate shown above.

See credit pricing for every model

How to generate with FLUX 3

Three steps from a prompt to footage with sound

Pick FLUX 3 and a mode

Choose text-to-video, image-to-video, or reference mode for storyboard keyframes; to extend a clip, pick the FLUX 3 Extend model.

Write the scene, sound included

Describe the shot and put dialogue in quotes — FLUX 3 renders the audio you write. Attach keyframes in reference mode, or hand FLUX 3 Extend a source clip, to direct the timeline.

Generate and export

FLUX 3 delivers picture and sound in one pass, 5 to 20 seconds at 720P or 1080P. Preview in place, then download the original file.

Write sound into your prompts

FLUX 3 renders the audio you describe — pin each sound cue to the physical event that causes it and the mix follows the picture. Straight from our full FLUX 3 prompt guide.

[Visual event] with [its sound] + optional [ambient bed]

Required

Visual Event

The on-screen action that makes noise: a door slams, glass clinks, footsteps splash.

Required

Sound Cue

The effect pinned to that event — named concretely, in the same sentence as the action.

Optional

Ambience

The background bed: wind, crowd murmur, room tone. One phrase is enough.

Read the full FLUX 3 prompt guide

FLUX 3 FAQ

What creators ask about Black Forest Labs' multimodal model.

Start Creating with Molyin Today

Sign up free and turn your first idea into a cinematic video in minutes.