FLUX 3: One Multimodal Model for Video, Image & Audio
Black Forest Labs' world model generates picture and sound in a single pass — write a prompt below and FLUX 3 renders 5–20 seconds of footage with dialogue and effects built in.
- Failed generations auto-refunded
- Credit packs never expire
- Every model's price is public
- Cancel anytime in one click
- 5–20 s
- Clip length
- 720P / 1080P
- Resolution
- Native, on/off
- Audio
- 3–10
- Storyboard keyframes
Real Prompts, Real FLUX 3 Footage
Every clip below was generated by FLUX 3 from the exact prompt shown — official samples and community work, sound and all. Hit Remix to load any prompt into the generator above and make it yours.
- FLUX 3
- Text to video
- 10s
- 1080P
A cozy ramen shop on a rainy Tokyo night: steam rising from the broth, neon reflections in the window puddles, the cook working calmly. The camera drifts slowly past the counter. Rain patter and quiet kitchen sounds.
- FLUX 3
- Text to video
- 20s
- 720P
ranting about ai
- FLUX 3
- Text to video
- 20s
- 720P
How pyramids are made? Visualize it.
- FLUX 3
- Text to video
- 15s
- 720P
claymation story about a man consistently getting smushed by a large giant hand that enters the scene so he cant live his life
- FLUX 3
- Text to video
- 20s
- 720P
Daily life of the Aztecs
One Photo In, a Fireworks Finale Out
FLUX 3's image-to-video mode animates a single still into motion. Here is the official example end to end: one fireworks photo plus a one-line prompt becomes a 10-second 1080P finale — crowd murmur and explosions generated natively with the picture.

Prompt
A fourth of July fireworks display concludes with a dazzling finale of red, white and blue fireworks.
What Is FLUX 3?
FLUX 3 is the multimodal foundation model announced by Black Forest Labs — the lab behind the FLUX image family — on July 23, 2026. Where FLUX 1 and FLUX.2 generated images, FLUX 3 learns from images, videos, and audio jointly inside one architecture. The bet is simple: a model that must render how objects move, collide, and sound is forced to learn how the world actually behaves — and every modality it generates benefits from that shared understanding.
That world-model foundation shows up in the footage. FLUX 3 generates clips up to 20 seconds long with native audio in a single pass — dialogue in multiple languages synced to lip movement, sound effects pinned to the physical events that cause them, styles ranging from handheld camcorder footage to claymation and cinematic shots. Give it three words and it stages a coherent scene; give it storyboard keyframes and it directs the transitions between your exact frames. On the image side it handles complex prompts markedly better than earlier FLUX generations and renders accurate multilingual text.
FLUX 3 video is live on Molyin today — the video side of the AI video generator above runs it directly, with storyboard keyframes and clip extension included. The image side is still in early access, so FLUX 3 stays marked coming soon in the AI image generator — the slot is ready, and both sides run on one credit system that covers every model.
FLUX 3 vs. Seedance 2.5
The two newest flagship video models on Molyin, side by side. They complement rather than replace each other: FLUX 3 leads on native dialogue and keyframe direction, Seedance 2.5 on 30-second takes and reference capacity — one credit balance covers both.
| Capability | FLUX 3 | Seedance 2.5 |
|---|---|---|
| Native audio | Dialogue + sound effects in the same pass, toggle on/off at the same price | Joint audio-visual generation, always on |
| Clip length | 5–20 s, any length in range | 4–30 s, any length in range |
| Resolution | 720P / 1080P | 480P / 720P |
| Reference inputs | 3–10 storyboard keyframes drive the shot sequence | Up to 50 reference assets — images, clips, and audio |
| Clip extension | FLUX 3 Extend model — continues an existing clip, 5–20 s | Seedance 2.5 Extend model — continues from the last frame, 4–30 s |
| Multilingual dialogue | Characters speak multiple languages, lip-synced | Multilingual lip-sync — eight languages demonstrated |
Both models are live in the generator above — switch between them without leaving the page.
FLUX 3 vs models live on Molyin
Every row in this table is live on Molyin right now. Native dialogue, storyboard keyframes, 30-second takes, 4K — pick the tool per shot, one credit balance for all of them.
| Model | Positioning | Release Date | On Molyin |
|---|---|---|---|
| FLUX 3 | Black Forest Labs' world model — native audio + storyboard keyframes | Aug 7, 2026 (video API) · open weights announced | Live — generate above |
| Seedance 2.5 | ByteDance flagship — 30s takes, directed second by second | Aug 7, 2026 (API) | Live — generate above |
| Wan 3.0 | Alibaba's All-in-One model, up to 30s takes | Aug 6, 2026 | Live — generate above |
| MiniMax H3 | Omni-modal flagship, native 2K + stereo | Jul 31, 2026 · open weights announced | Live — generate above |
| HappyHorse | Expressive generation + instruction video editing | Jun 2026 (1.1) | Live — generate above |
| Seedance 2.0 family | Standard / Fast / Mini workhorses, up to 4K | Apr 2026 (API) | Live — generate above |
| Kling 3.0 | Element references + native audio, up to 4K | Feb 4, 2026 | Live — generate above |
| Veo 3.1 | Google's cinematic model with audio | Oct 2025 | Live — generate above |
One subscription, one credit balance — every live model above is included.
FLUX 3 Full Specifications
Every FLUX 3 parameter on Molyin, straight from the model registry and billing engine — what this table says is what the generator enforces.
| Parameter | Value | Notes |
|---|---|---|
| Generation modes | Text-to-video · Image-to-video (first/last frame) · Storyboard keyframes · Clip extension | Storyboard keyframes run in reference mode; clip extension is its own model — pick FLUX 3 Extend in video-extend mode. |
| Resolution | 720P / 1080P | Both tiers generate with native audio; each tier has its own per-second rate. |
| Clip length | 5–20 seconds | Any length in the range — billing is per second of output. |
| Aspect ratios | auto · 16:9 · 9:16 · 4:3 · 3:4 · 1:1 · 21:9 · 2:1 | auto lets the model pick from the prompt and inputs; 2:1 widescreen is a FLUX 3 exclusive on Molyin. |
| Audio | Native — dialogue and sound effects, toggle on/off | Generated in the same pass as the picture; same price with audio on or off. |
| Languages | Multilingual dialogue and on-screen text | Characters speak multiple languages lip-synced, and rendered signs or titles stay legible. |
| Storyboard keyframes | 3–10 images | Keyframes are spread across the timeline; the model generates the motion between your exact frames. |
| Clip extension | 1 source clip | The FLUX 3 Extend model continues an existing clip; new footage bills at the extension rate below. |
| Seed | Not supported | Vary the wording to explore alternatives — identical prompts still differ between runs. |
| Credits per second | 38 (720P) · 65 (1080P) | Text-to-video, image-to-video, and storyboard keyframes all bill at this rate. |
| Extension credits per second | 91 (720P) · 118 (1080P) | Applies only to newly generated footage when extending a clip — the source clip itself is free. |
Everything listed here is already supported on the site, and stays in sync with the model's latest official capabilities.
What FLUX 3 Costs in Credits
Transparent per-second pricing from the same billing engine that charges the generator — audio adds nothing, and failed generations are refunded automatically.
| Model | Resolution | Credits/sec | 5s Video | 5s Video + 3s Reference |
|---|---|---|---|---|
| 720p | 38 | 190 credits | Reference images only | |
| 1080p | 65 | 325 credits | Reference images only | |
| 720p | 91 | 455 credits | 455 credits (source video free) | |
| 1080p | 118 | 590 credits | 590 credits (source video free) |
- Video-extend models (FLUX 3 Extend) never bill the source video — only the newly generated seconds are billed, at the extend rate shown above.
How to generate with FLUX 3
Three steps from a prompt to footage with sound
Pick FLUX 3 and a mode
Choose text-to-video, image-to-video, or reference mode for storyboard keyframes; to extend a clip, pick the FLUX 3 Extend model.
Write the scene, sound included
Describe the shot and put dialogue in quotes — FLUX 3 renders the audio you write. Attach keyframes in reference mode, or hand FLUX 3 Extend a source clip, to direct the timeline.
Generate and export
FLUX 3 delivers picture and sound in one pass, 5 to 20 seconds at 720P or 1080P. Preview in place, then download the original file.
Write sound into your prompts
FLUX 3 renders the audio you describe — pin each sound cue to the physical event that causes it and the mix follows the picture. Straight from our full FLUX 3 prompt guide.
[Visual event] with [its sound] + optional [ambient bed]
Visual Event
The on-screen action that makes noise: a door slams, glass clinks, footsteps splash.
Sound Cue
The effect pinned to that event — named concretely, in the same sentence as the action.
Ambience
The background bed: wind, crowd murmur, room tone. One phrase is enough.
FLUX 3 FAQ
What creators ask about Black Forest Labs' multimodal model.
Start Creating with Molyin Today
Sign up free and turn your first idea into a cinematic video in minutes.