Extended 7 Days

Start creating
Black Forest Labs · Announced July 23, 2026 · Coming soon on Molyin

FLUX 3: One Multimodal Model for Video, Image & Audio

FLUX 3 is Black Forest Labs' new multimodal foundation model — one architecture that generates video with native audio, synthesizes and edits images, and even drives robots. Here's what it does, when you can use it, and how to create with its closest rivals while you wait.

What it is
Black Forest Labs' multimodal foundation model: video, image, and audio generated by a single architecture.
Video
Up to 20 seconds per generation with native audio — text-to-video, image-to-video, video-to-video, and keyframe transitions.
Image
Synthesis and editing with accurate multilingual text rendering; image early access opens in the weeks after announcement.
On Molyin
Coming soon — already in both the video and image model pickers below. Create with Seedance 2.0 or GPT Image 2 today.

Facts on this page verified as of 2026-07-25.

Try the FLUX 3 Workflow Today

FLUX 3 is preselected in both generators below, marked coming soon — it activates the day the API opens. Switch between AI video and AI image with one click, and until FLUX 3 goes live, the models it's benchmarked against — Seedance 2.0, HappyHorse 1.1, GPT Image 2 — are ready: same prompt box, same credits.

FLUX 3 isn't generating yet — it activates the day access opens.

What Is FLUX 3?

FLUX 3 is the multimodal foundation model announced by Black Forest Labs — the lab behind the FLUX image family — on July 23, 2026. Where FLUX 1 and FLUX.2 generated images, FLUX 3 learns from images, videos, and audio jointly inside one architecture. The bet is simple: a model that must render how objects move, collide, and sound is forced to learn how the world actually behaves — and every modality it generates benefits from that shared understanding.

That world-model foundation shows up across both creative surfaces. On the video side, FLUX 3 generates clips up to 20 seconds long with native audio in a single pass — dialogue in multiple languages, sound effects synced to physical events, styles ranging from handheld camcorder footage to animation and cinematic shots — and chains clips into multi-shot sequences that keep characters consistent. On the image side it handles complex prompts markedly better than earlier FLUX generations and renders accurate text in multiple languages. The same backbone even powers FLUX-mimic, a video-action model running robots on automotive production lines.

For now, FLUX 3 Video is in an invite-only early access phase and the image API opens in the following weeks — there is no public API yet. That's why FLUX 3 appears as coming soon in Molyin's AI video generator and AI image generator: the slot is ready on both sides, and until it activates you can create with the very models FLUX 3 is benchmarked against, on one credit system that covers them all.

What FLUX 3 Can Do

One model, two creative surfaces — and every video capability ships with audio generated natively, not stitched on afterwards.

Video With Native Audio

Every video generation includes sound — multilingual dialogue synced to lip movement and effects synced to the physical events that cause them.

Up to 20 Seconds — and Beyond

Single generations run up to 20 seconds, and clips chain into multi-shot sequences lasting minutes, with visual references keeping characters consistent across scenes.

Image & Video References

Animate a starting frame, guide a scene with reference images, or carry a character from an existing clip into a completely new setting.

Keyframe-to-Video Transitions

Define the moments that matter and let the model generate controlled transitions between them — storyboard-level control over motion.

Multilingual Dialogue & Typography

Characters speak multiple languages, and both video and image outputs render accurate on-screen text — including animated title designs.

Image Synthesis & Editing

Generates and edits images across a wide range of styles, aspect ratios, and resolutions, with markedly better complex-prompt handling than earlier FLUX generations.

How FLUX 3 Compares in Early Testing

In early human-preference evaluations of 10-second 720p text-to-video clips with audio, FLUX 3 was preferred over every major video model tested — Luma Ray 3.2, Runway Gen-4.5, Grok Imagine Video, Kling v3 Pro, Happy Horse, Seedance 2.0, and Gemini Omni Flash. The rates below are the share of comparisons where evaluators chose FLUX 3.

Compared againstFLUX 3 preferred
Luma Ray 3.293%
Runway Gen-4.577%
Grok Imagine Video69%
Kling v3 Pro60%
Happy Horse v159%
HappyHorse 1.157%
Seedance 2.052%
Gemini Omni Flash52%

Preliminary results published at announcement, before general release — expect movement as the model matures. Independent head-to-head comparisons are still emerging.

Official FLUX 3 Samples

Sample outputs published by Black Forest Labs at the FLUX 3 announcement. Turn the sound on for the videos — native audio is the headline feature.

Video — generated with native audio

Image — synthesis & editing

FLUX 3 example — blue tubular chair with red cushion, product design renderFLUX 3 example — lava waterfall pouring into the sea at nightFLUX 3 example — oil painting of a stool in a sunlit roomFLUX 3 example — flat illustration of a lighthouse beam under a crescent moonFLUX 3 example — macro photograph of an octopus eyeFLUX 3 example — black sedan drifting across an empty lot, photographic style

FLUX 3 at a Glance

The verified state of FLUX 3: what's announced, what's accessible, and who can build on it.

Announced

July 23, 2026

Unveiled alongside FLUX-mimic, a robotics model built on the same backbone and tested on automotive production lines.

FLUX 3 Video

Early access (invite-only)

Video + audio generation and editing, via application. APIs and private weight access roll out after the early-access phase.

FLUX 3 Image

Early access opens within weeks

Image synthesis and editing through APIs and private weight access; early results already beat earlier FLUX generations.

FLUX 3 Dev

Open weights announced

An open-weight multimodal backbone for content creation and action prediction is on the launch plan; timing not yet announced.

The moment Black Forest Labs opens FLUX 3 access, it activates in both generators above — same credits, same workflow, nothing to migrate.

FLUX 3 vs. the Models You Can Use Today

FLUX 3's early numbers put it ahead of the field, but it isn't publicly usable yet. Every other model in this table is live on Molyin right now — including Seedance 2.0 and HappyHorse 1.1, two of the strongest models FLUX 3 was benchmarked against.

ModelModalityAvailabilityOn Molyin
FLUX 3 (Black Forest Labs)Video + audio · ImageInvite-only early accessComing soon — already in both model pickers
Seedance 2.0VideoGenerally availableLive
HappyHorse 1.1VideoGenerally availableLive
GPT Image 2ImageGenerally availableLive
Nano Banana ProImageGenerally availableLive

Win rates referenced from evaluations published at FLUX 3's announcement, July 2026. Independent comparisons are still emerging.

FLUX 3 FAQ

What creators and developers ask about Black Forest Labs' multimodal model.

Start Creating with Molyin Today

Sign up free and turn your first idea into a cinematic video in minutes.