One Model. Four Outputs. No Stitching Required.
Black Forest Labs just launched FLUX 3, and it does something no other publicly available model does right now: generate images, video with native audio, and predict physical actions from the same architecture. Not three separate models wired together. One network, trained on all of it at once.
For anyone who has been cobbling together a text-to-image model for stills, a video generator for clips, and a separate audio tool for sound effects, this is the thing that replaces all three. FLUX 3 generates up to 20 seconds of video with synchronised dialogue, sound effects, and music from a single text prompt. It can also continue from a reference image, swap scenes while keeping a character consistent, and chain multiple clips into longer sequences.
FLUX 3 Video is available now through an early access program. FLUX 3 Image follows in the coming weeks. An open-weight FLUX 3 Dev backbone is confirmed for later this year.
What FLUX 3 Actually Generates
The headline capability is video. Here is what that includes.
Text-to-video with native audio. Type a prompt, get a video clip up to 20 seconds long with dialogue, ambient sound, and music baked in. No separate audio pass. No lip-syncing after the fact. The model generates picture and sound together because it learned both at the same time.
Image-to-video. Feed it a starting frame and it animates from there. You can also provide reference images that the model uses as visual guides without making the output a direct copy of the input.
Video-to-video. Provide a source clip and a new scene description, and FLUX 3 carries central elements like character appearance into a different context. Same person, new environment, new action.
Keyframe control. Define specific moments you want the video to hit, and FLUX 3 generates the transitions between them. Useful for storyboarding where you know the start and end but need the middle filled in.
Multi-shot chaining. Individual clips can be linked into longer sequences with visual consistency across scenes. Characters maintain their appearance. Settings stay coherent.
Style range. FLUX 3 is not limited to cinematic. Early outputs show candid camcorder footage, animation, stylised design sequences, and product showcases. The model handles typography and animated design elements alongside footage-style generation.
Multilingual dialogue. The audio generation includes spoken dialogue in multiple languages, synced to the video it generates.
How It Compares
Black Forest Labs published preference testing results. In 10-second, 720p text-to-video clips, FLUX 3 was preferred over:
- Luma Ray 3.2 in 93% of comparisons
- Runway Gen-4.5 in 77%
- Grok Imagine Video in 69%
- Kling v3 Pro in 60%
- Happy Horse v1 in 59%
- Happy Horse 1.1 in 57%
- Seedance 2.0 and Gemini Omni Flash in 52% each
Those are early numbers from a pre-release checkpoint, not the shipping model, so treat them as directional rather than final. The 52% figure against Gemini Omni Flash is essentially a tie, and Seedance 2.0 is not widely available outside China after major studios threatened legal action over copyright concerns.
The more interesting comparison is structural. Most video generators produce clips you then have to add sound to in post. FLUX 3 generates both simultaneously. That removes a whole production step and the quality gap that comes from bolting audio onto video after the fact.
What to Look Out For
Pricing and access are not published yet. FLUX 3 Video is in early access via application. There is no public API, no per-generation pricing, and no service-level commitments. You apply, BFL approves you, and then you use it. That limits who can start building with it right now.
The open-weight version is coming but not here yet. FLUX has always shipped a Dev tier with downloadable weights. FLUX 3 Dev is confirmed for later this year, but timing and licence terms have not been announced. For developers who run models locally, this is the version that matters most, and it is not available yet.
Canva, Burda, Magnific, Krea, and Picsart are already testing it. These companies are BFL’s launch partners for creative tooling. If you use any of those platforms, FLUX 3 capabilities will likely surface there before they are available as a standalone API.
The robotics angle is real but early. BFL partnered with mimic robotics to build FLUX-mimic, a video-action model trained on the FLUX 3 backbone for dexterous manipulation tasks. It is being tested at Audi. For creative professionals, this does not matter today. For the trajectory of where multimodal models are going, it confirms that the same architecture that generates your video can also learn physical tasks, which is a signal worth watching.
What This Unlocks
For anyone making content, the immediate opportunity is simple: one prompt produces video with sound. No more generating a clip, realising the audio is wrong, regenerating in a separate tool, and spending hours syncing them. The model handles both because it learned they belong together.
For product brands, the video-to-video capability means you can shoot a single product clip and generate variations that place the same product in different scenes, different lighting, different contexts, while keeping visual consistency. That is the kind of thing that used to require a full reshoot.
For designers working in Canva or Picsart, FLUX 3 is likely to show up as a built-in generation option within those tools. When it does, you will have video with audio accessible inside your existing workflow without switching to a separate platform.
The bigger shift is architectural. When one model learns images, video, audio, and physical prediction together, it develops a richer understanding of how the world works than any single-modality model can. BFL’s own framing is blunt: “You cannot cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.”
FLUX 3 is the first model from a major lab to ship on that principle. Whether the output quality fully delivers on the promise is something early-access testers will confirm over the coming weeks. But the direction is clear: generation is moving from separate tools for separate outputs to one model that understands all of them at once.
Want more like this? Join MOKU Club for free. Weekly resources, early access to new guides, and occasional templates you can actually use. Join below.



