Black Forest Labs Releases Flux 3 for Video Generation With Native Audio Up to 20 Seconds
Key Takeaways
- •Flux 3 is a multimodal foundation model from Black Forest Labs that learns from images, video, and audio simultaneously and can generate video clips of up to 20 seconds with native synchronized audio.
- •In BFL's preliminary internal evaluations using 10-second 720p clips, Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons and Runway Gen-4.5 in 77 percent, but margins narrowed to 52 percent against stronger competitors such as Seedance 2.0 and Gemini Omni Flash.
- •BFL collaborated with Mimic Robotics to develop Flux-mimic, a video-action model for robotics applications that is already being tested on production tasks at Audi.
- •The model is built on BFL's Self-Flow approach, which uses a multimodal transformer to convert images, video, and audio into a shared internal representation for both generation and understanding tasks.
- •BFL is releasing Flux 3 in phases, with Flux 3 Video available now, Flux 3 Image scheduled for the coming weeks, and open-weight access to the backbone planned under the name Flux 3 Dev.

German AI company Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model designed to learn from images, video, and audio at the same time. The company says the model can generate video clips up to 20 seconds long with native audio, and that early internal evaluations showed it outperforming several rival video-generation systems.
BFL, founded by researchers who previously developed the Stable Diffusion models at Stability AI, earlier built its reputation on the Flux series of open-weight image generation models. Flux 3 marks the company's expansion from image generation into full multimodal video and audio, areas where competitors including Runway, Luma Labs, Google, and OpenAI have been racing to establish dominance.
BFL describes Flux 3 as a step toward "real-world visual intelligence," which it defines as models able to "perceive, predict, and act across physical and digital environments." The company frames the release as part of a wider industry effort to build so-called world models, an approach being pursued by multiple major AI labs aiming to create systems that understand physical dynamics rather than simply pattern-matching pixels.
According to BFL, no single modality fully captures reality. Images provide spatial structure, video shows how that structure changes over time, and audio can expose relationships between physical or mechanical events and the sounds they produce. The company argues that training on images, video, and audio together allows the modalities to fill gaps for one another, giving the model richer information than separate training on each data type.
Flux 3 adds native audio to generated videos
Flux 3 introduces native audio generation for BFL video models, with clips of up to 20 seconds. Synchronized audio has remained a gap for many video generation systems, which often produce silent clips or rely on separate audio models. The model supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven connections between clips for longer multi-shot sequences. BFL says Flux 3 performs especially well on human facial expressions and on matching audio to physical events.
In early evaluations using 10-second clips at 720p, BFL reported that Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent.
The reported margins were smaller against stronger competing systems. BFL said Flux 3 was preferred over Kling v3 Pro 60 percent of the time, over Happy Horse v1 in 59 percent of comparisons, over Happy Horse 1.1 in 57 percent, and over both Seedance 2.0 and Gemini Omni Flash in 52 percent each.
BFL says these results are preliminary, and no independent tests are currently available. If the model performs comparably to leading systems such as Seedance, which has already reached Hollywood, and Gemini Omni Flash, Flux 3 would be positioned among the top video models.
The company also expects Flux 3 to improve image generation, particularly for complex prompts and accurate text rendering across multiple languages. BFL plans to release Flux 3 Image in early access within the next few weeks.
Flux-mimic targets robotics applications
BFL says Flux 3 can also predict actions based on its understanding of the world. Working with Mimic Robotics, the company developed Flux-mimic, a video-action model for robotics applications that is already being tested on production tasks at Audi. The work reflects a broader industry trend of applying video and world models to embodied AI, where companies including Google and Tesla have explored similar approaches for robot control and autonomous systems.
Flux 3 is based on Self-Flow, BFL's approach for training one model to both generate and understand content. The system uses a multimodal transformer with dedicated components that convert images, video, and audio into a shared internal representation and then convert that representation back into outputs.
An additional action component provides the basis for robotics applications. BFL says this unified learning process performs better than the previously standard flow-matching method, both in generation quality and in the model's understanding of the physical world.
BFL plans staged rollout and open-weight access
BFL is releasing Flux 3 capabilities in phases, with early-access periods intended to gather feedback and support safety testing. Flux 3 Video is already available, while Flux 3 Image is scheduled to follow in the coming weeks. Action prediction will initially be available through selected partners.
The company also plans to release open-weight access to the multimodal backbone under the name "Flux 3 Dev." Open-weight releases have been a hallmark of BFL's strategy, consistent with its earlier Flux image models that were widely adopted by developers and researchers. Over the longer term, BFL says it is working on next-generation models intended to combine perception, action, and language prediction in a single model.