NewsStocksBlack Forest Labs Launches FLUX 3: First Video Model Powering Robotics With Audi Partnership

Black Forest Labs Launches FLUX 3: First Video Model Powering Robotics With Audi Partnership

Author: Decrypt·

Key Takeaways

  • FLUX 3 is Black Forest Labs' first model capable of generating video clips up to 20 seconds long with synchronized audio including dialogue, sound effects, and ambient noise.
  • In preference testing, human reviewers favored FLUX 3 over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons, but evaluations did not include OpenAI's Sora or Google's Veo.
  • The same multimodal architecture underlying FLUX 3 powers FLUX-mimic, a robotics model co-developed with mimic robotics that Audi is testing on its production line for tasks such as fitting flexible door seals.
  • FLUX-mimic achieves a system response time of approximately 101 milliseconds, which BFL describes as comparable to human visual reflexes.
  • The open-weight Dev version of FLUX 3, the only tier planned for local deployment, is not expected until later in 2026, while video and action capabilities are currently limited to API and select partner access.
Black Forest Labs Launches FLUX 3: First Video Model Powering Robotics With Audi Partnership

Black Forest Labs has released FLUX 3 in early access, marking the company's first model capable of generating video rather than only still images. The new system produces clips up to 20 seconds long with synchronized audio and is built on a multimodal architecture trained jointly on images, video, and sound. The launch places BFL in one of AI's most competitive categories, where OpenAI's Sora, Google's Veo, Runway, and Luma are all racing to produce commercially viable AI-generated video.

The same underlying technology also powers FLUX-mimic, a robotics model developed in collaboration with Zurich-based mimic robotics, which German automaker Audi is already testing on its production line. The open-weight "Dev" version—the only tier planned for local deployment—is not expected until later in 2026. Video and Action capabilities remain available solely through APIs and select partner access, with image generation slated to follow "in the coming weeks," according to BFL.

FLUX 3's video functionality represents its flagship feature. The model generates clips with audio synced to on-screen events, including dialogue, sound effects, and ambient noise. In early head-to-head evaluations, human reviewers expressed a preference for FLUX 3's output over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons. The model also edged out Gemini Omni and Seedance, prevailing in 52% of those evaluations. BFL notes that these results reflect preference testing rather than a fixed scoring rubric—evaluators watch paired clips and select the more convincing one, with BFL tallying how frequently FLUX 3 wins. Notably, the evaluations did not include direct comparisons against OpenAI's Sora or Google's Veo, leaving FLUX 3's standing against those prominent systems untested in BFL's reported results.

The model also maintains strong still-image capabilities consistent with the FLUX line's heritage. BFL shared sample images demonstrating versatility across a broad range of styles beyond photorealism.

BFL positions FLUX 3 as more than a creative content tool. "A model that only learns images can only generate images," said co-founder and CEO Robin Rombach. The company's thesis holds that learning to predict video entails learning the underlying physics—weight, contact, timing—which is precisely what a machine requires to navigate the physical world. The idea of using large-scale generative models as a foundation for robotic control has drawn growing interest across the AI and robotics fields, with companies including Google DeepMind and Tesla exploring related approaches.

That thesis underpins FLUX-mimic. Developed with mimic robotics, the system adds a lightweight "decoder" to FLUX 3's video-prediction engine, translating the model's internal understanding of motion into actual robotic actions. Audi is already testing the system on tasks such as fitting flexible door seals—work that conventional automation has historically struggled to handle because deformable materials behave unpredictably and resist the rigid programming traditional robots rely on.

"Audi represents the kind of manufacturing partner we built FLUX-mimic for," said mimic co-founder Stephan-Daniel Gravert. Audi's Christoph Schneider stated that the robots now "solve complex soft-body manipulation work" that earlier machines could not address. BFL reports that the full system responds in approximately 101 milliseconds, comparable to human visual reflexes.

FLUX 3 arrives as BFL's latest bid to reclaim momentum. Founded in August 2024 by researchers who helped build the original Stable Diffusion models at Stability AI, Black Forest Labs quickly established the FLUX line as a dominant force. Its open-source Flux Dev and Schnell models captured recognition as the best open-source image generators, surpassing MidJourney and outperforming Stability AI's own Stable Diffusion 3. FLUX 1.1 Pro subsequently topped the Artificial Analysis image arena in October, though that version was not open source.

BFL released FLUX.2 in November 2025, but it failed to achieve the same popularity. The open-source crown held by the original Flux models persisted until Alibaba's Z-Image Turbo matched their quality on lower-end consumer graphics cards in late 2025. "This is what SD3 was supposed to be," a CivitAI user commented at the time.

FLUX 3 represents BFL's return to form, though it is not yet fully open. Video and Action capabilities are available now in early access through APIs and select partners, including mimic robotics. Image generation will follow in the coming weeks. The open-weight Dev version is expected later in 2026.