Black Forest Labs Launches FLUX 3: A Multimodal Frontier Model for Images, Video, Audio, and Robotic Action
Key Takeaways
- •FLUX 3 is jointly trained across image, video, audio, and robotics action prediction within a single architecture using Black Forest Labs' Self-Flow method, rather than combining separate specialized models.
- •FLUX 3 Video can generate clips up to 20 seconds with synchronized native audio from a single prompt, matching the duration previously achieved by OpenAI's discontinued Sora model.
- •Preliminary preference benchmarks show FLUX 3 outperforming Luma Ray 3.2 and Runway Gen-4.5, but tying Google's Gemini Omni Flash at 52% on 10-second text-to-video quality comparisons.
- •Through the FLUX-mimic partnership with Mimic Robotics, the model can be fine-tuned for new robotic manipulation tasks with as little as 30 minutes of demonstration data, a 60x reduction from the 30-plus hours previously required.
- •Black Forest Labs has not disclosed pricing, service-level commitments, evaluation methodology, or open-weight licensing terms at launch, leaving enterprise buyers unable to assess total cost of ownership.

Black Forest Labs (BFL), the Freiburg, Germany-based AI lab, has launched FLUX 3, a multimodal frontier model designed to understand and generate images, combined audio/video clips of up to 20 seconds from a single text prompt, and physical actions for robotics. The company describes FLUX 3 as jointly trained across all these modalities within a single architecture, rather than assembling separate image, video, and audio models behind a common interface.
That architectural distinction is central to BFL's positioning. The company wants enterprises to treat creative generation, simulation, computer use, and robotics as interconnected applications of what it calls "visual intelligence" — models, in BFL's words, "that can perceive, predict, and act across physical and digital environments." The launch also marks BFL's first publicly available video generation model and positions the company, at roughly 100 employees, directly against far larger U.S. platforms including Google's Gemini Omni and a field of specialized video startups that have collectively raised hundreds of millions in the past two years.
Product Lines and Availability
FLUX 3 will be delivered through four product lines:
- FLUX 3 Video — with optional native audio generation
- FLUX 3 Image — rolling out in the coming weeks
- FLUX 3 Action — for physical AI and robotics
- FLUX 3 Dev — an upcoming open-weight release
FLUX 3 Video and FLUX 3 Action are now entering a gated Early Access program that anyone can apply for, but BFL must approve each applicant. There is no public API access yet through BFL or its partners. The company says FLUX 3 Image will roll out in the coming weeks, followed by broader general availability.
The phased release strategy mirrors recent approaches taken by other U.S. frontier labs including Anthropic and OpenAI, though those restrictions were attributed to security concerns and government requests. For BFL, the gating appears driven by infrastructure readiness and ongoing development rather than safety classifications.
What BFL Has Not Announced
Several critical details are absent from the launch. BFL has not disclosed pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts, or any image-model benchmarks. Enterprise buyers cannot yet calculate total cost of ownership or independently reproduce the video comparisons the company has published.
FLUX 3 is also not launching with downloadable weights or an open source license. BFL says faster and open-weight versions will arrive later this year. Its technical blog names FLUX 3 Dev as "open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction" — a broader commitment than any previous FLUX Dev release, all of which covered images only. However, FLUX 3 Dev arrives last in the release sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside or shortly after a major model announcement will have to wait.
The delay is notable because BFL's open-weight releases have been a primary adoption driver. Locally deployable FLUX models are integrated into ComfyUI and Hugging Face Diffusers workflows, and the absence of a downloadable FLUX 3 variant at launch means that developer ecosystem cannot yet test the new architecture on their own hardware.
Preliminary Benchmark Results
BFL has published several benchmark comparisons, all qualified as preliminary. Full results and methodology are promised for the broader general availability phase.
In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, BFL reports FLUX 3 was preferred over:
- Luma Ray 3.2 in 93% of comparisons
- Runway Gen-4.5 in 77%
- Grok Imagine Video in 69%
- Kling v3 Pro in 60%
- Happy Horse v1 in 59%
- Happy Horse 1.1 in 57%
- Seedance 2.0 in 52%
- Google's Gemini Omni Flash in 52%
A significant caveat accompanies these figures. The chart carrying the results is labeled a "preliminary evaluation of an early FLUX 3 candidate," meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. The shipping model could perform better or worse; nothing published today measures what customers will actually use.
Luma Ray 3.2 and Runway Gen-4.5, against which FLUX 3 posted its strongest results, are established products but not the models currently leading independent video rankings. Seedance 2.0, tied at 52%, is a product most Western enterprises cannot currently procure — ByteDance indefinitely postponed its international rollout after Netflix, Warner Bros., Disney, Paramount, and Sony issued legal threats over alleged systematic copyright infringement, and that suspension remains in place.
Gemini Omni Flash, also at 52%, is the more consequential comparison. Omni represents the closest large-platform analogue to FLUX 3's multimodal approach, and by BFL's own measurement, the two are indistinguishable on 10-second text-to-video quality. Google's advantage is availability: Omni is accessible via the Gemini API at $0.10 per second of generated 720p video, or roughly $1.00 for a 10-second clip. However, editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland, and the United Kingdom — though editing video the model itself generated is permitted. A European enterprise wanting to process existing footage through a generative editing pass cannot currently do so on Omni Flash. That geographic gap is an opening BFL, headquartered in Germany, is positioned to exploit if FLUX 3 ships without similar regional restrictions.
Enterprise Video Model Comparison
| Model | Max Generation Duration | Max Resolution | Key Constraints | Price per 10s Clip (720p) | Price per 10s Clip (1080p) | Price per 10s Clip (4K) |
|---|---|---|---|---|---|---|
| FLUX 3 Video | 20 seconds | Not stated; evaluations at 720p | Early access; no published SLA or pricing | Not announced | Not announced | Not announced |
| HappyHorse 1.1 | 15 seconds | 1080p | No 4K; closed weights | Not published (v1.0 reseller rate ~$1.82) | Not published (v1.0 reseller rate ~$3.12) | n/a |
| Veo 3.1 | Per-second billing | 4K | Supports clip extension; preview | $4.00 | $4.00 | $6.00 |
| Veo 3.1 Fast | Per-second billing | 4K | Preview | $1.00 | $1.20 | $3.00 |
| Veo 3.1 Lite | Per-second billing | 1080p | No 4K, no clip extension; preview | $0.50 | $0.80 | n/a |
| Gemini Omni Flash | 10 seconds (3s minimum) | 720p at 24 FPS | Preview; no EU access for uploaded video editing | $1.00 | n/a | n/a |
Unified Architecture: Self-Flow
FLUX 3 builds on Self-Flow, BFL's method for aligning multimodal understanding and generation within a single architecture, first publicized in March 2026. The company says it significantly scaled up compute and data to train across video, images, and audio simultaneously, and that testing demonstrated video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.
"We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture," said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. "True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express."
Rombach added elsewhere in the announcement: "You can't cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds."
BFL says FLUX 3 targets creative tooling, media, design, e-commerce, and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation, and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea, and Picsart.
For creative software companies, the appeal is consolidation — a single foundation model that could support storyboarding, image editing, product rendering, video variation, and localization without repeatedly translating assets between disconnected models. For robotics teams, the potential value lies in data efficiency: models that already encode motion, object behavior, and physical change may require less task-specific robot training than systems starting from raw demonstrations. This convergence of generative media and robotics foundation models reflects a broader industry shift, as companies including Google DeepMind and multiple robotics startups have begun applying large pretrained models to manipulation and locomotion tasks.
FLUX 3 Video Capabilities
The video tier is the most concretely specified part of the launch. FLUX 3 generates clips of up to 20 seconds with native audio from a single prompt. Every video output includes synchronized audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio. BFL has not stated what resolution its 20-second clips run at, and published evaluations were conducted at 720p. A 20-second single-prompt generation is among the longest yet achieved, matching OpenAI's discontinued Sora model.
The published capability list includes:
- Text-to-video generation
- Image-to-video generation, animating from a starting frame or using images as visual references
- Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context
- Generative video-audio continuation from existing video and audio input
- Keyframe-to-video generation for controlled transitions between defined moments
- Multilingual dialogue
- A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics
- Typography generation and animated design
- Agentic chaining of individual clips into longer, multi-shot sequences
The last item — agentic chaining — is the capability enterprise video teams should examine most closely. BFL claims the combined features can produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it would address the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.
Competition in character consistency is intense. HappyHorse 1.1's headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to maintain identity stability across generated footage. Alibaba claims zero-drift lip sync and has specifically targeted artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening.
BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, preliminary midtraining evaluations show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. No image benchmarks or win rates have been published.
FLUX-mimic: From Video Models to Robot Models
BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.
The technical blog describes two routes to action prediction: integrating native action prediction directly into FLUX 3 by scaling up the initial Self-Flow work, and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic takes the second route, combining the FLUX 3 backbone with Mimic Robotics' expertise in robot learning and production deployment for dexterous manipulation.
FLUX-mimic is designed for general-purpose robotic manipulation, helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with minimal task-specific data. BFL and Mimic Robotics say the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.
"The hardest part of robotics is data," said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. "Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning."
The 60x reduction in required demonstration data — from 30 hours to 30 minutes — targets what has been one of the most persistent bottlenecks in applied robotics, where each new manipulation task typically demands extensive physical data collection that does not transfer across tasks or environments.
The World-Model Claim
BFL argues that a model trained only on images cannot understand a world that "moves, sounds, changes, and responds," and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni. Its developer documentation cites "world knowledge" combining "an understanding of physics" with Gemini's grasp of history, science, and cultural context. Google's marketing states: "Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different," crediting the model with "an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic."
The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it. There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of this indirectly; nothing else on offer captures it at all.
Open Weights and Company Background
BFL officially launched in summer 2024 and built its reputation on open sourcing high-quality AI image models widely adopted by developers, creatives, and enterprises. The company's founders — Rombach, Andreas Blattmann, and Patrick Esser — previously helped create VQGAN, latent diffusion, and Stable Diffusion, the open source technology that catalyzed broad AI generation capabilities for consumers and is currently used by many AI image generators and companies.
That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart, and Nous Research's Hermes Agent, among other platforms. The company cites film director Martin Scorsese among professional users. Wired magazine described Black Forest Labs as a relatively small company that became a leading competitor to Silicon Valley's largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on Hugging Face. BFL says it now runs a 100-person team across Freiburg and San Francisco.
FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev, and related control models — released shortly after the firm's launch — gave researchers and creative-tool developers access to downloadable checkpoints, local inference, and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.
The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. BFL called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code, and optimized implementations for consumer Nvidia GPUs.
FLUX 3 Dev raises the stakes. Previous Dev releases were image models. FLUX 3 Dev is described as a multimodal backbone spanning video, audio, image, and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL has not yet shared information about its license, parameter count, quantizations, or hardware requirements. The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and allow teams to adapt FLUX 3 to their own data, products, and workflows. Whether the open-weight license will permit commercial robotics use — where physical safety and liability considerations differ from image generation — remains an open question the company has not yet addressed.
Financial Backing
Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva, and Deutsche Telekom's T.Capital. The investor roster spans both capital providers and commercial partners — Adobe, Figma, and Canva are also integration partners — which aligns BFL's funding with distribution channels for FLUX 3's creative and media capabilities.