Black Forest Labs (BFL) has officially announced FLUX 3, a next-generation multimodal foundation model capable of generating up to 20-second videos with native audio from a single prompt. This breakthrough marks the first time the startup behind Stable Diffusion has expanded its FLUX model line from static images into video, audio, and robotic control. Currently, FLUX 3 is only available through a limited Early Access program before a wider rollout in the coming weeks.
Key Developments
The launch of FLUX 3 comes amid fierce competition in the multimodal AI space among major research labs. According to VentureBeat, FLUX 3 will be offered across four main product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and the highly anticipated open-source version, FLUX 3 Dev. In this release, FLUX 3 Video and FLUX 3 Action are entering a gated Early Access program, while the image-generation model FLUX 3 Image is expected to launch in the coming weeks. However, developers who prefer running model weights locally will have to wait patiently, as FLUX 3 Dev is scheduled for release late this year.
Technical Analysis & Technology
At the core of FLUX 3's capabilities is joint training on a unified architecture, rather than stitching together separate image, video, and audio models. BFL built this model using Self-Flow—a multimodal alignment technique published by the company in March 2026—combined with massive upgrades to computing resources and training data. As a result, the model can generate 20-second video clips at 720p resolution perfectly synchronized with natural audio. Furthermore, this architecture extends to computer vision and robotic action tasks through FLUX-mimic, an experimental version developed in partnership with Switzerland-based Mimic Robotics.
Expert Opinions & Insights
BFL has released preliminary test results showing that users prefer FLUX 3 over competitors like Runway Gen-4.5 (77%), Luma Ray 3.2 (93%), and on par with Google's Gemini Omni Flash (52%). However, enterprise analysts remain cautious, as BFL has yet to publish detailed pricing or independent benchmarks. Robin Rombach, co-founder and CEO of BFL, emphasized the company's philosophy: 'A model trained only on images can only create images. But the world is not made of static frames. It moves, makes sound, changes, and responds.'
Impact & Future Outlook
The arrival of FLUX 3 with unified 'visual intelligence' promises to reshape how enterprises produce digital content and develop robotics. For the technology community, the optimization capabilities of FLUX-mimic—which enables a robot to learn a new pick-and-place task with just 30 minutes of data instead of the traditional 30 hours—will be an invaluable tool to accelerate automation testing. Once the fully open-source FLUX 3 Dev launches later this year, it is bound to serve as a launchpad for many groundbreaking AI applications.