Black Forest Labs Releases Open-Weight FLUX 3 Action
The 7B model predicts video and robot actions together, placing a visual-generation lab inside the race to build open physical-AI systems.
Black Forest Labs has released FLUX 3 Action, a 7-billion-parameter open-weight world-action model for robotics. The system takes camera frames, robot state, and a textual instruction, then predicts both future visual frames and the next sequence of physical actions.
From visual generation to robot control
The release extends the FLUX family beyond image and video generation. BFL’s approach treats visual prediction as part of a robot-control problem: a model that can represent how scenes evolve may also be able to infer what a robot should do next.
On the RoboLab-120 benchmark, the released system reportedly achieved a 42.92% task-success rate, ahead of the 36.8% score cited for NVIDIA’s Cosmos 3 Nano. Those results are useful but narrow. They measure a defined set of simulated or laboratory tasks, not broad reliability in homes, factories, or unstructured environments.
The model is open-weight, but access does not mean unrestricted commercial deployment. BFL’s license permits non-commercial and limited qualifying commercial use while excluding military applications, surveillance, biometric processing, and other prohibited uses. The DROID policy reportedly needs about 32GB of GPU memory in BF16, although quantized configurations can reduce the hardware requirement.
Why the release matters
FLUX 3 Action is strategically important because it connects two markets that have usually developed separately. The same foundation-model lineage can support creative generation and physical action prediction, while researchers receive weights, policy variants, and recipes they can inspect and adapt.
The constraints are equally important. Hardware requirements remain substantial, the license limits some real-world deployments, and benchmark leadership does not establish robustness under changing objects, lighting, contact forces, or safety constraints.
Still, the release lowers the barrier for researchers to experiment with visual world models as robot policies. It also signals that major generative-media labs increasingly view physical AI as a natural extension of multimodal modeling rather than a separate discipline.