Black Forest Labs Releases FLUX 3 Action, an Open Robotics World-Action Model
- Source
- Black Forest Labs
- Date

Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical…
- Black Forest Labs says FLUX 3 Action, an open-weights 7B world action model, ranks first on RoboLab, beating the previous best open model by 6.1 points while using 56% fewer parameters and running up to 3.95x faster.
- It still predicts video and actions together, and BFL says it plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA.
- The model reuses FLUX 3's image, video and audio in a smaller architecture built for deployment, with a midtraining stage that teaches it to predict actions and future frames jointly.
- Teams can it on their own demonstrations. It is integrated into Hugging Face LeRobot with NVIDIA, ships with fine-tuning recipes, and supports edge deployment on NVIDIA Jetson.
- benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
- fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
- pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
Gives roboticists and builders an open, fine-tunable model that jointly predicts video and actions with a stronger performance-per-compute tradeoff than existing open VLAs, plus a path to edge deployment.
Checking sign-in…
Loading comments…


