FLUX 3 Action by @bfl_ai is super cool!
FLUX 3 Action is an open-source 7B parameter world action model that achieves first place on the RoboLab benchmark. It outperforms previous models by 6.1 percentage points with 56% fewer parameters and runs 3.95x faster. The model supports fine-tuning for specific robots and tasks and is integrated into LeRobot with deployment on NVIDIA Jetson.
FLUX 3 Action is a 7B parameter world action model that achieves state-of-the-art performance on the RoboLab benchmark, outperforming previous models with fewer parameters and faster inference. It integrates action prediction with video generation and supports fine-tuning for specific robots and tasks.
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
- Environment:
- Not reported
- Control:
- Not reported
- Data origin:
- Not reported
| Metric | Value / unit | Basis / context | Evidence |
|---|---|---|---|
| percent | 42.9 percent | Basis not reported Source wording: “42.9% vs. 36.8%” | Source E1 |
| percent | 6.1 percent | Basis not reported Source wording: “6.1 percentage points” | Source E2 |
| percent | 29 percent | Basis not reported Source wording: “29% and runtimes by 40%” | Source E3 |
| percent | 40 percent | Basis not reported Source wording: “29% and runtimes by 40%” | Source E3 |
| robots | 1 robots | Basis not reported Source wording: “for a particular robot and task” | Source E4 |
- dataset: Available Artifact linkSource E5
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “Beats Cosmos 3 Nano (16B) on RoboLab-120 (42.9% vs. 36.8%)”
Source E1
Original source quotation: “It outperforms the previous best open model by 6.1 percentage points”
Source E2
Original source quotation: “pairing F3A with high-level reasoners cuts execution costs by 29% and runtimes by 40%”
Source E3
Original source quotation: “Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task”
Source E4
Original source quotation: “We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples”
Source E5
FLUX 3 Action by @bfl_ai is super cool!
Everyone is pivoting into a robotics neolab
F3A is SOTA without the WAM latency: Beats Cosmos 3 Nano (16B) on RoboLab-120 (42.9% vs. 36.8%) while running at less than half the parameter count (7B) and up to 3.95x faster.
Flow matching tuned for physical control: Uses separate noise-to-signal transitions for video vs. actions and asymmetric classifier-free guidance, completely avoiding the need to decouple action heads from video synthesis.
And notably, pairing F3A with high-level reasoners cuts execution costs by 29% and runtimes by 40% over pure reasoning. This lets the fast policy handle motor primitives and calling the LLM only to correct course.
Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).View the quoted post on X
Source:Animesh Garg · x.com