Skip to content
Source
Animesh Garg· @animesh_garg · X·· 11 days agoSignalEditorial score85

FLUX 3 Action by @bfl_ai is super cool!

Summary

FLUX 3 Action is an open-source 7B parameter world action model that achieves first place on the RoboLab benchmark. It outperforms previous models by 6.1 percentage points with 56% fewer parameters and runs 3.95x faster. The model supports fine-tuning for specific robots and tasks and is integrated into LeRobot with deployment on NVIDIA Jetson.

Editorial context

FLUX 3 Action is a 7B parameter world action model that achieves state-of-the-art performance on the RoboLab benchmark, outperforming previous models with fewer parameters and faster inference. It integrates action prediction with video generation and supports fine-tuning for specific robots and tasks.

What the source reports

Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.

What remains unknown

Not established in the collected evidence: Environment, Control, Data origin.

Reported performance applies to the described task. It does not establish general autonomy or deployment readiness.

Source excerpts and review record

Automatically extracted; no manual editorial approval recorded.

Beats Cosmos 3 Nano (16B) on RoboLab-120 (42.9% vs. 36.8%)

Open source E1

It outperforms the previous best open model by 6.1 percentage points

Open source E2

pairing F3A with high-level reasoners cuts execution costs by 29% and runtimes by 40%

Open source E3

Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task

Open source E4

We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples

Open source E5

Implications for data suppliers

RoboSignal interpretation and collection questions, not statements of buyer demand.

  • Confirm the required data type and collection setting with the buyer; this source does not establish a complete collection specification.
  • Validate demand and acceptance criteria with a buyer before scaling. Publication, popularity and a research result do not establish a purchase commitment.
Text

FLUX 3 Action by @bfl_ai is super cool!
Everyone is pivoting into a robotics neolab

F3A is SOTA without the WAM latency: Beats Cosmos 3 Nano (16B) on RoboLab-120 (42.9% vs. 36.8%) while running at less than half the parameter count (7B) and up to 3.95x faster.

Flow matching tuned for physical control: Uses separate noise-to-signal transitions for video vs. actions and asymmetric classifier-free guidance, completely avoiding the need to decouple action heads from video synthesis.

And notably, pairing F3A with high-level reasoners cuts execution costs by 29% and runtimes by 40% over pure reasoning. This lets the fast policy handle motor primitives and calling the LLM only to correct course.

Quoted postBlack Forest Labs@bfl_ai
Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).
View the quoted post on X

Source:Animesh Garg · x.com