Black Forest Labs Releases FLUX 3 Action 7B Model
Black Forest Labs launches FLUX 3 Action, a 7B open-weights world action model scoring 42.92% on RoboLab-120 with high processing speed.

Stock photo for illustration only, not from the actual event
- Black Forest Labs has released FLUX 3 Action, a 7B open-weights robot policy model.
- It achieves a top leaderboard score of 42.92% on the RoboLab-120 benchmark.
- The model delivers significantly faster inference speeds compared to existing alternatives.
- It supports flexible edge deployment and custom fine-tuning for robotic research teams.
Black Forest Labs (BFL) has officially introduced FLUX 3 Action, a 7-billion-parameter open-weights robot action model designed to bridge the gap between high-accuracy world models and fast vision-language-action (VLA) architectures.
Derived from the multimodal FLUX 3 backbone, the model's pretraining utilized a diverse mix of image, video, and audio data, with video comprising over 95% of training tokens. During midtraining, pretraining data was combined with action-aligned video covering various robotic embodiments and human hand interactions.
On the RoboLab-120 benchmark featuring 120 tabletop tasks in Isaac Sim, FLUX 3 Action achieved a guidance-distilled FP8 multi-seed mean score of 42.92%. This represents a 6.1 percentage point lead over NVIDIA's Cosmos 3 Nano while using 56% fewer parameters.

Stock photo for illustration only, not from the actual event
Real-world hardware evaluations conducted by Positronic Robotics on a Franka robot arm further demonstrated its capability. FLUX 3 Action successfully completed 28 out of 30 blind evaluation attempts, achieving a 93.3% success rate that outpaced rival models.
"FLUX 3 Action keeps joint video and action prediction. BFL closes the speed gap with a smaller backbone and distillation instead."
Black Forest Labs
World Action Models have traditionally struggled with heavy computational overhead when predicting future video frames. BFL's strategy of combining a streamlined backbone with advanced distillation techniques effectively overcomes this latency barrier, paving the way for real-time robotic deployment on standard workstation and datacenter GPUs.
Regarding deployment requirements, the DROID policy requires approximately 32 GB of GPU memory in BF16 on an H200, or fits within 24 GB cards utilizing FP8 quantization and text encoder offload under the FLUX Kommunity License for non-commercial use.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment