Skip to main content

Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action

Black Forest Labs introduces FLUX 3, a multimodal foundational model integrating image, video, audio, and robot action prediction with strong human preference results.

AI-written
Inewgen
27 Jul 2026Source: MarkTechPost3 min read (0 views)Last updated 03 Aug 2026
Share
Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action

Stock photo for illustration only, not from the actual event

Font size
  • Black Forest Labs launches FLUX 3, simultaneously learning across image, video, audio, and robot actions.
  • Built on the Self-Flow architecture to align generative creation and comprehension.
  • Generates video clips up to 20 seconds long in a single pass with native audio.
  • Outperformed competitors like Luma Ray 3.2 and Runway Gen-4.5 in preliminary human evaluations.

The research team at Black Forest Labs (BFL) argues that no single modality provides a complete representation of reality. Images capture spatial structure instantly, video restores time and physical dynamics, and audio reveals causal relationships. Each is viewed as a lossy projection of the same underlying reality. Training on all of them simultaneously ensures the modalities constrain one another, such as matching sound to impact and motion to mass, making FLUX 3 the first model built entirely on this principle.

FLUX 3 builds upon Self-Flow, BFL's architecture for aligning multimodal generation and understanding. The approach combines the flow matching objective with a self-supervised feature reconstruction objective. The reference implementation on GitHub is available under the Apache-2.0 license and utilizes SiT-XL/2 with per-token timestep conditioning.

93%Preferred over Luma Ray 3.2
77%Preferred over Runway Gen-4.5
20sMax video generation length

FLUX 3 Video generates clips up to 20 seconds long in a single generation complete with native audio. Supported modes include text-to-video, image-to-video, reference-based video-to-video, keyframe-to-video, and generative audio-video continuation. BFL also highlights multilingual dialogue, agentic clip chaining, strong typography generation with animated designs, and exceptional capabilities in rendering human facial expressions and associating sounds with physical events.

"video is the expensive modality because rendering the world correctly forces the model to learn contact, motion, weight and cause. Audio and actions are low-dimensional signals that attach to a world model already paid for."

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

Black Forest Labs

Training an AI model across multiple foundational modalities simultaneously allows the system to build a robust internal physics engine of the real world. Because rendering video requires mastering contact, weight, motion, and causality, it serves as an expensive foundational anchor. Once the model grasps video dynamics, lower-dimensional signals like audio and physical actions can be seamlessly integrated into an already established world model.

multimodal ai video generation technology

Stock photo for illustration only, not from the actual event

In preliminary human preference evaluations using 10-second 720p text-to-video clips with audio, FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77% of evaluations. Furthermore, preview builds tested by Mimic Robotics on physical robots demonstrated robust performance in structured tasks like kitting and assembly, with Audi currently testing and deploying the system on production lines.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article