FLUX 3 unifies video, audio, and action in one architecture — but the launch is incomplete
Black Forest Labs released FLUX 3 today, a multimodal model the company trains jointly across images, video clips up to 20 seconds, audio, and robotic action prediction. The architecture, called Self-Flow, is the same foundation for all four modalities. FLUX 3 is BFL's first public video generation