Black Forest Labs launches FLUX 3 for video, audio, and robotics

Black Forest Labs launches FLUX 3 for video, audio, and robotics

The unified multimodal model can generate videos with native audio, edit images, and provide the foundation for robotic action prediction.

Black Forest Labs has unveiled FLUX 3, a multimodal foundation model trained jointly on images, video, and audio to support content generation and robotic action prediction.

The model uses a unified architecture designed to learn spatial structure, movement, sound, and physical interactions together rather than treating each format as a separate task. It builds on Self Flow, the company’s method for aligning multimodal generation and understanding within the same system.

FLUX 3 Video can generate clips with native audio lasting up to 20 seconds from text, images, or existing videos. It also supports video continuation, keyframe transitions, multilingual dialogue, typography, and the chaining of clips into longer sequences.

Advertisement

In preliminary tests conducted by Black Forest Labs, FLUX 3 was preferred over Grok Imagine Video in as many as 69% of comparisons, Kling v3 Pro in 60%, Runway Gen 4.5 in 77%, and Luma Ray 3.2 in 93%.

The company cautioned that the model and evaluation system remain in development and said full benchmark results and methodology will be published with broader availability.

FLUX 3 can also generate and edit images across multiple styles, formats, and resolutions, with improvements in complex prompt following and multilingual text rendering. Early access for FLUX 3 Image is expected to open in the coming weeks.

The model’s video backbone is also being adapted for robotics through FLUX mimic, developed with mimic robotics. Audi is testing the system for production tasks involving robotic manipulation, with Black Forest Labs saying some tasks can be fine tuned using as little as 30 minutes of robot data.

FLUX 3 Video and FLUX 3 Action are available through early access. Black Forest Labs plans to release API access, private model weights, and an open weight FLUX 3 Dev version later this year.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
Black Forest Labs launches FLUX 3 for video, audio, and robotics
Black Forest Labs launches FLUX 3 for video, audio, and robotics

The unified multimodal model can generate videos with native audio, edit images, and provide the foundation for robotic action prediction.

Share

Add us on Google

Black Forest Labs has unveiled FLUX 3, a multimodal foundation model trained jointly on images, video, and audio to support content generation and robotic action prediction.

The model uses a unified architecture designed to learn spatial structure, movement, sound, and physical interactions together rather than treating each format as a separate task. It builds on Self Flow, the company’s method for aligning multimodal generation and understanding within the same system.

FLUX 3 Video can generate clips with native audio lasting up to 20 seconds from text, images, or existing videos. It also supports video continuation, keyframe transitions, multilingual dialogue, typography, and the chaining of clips into longer sequences.

Advertisement

In preliminary tests conducted by Black Forest Labs, FLUX 3 was preferred over Grok Imagine Video in as many as 69% of comparisons, Kling v3 Pro in 60%, Runway Gen 4.5 in 77%, and Luma Ray 3.2 in 93%.

The company cautioned that the model and evaluation system remain in development and said full benchmark results and methodology will be published with broader availability.

FLUX 3 can also generate and edit images across multiple styles, formats, and resolutions, with improvements in complex prompt following and multilingual text rendering. Early access for FLUX 3 Image is expected to open in the coming weeks.

The model’s video backbone is also being adapted for robotics through FLUX mimic, developed with mimic robotics. Audi is testing the system for production tasks involving robotic manipulation, with Black Forest Labs saying some tasks can be fine tuned using as little as 30 minutes of robot data.

FLUX 3 Video and FLUX 3 Action are available through early access. Black Forest Labs plans to release API access, private model weights, and an open weight FLUX 3 Dev version later this year.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.