
Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs
Quick Answer
Black Forest Labs has launched Flux 3, a multimodal model capable of generating 20-second videos with native audio, outperforming competitors like Luma Ray 3.2 in 93% of comparisons.
Quick Take
This model enhances video generation by integrating images, audio, and video, paving the way for advanced visual intelligence applications.
Key Points
- Flux 3 generates videos with native audio for the first time, up to 20 seconds long.
- In early tests, Flux 3 was preferred over Luma Ray 3.2 in 93% of comparisons.
- The model uses a multimodal transformer for better integration of images, video, and audio.
- BFL plans to release Flux 3 Image and open-weight access in the coming weeks.
- Flux 3's action prediction capabilities are being tested in collaboration with Mimic Robotics.
Source Excerpt
Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can generate video with native sound for the first time. BFL's own tests put it just ahead of market leader Seedance 2. 0, though independent results aren't yet available. The company ultimately wants to build a world model and is already testing Flux 3 on robotics tasks.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

