News

Black Forest Labs’ FLUX3 multi-modal model debuts: a single generation of 20-second audio and video, with a winning rate that beats Grok and Seedance. Black Forest Labs announced the launch of the FLUX3 multi-modal basic model in Early Access mode, using a unified architecture to jointly learn images, videos and audio

2 min read
Black Forest Labs’ FLUX3 multi-modal model debuts: a single generation of 20-second audio and video, with a winning rate that beats Grok and Seedance. Black Forest Labs announced the launch of the FLUX3 multi-modal basic model in Early Access mode, using a unified architecture to jointly learn images, videos and audios. This model is extended based on the Self-Flow (self-supervised flow matching) learning framework, and simultaneously trains videos, images and audios. Based on the FLUX. 1 and FLUX. 2 series, it is further extended to multi-modal generation and understanding tasks. Comprehensive performance, native audio synchronous output FLUX3 can output up to 20 seconds of video with native audio in a single generation, and supports multiple creative modes such as Vincent video, Tuxed video, Video-generated video, input video and audio continuation, key frame to video, multi-language dialogue, and multi-shot concatenation. In the manual evaluation of a 10-second 720p video with sound, FLUX3 has a winning rate of 69% compared to Grok Imagine Video, and a winning rate of 52% compared to Seedance 2. 0 and Gemini Omni Flash, showing a comprehensive lead. Comprehensive image capabilities, entering into robot behavior prediction. In terms of image capabilities, officials say FLUX3 can complete image generation and editing, covering a variety of styles, aspect ratios and resolutions. What is more noteworthy is that Black Forest Labs is collaborating with Mimic Robotics to study the use of FLUX3 as a robot behavior prediction model, extending multi-modal AI capabilities from content generation to the field of physical robots. This layout means that FLUX3 is not only a creative tool, but also expected to become a universal multi-modal engine that connects the digital world and the physical world. via AI News (author: AI Base)