Black Forest Labs has unveiled FLUX 3, its first video-generation artificial intelligence (AI) model that creates video and audio at the same time. The company also set out a blueprint to use the model as a foundational technology for robot control, going beyond a simple video-generation tool.
On July 23 local time, blockchain media outlet Decrypt reported that FLUX 3 is a multimodal AI model that can generate up to 20 seconds of video, along with dialogue, sound effects and background music suited to the scene. It is the first case in the FLUX series, which had focused on still-image generation, to expand into video and audio and further into robotics.
FLUX 3 adopts a structure that trains images, video and audio within a single common system. That allows it to naturally generate speech, sound effects and background music to match the on-screen scene.
For now, it is offered in an early-access form. Video generation and action features are being provided first via an API and to selected partners, and image-generation features are to be added in the coming weeks. A public-weights version that can be used locally will be released only for the FLUX Dev model in the second half of 2026.
Black Forest Labs also stressed competitiveness on performance. In an evaluation where human raters compared two outputs and chose the better video, FLUX 3 recorded preference rates of 77 percent versus Runway Gen-4.5 and 93 percent versus Luma Ray 3.2. The company said it also had an advantage in 52 percent of evaluations when compared with Google Gemini Omni and Seedance. It added the results were not a quantitative benchmark but a comparison based on human preference.
Black Forest Labs also unveiled a strategy to use the model as a foundation for "Physical AI", beyond a content creation tool. Co-founder and Chief Executive Robin Rombach (로빈 롬바흐) said, "A model trained only on images can only generate images." He said, "The process of predicting video is the process of learning principles of the real physical world such as weight, contact and timing."
The technology led to the robot AI model FLUX-mimic. Co-developed with Zurich-based startup Mimic Robotics, FLUX-mimic combines a lightweight decoder with FLUX 3's video-prediction engine to convert movement in video into actual robot motions.
Audi is already testing the technology on a production line. The target is a door seal installation process, a soft-material assembly task that existing automated equipment has struggled to handle. Mimic Robotics co-founder Stefan-Daniel Graber (스테판-다니엘 그라베르트) said, "Audi is a representative case of a manufacturing site where FLUX-mimic will be applied." Audi's Christoph Schneider (크리스토프 슈나이더) also said robots can now perform complex soft-object manipulation tasks that have been difficult for existing automation equipment. Black Forest Labs said overall system response time is about 101 milliseconds, close to human visual reflex speed.
The announcement is also seen as a signal of a shift in Black Forest Labs' business strategy. The company is a startup founded in 2024 by early Stable Diffusion developers at Stability AI, and drew attention as it competed with Midjourney using early FLUX models. In particular, the openly available FLUX Dev and FLUX Schnell models received high evaluations in the open-source image-generation market, but FLUX.2, released later, did not show the same influence as before.
Black Forest Labs is seeking to secure new growth drivers by expanding a business centered on image generation into video generation and robotics through FLUX 3.
Key functions are still being run on a limited basis. Video generation and action features are being offered only through an API and to selected partners, and the public-weights model is also due to be released in the second half of 2026, not this year. That is expected to leave the pace of FLUX 3's market spread dependent on broader API availability and whether a partner ecosystem is built.