One Model to Sense the World: Why Black Forest Labs Trained FLUX 3 on Everything at Once

Black Forest Labs has released FLUX 3, a multimodal foundation model that abandons the industry convention of training separate models for different media types. Instead, it learns from images, video, audio, and robot action sequences within a single unified architecture, built on the premise that no single modality can fully describe how the physical world behaves.

The German company, which made its name with the open-weight FLUX.1 image generation models, is making a significant strategic pivot. FLUX 3 is not a bigger image generator, it is a bet that the same underlying representation that allows a model to render a convincing wave can also help a robot predict how to install a car door.

The technical foundation is Self-Flow, a training method the lab introduced earlier this year that combines a flow-matching objective with self-supervised feature reconstruction. FLUX 3 scales this approach across multiple modalities simultaneously, and the company reports preliminary human preference results that show the model outperforming competitors across the board. In head-to-head comparisons of 10-second text-to-video clips at 720p with audio, reviewers preferred FLUX 3 over Luma Ray 3.2 in 93 percent of cases, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent.

The headline capability is video generation up to 20 seconds in a single pass with native, synchronized audio, dialogue, sound effects, and ambient noise all produced together, not layered on afterward. The model supports text-to-video, image-to-video, video-to-video with character consistency, keyframe-controlled transitions, multilingual dialogue, and agentic chaining that can link clips into longer multi-shot sequences. Black Forest Labs highlights particular strength in human facial expressions and in associating sounds with physical events, a glass breaking sounds like glass because the model has learned the causal relationship between impact and acoustic signature.

We're building an independent news platform focused on facts over headlines. Join us by supporting our work.

Support independent reporting

Video generation consumes more than 95 percent of the model’s training compute. Audio, by contrast, accounts for less than 0.5 percent of tokens, a figure that underscores how cheaply the audio modality attaches once the visual backbone is in place.

The robotics component, called FLUX-mimic, reuses FLUX 3’s video-prediction engine with a lightweight decoder that translates predicted frames into robot motion sequences. Audi is already testing the system for flexible door-seal installation, a soft-body manipulation task that the development team says was difficult to automate with traditional programming. FLUX-mimic responds in roughly 101 milliseconds on a single RTX 5090, and certain tasks can be fine-tuned with as little as 30 minutes of robot demonstration data.

Black Forest Labs is staging the release in phases. Video and action capabilities are available now through early access APIs and private weight access to selected partners. Image synthesis and editing will follow in the coming weeks. An open-weight version, FLUX 3 Dev, is planned for later in 2026, continuing the lab’s pattern of eventually releasing model weights to the community.

The broader signal FLUX 3 sends is about architectural direction rather than any single benchmark score. By folding image, video, audio, and robot action into one backbone, Black Forest Labs is arguing that the most efficient route to general visual intelligence is not to optimize each modality in isolation, but to force them to constrain each other during training. Whether that unified approach holds up outside curated demonstrations will be tested when independent benchmarks and the open-weight release arrive later this year.

Sources: FLUX 3 – Real World Models (Black Forest Labs, July 23, 2026); Black Forest Labs Unveils FLUX 3 (NYU Shanghai RITS, July 24, 2026); Black Forest Labs Releases FLUX 3 (MarkTechPost, July 26, 2026)

Scroll to Top