Logic of Logic
thursday, august 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
brief productsresearch

FLUX 3 unifies image, video, audio, robots

Black Forest Labs' FLUX 3 trains one model jointly on images, video, and audio, and is already driving soft-body manipulation in Audi's production lines.

Black Forest Labs introduced FLUX 3 on July 23, and unlike prior FLUX releases, it is one model trained jointly on images, video, and audio, extendable to predicting physical actions, rather than separate models bolted together. CEO Robin Rombach’s rationale is that “you can’t cheat reality”: a static image model can only ever produce stills, so FLUX 3 learns from footage and sound that actually move and change over time. The model ships in four variants, FLUX 3 Video, Image, Action, and Dev, with Video and Action in early access now, Image following in the coming weeks, and full benchmark results due alongside general availability.

The clearest proof point is industrial, not creative. Black Forest Labs and mimic robotics launched FLUX-mimic, a video-action model built on FLUX 3 for robotic manipulation, and Audi is already testing it in production. Mimic CTO Elvis Nava says the model can pick up a brand-new task after only a half hour of hands-on robot practice, well under the multi-day training prior approaches needed, because it inherits a physical-world understanding from FLUX 3’s video training rather than learning each task from scratch. Audi’s production lab credits it with soft-body manipulation work that conventional robotics could never have pulled off.

What it means for operators

FLUX already powers generative features inside Adobe Photoshop, Picsart, and Nous Research’s Hermes Agent, so a joint-architecture successor is a real migration question for anyone building on the family, not a lab curiosity. Black Forest Labs says open-weight and faster versions of FLUX 3 are coming later this year, continuing the open-access pattern that has differentiated FLUX from closed rivals. If your roadmap touches image or video generation, the near-term move is to apply for FLUX 3 Video/Action early access and watch for the Image and open-weight rollouts before committing production traffic; if you touch physical automation, FLUX-mimic’s 30-minutes-of-data claim is worth testing against your own task before dismissing vision-first robotics as still too far from the plant floor.

sources 3 cited
1 globenewswire.com Black Forest Labs Unveils FLUX 3, A New Multimodal Frontier Model For Visual Intelligence 2 bfl.ai FLUX 3 model page 3 latent.space [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine
next