July 25, 2026AgentsResearch

FLUX 3 x mimic: a video model that learned to move robots

Black Forest Labs makes image and video models. So it's a little surprising their new one is running on Audi's factory floor, moving actual robot arms.

FLUX 3 x mimic, built with mimic robotics, is a video-action model. The bet underneath it: a model trained to predict video already learned a ton of physical intuition, how objects fall, bend, catch, deform. So instead of training a robot policy from zero, they bolt a lightweight action decoder onto the intermediate features from FLUX 3's video-prediction path, and read robot actions straight out of the world representation the video model already built.

What that buys them is the hard stuff. Soft-body manipulation, flexible materials, cables, fabric, the things conventional robotics chokes on. Natural failure recovery the model was never explicitly trained to do, because the video prior already knows what a recovered grasp should look like. State-of-the-art success rates, and about 101ms reaction time, which is the difference between a demo and a production line. It's already kitting parts, inserting components, and handling flexible materials at Audi.

This is the same move as the Qwen-VLA and Xiaomi robotics stacks, but from the other direction, starting from a frontier video model instead of a language one. The pattern to watch: the more physics a generative model absorbs from pixels, the less a robot has to learn from scratch. The world model becomes the policy.

https://bfl.ai/blog/flux-3-mimic
← Previous
AREX: a deep-research agent that grades its own homework
Next β†’
Super User Daily: July 25, 2026
← Back to all articles

Comments

Loading...
>_