

FLUX 3
#2 in Text-to-VideoBlack Forest Labs · v3 · 21× · last seen Jul 25, 2026
FLUX 3 is the first multimodal Foundation Model from Black Forest Labs (BFL) that goes beyond pure image generation. It learns jointly from images, videos, and audio in a unified architecture and can generate text-to-video, image-to-video, video-to-video, and keyframe-to-video with natively synchronized audio, up to 20 seconds per generation. Additionally, the architecture is extended to action prediction for robotics (e.g., collaboration with mimic robotics under the name FLUX-mimic). The model was released on July 23, 2026 as "FLUX 3 Video" in Early Access; an open weight set ("FLUX 3 Dev") as well as additional modes such as FLUX 3 Image and FLUX 3 Action are expected to follow in the coming weeks. Specific information on pricing, maximum resolution, and generation time will