Black Forest Labs launches FLUX 3 with video, audio, and robotics path

Official FLUX 3 launch image from Black Forest Labs.Black Forest Labs
Official FLUX 3 launch image from Black Forest Labs.Black Forest Labs
AI & Automation

Black Forest Labs has put FLUX 3 into early access, expanding its image-model line into a unified multimodal system for video with native audio, image generation, and action prediction for robotics partners.

Black Forest Labs has moved FLUX beyond still-image generation. FLUX 3 is now in early access as a multimodal foundation model trained across images, video, and audio, with a stated path into action prediction for robotics.

The important qualifier is access. This is not a general release of every FLUX 3 capability. Black Forest Labs says FLUX 3 Video is available through early access first, while image synthesis, private weights, APIs, robotics access, and an open-weight FLUX 3 Dev backbone are planned in phases after safety testing and feedback.

FLUX 3 starts with video and native audio

The flagship launch lane is FLUX 3 Video. Black Forest Labs says the model can generate video with native audio up to 20 seconds in a single generation, using text prompts or references such as images, video clips, keyframes, and existing audio.

That puts FLUX 3 in the same competitive field as the current wave of world-model and video systems, not just the image-generation market where FLUX first became widely used. The supported tasks listed by Black Forest Labs include text-to-video, image-to-video, video-to-video, video-audio continuation, keyframe-to-video, multilingual dialogue, animated typography, and chaining clips into longer multi-shot sequences.

The company is also making a stronger claim than "better clips." Its launch post frames FLUX 3 as a unified model that learns how objects look, move, and sound together. In practical terms, the bet is that training across modalities should help with problems that often break generated media: matching sound to events, keeping motion physically plausible, and preserving references across scenes.

Early evaluations are promising, but still vendor-run

Black Forest Labs published preliminary preference comparisons for 10-second 720p text-to-video clips with audio. In those internal evaluations, it says FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Seedance 2.0 and Gemini Omni Flash in 52%, Runway Gen-4.5 in 77%, and Luma Ray 3.2 in 93%.

Those numbers are useful as a directional signal, not as a settled public benchmark. The company says the model and harness are still in development and that results should improve during early access. Teams deciding between video APIs should wait for hands-on tests against their own prompts, licensing constraints, latency targets, and review workflows.

The image lane is not live on the same footing yet. Black Forest Labs says FLUX 3 can synthesize and edit images across styles, aspect ratios, and resolutions, with improved prompt handling and multilingual text rendering, but the image early access phase is planned for the following weeks.

The most unusual part of the announcement is the robotics branch. Black Forest Labs says FLUX 3 extends to action prediction, and mimic robotics separately announced FLUX-mimic, a video-action model trained on top of the FLUX 3 backbone.

mimic says FLUX-mimic is already being tested and deployed with manufacturing partners including Audi. Its post describes soft-body manipulation tasks, car-door assembly work, and a preview benchmark where FLUX-mimic reached a 95% success rate on a soft-body kitting task without single-task fine-tuning or post-training. mimic also says the system runs locally on a robot using a single NVIDIA RTX 5090 GPU, with quantization, chunked action prediction, and cache reuse to keep inference practical.

That is partner-provided evidence, so it needs independent validation before anyone treats it as a new robotics baseline. Still, it shows why FLUX 3 matters outside creator tools: Black Forest Labs is trying to turn the same multimodal backbone into a foundation for content generation and physical action.

Access is staged, and the open-weight promise comes later

The launch plan names four tracks: FLUX 3 Video for video and audio generation and editing through APIs and private weight access; FLUX-mimic and FLUX 3 Action through selected research and commercial partners; FLUX 3 Image for image synthesis and editing; and FLUX 3 Dev as open-weight access to a multimodal backbone for content creation and action prediction.

For builders, that means the practical next step is not migration. It is watch-listing the access lane that matches the job: video API, image editing, private deployment, robotics research, or open-weight experimentation. Pricing, final model cards, licenses, public API details, and independent benchmarks are still the missing pieces.

Sources and methodology

This post treats the Black Forest Labs launch post as the primary source and uses mimic robotics as a first-party partner source for the robotics branch. VentureBeat coverage was used as independent media corroboration for the early-access framing and launch scope.