⚡ AI Focus Bulletin
AIGCModels

Black Forest Labs debuts FLUX 3, its first video model with audio

Black Forest Labs' new flagship generates 20-second clips with native audio, claims wins over Kling and Runway, and promises open weights later in 2026.

Black Forest Labs, the Freiburg-based lab behind the FLUX image models, has announced FLUX 3 — and it is no longer just an image company. FLUX 3 is described as a single multimodal foundation model jointly trained across image, video, audio and action prediction, and its first shipping product is video: clips up to 20 seconds in one generation, with a native soundtrack — dialogue, footsteps, impacts — produced by the same model rather than a bolted-on audio stage.

What it does

FLUX 3 Video supports text-to-video, image-to-video, video-to-video and keyframe-driven transitions, with multilingual dialogue. BFL's internal human-preference tests on 720p, 10-second clips claim win rates of 93 percent against Luma's Ray 3.2, 77 percent against Runway Gen-4.5, 60 percent against Kling v3 Pro and roughly even against ByteDance's Seedance 2.0. Notably absent from the comparison set: Google's Veo and OpenAI's Sora. The numbers are the company's own, with no independent verification yet.

Availability and the open question

Access starts as a gated early-access API with private weights; pricing is undisclosed. A refreshed FLUX 3 Image model follows in the coming weeks, and — most consequentially — BFL says FLUX 3 Dev, an open-weight multimodal backbone, will land on Hugging Face and GitHub later in 2026. Early testers include Canva, Krea, Magnific, Burda and Picsart. There is also a robotics arm: FLUX-mimic, a video-action variant built with partner mimic robotics and trialed at Audi's production lab, which BFL claims cuts task fine-tuning data from 30-plus hours of robot demonstrations to about 30 minutes.

Why it matters: FLUX.1's open weights reshaped image generation by giving the ecosystem a frontier-class model to build on. If the lab follows through on an open-weight video-capable backbone, the same disruption could hit video generation, where every serious contender today is closed. The audio-native, action-trained design also signals that BFL is positioning FLUX 3 as an early world model, not merely a clip generator.

Sources