Text-to-Video / Native Audio
For the first time,
video was born with sound attached.
Black Forest Labs, until now an image-only specialist, has made "FLUX 3 Video" — which generates video and audio together — generally available. In the company's own preliminary evaluation, it claims a 52% preference rate against Seedance 2.0 and Gemini Omni Flash — a narrow lead.
The Launch
Breaking its own "image-only" identity
FLUX 3 Video ships as BFL's first video model with native audio.
Germany's Black Forest Labs announced on its official blog that FLUX 3 Video is now generally available. Generating audio synchronized to on-screen action, alongside the video itself, is a first for the company — and clips can run up to 20 seconds. THE DECODER's coverage frames this as a milestone: BFL's first video model with native audio.
Until now, BFL built its reputation on the FLUX line of still-image models, while video generation was the territory of Sora 2, Runway Gen-4, Veo, and Kling AI. FLUX 3 Video is positioned as a multimodal foundation model trained across image, video, and audio together — so this isn't just BFL entering video, it's choosing to compete specifically on simultaneous audio generation, a battleground few rivals have fully staked out.
By The Numbers
The "narrow win," in numbers
A 52% figure also means the other model was preferred in roughly half of comparisons. BFL itself states plainly that this evaluation is preliminary and has not been independently verified. Taken with that caveat, BFL says its edge comes from facial-expression fidelity, tighter sound-to-event correspondence, and multilingual support.
The Landscape
How the video-generation field shifts
| FLUX 3 Video | Established players |
|---|---|
| Generates video and audio together, synced | Sora 2 / Runway Gen-4 / Veo / Kling are video-first |
| Up to 20s; consistency across multi-scene sequences lasting minutes | Length and consistency specs vary by vendor |
| A Dev edition is planned to go open-weight | Major players mostly stay closed |
BFL's official blog credits FLUX 3 Video's strengths to three things: facial-expression fidelity, the correspondence between sound effects and on-screen events, and multilingual support. The claim that "given a reference image, the model keeps a character's appearance consistent across multiple scenes for minutes-long sequences" speaks directly to ad and short-film production teams who need to preserve a character's identity across cuts.
Who It Hits
Who this affects, and how
Designers and video producers
The separate pass of adding sound effects and ambience in another tool could collapse into a single generation step — though quality still needs to be verified case by case.
Marketers
For short ad or social clips, generating dialogue and ambient sound in one pass could shorten production lead time. Multilingual support also fits international rollouts.
PMs and tool evaluators
Treat this as one more candidate on the shortlist. Since BFL's numbers await independent verification, run your own A/B test against your actual footage before deciding.
What's Next
What's coming, and what to discount
For the near-term outlook, BFL says FLUX 3 Image (the still-image counterpart) will follow "within weeks," with an open-weight "FLUX 3 Dev" edition planned later in 2026. The recommended action is to test it against Seedance 2.0 and your current tools on your own use case — clip length, language, character-consistency needs — first. Don't take BFL's 52% at face value; verifying it on your own footage is the only reliable path.
As for the counter view and risk: that 52% preference figure is BFL's own internal evaluation, and results like this can shift depending on the comparison set, the raters, and how prompts are designed. Until an independent third-party benchmark appears, it's more accurate to read this as "claims a narrow lead" than as a settled fact. FLUX 3 Video's general availability is also limited for now, and both the Image and Dev release timelines remain loosely worded — "within weeks" and "later in 2026." The video-generation field is already crowded with proven players — Sora 2, Runway Gen-4, Veo, Kling AI — and whether BFL can actually move market share depends on how it performs in real-world use going forward.
BFL has an established position in image generation, but in this new territory of video and audio it remains, by definition, a challenger — and this general release is only the first test of whether the claims hold up.