Media GenerationModels 🇩🇪 05.08.2026 16:01

Black Forest Labs releases FLUX 3 Video with native audio and lip-synced dialogues

Black Forest Labs has released its FLUX 3 Video model, available via API and partners. It generates clips up to 20 seconds in HD/Full HD with native audio, supports text-to-video and image-to-video, and ranks first in ELO rankings.
Black Forest Labs has made its FLUX 3 Video model generally available through its API and selected partners. The model generates clips of up to 20 seconds in HD and Full HD, including native audio with dialogues, sound effects, and ambient sounds. It supports text-to-video, image-to-video, keyframes, video continuation, and multiple scenes and camera angles in one clip. FLUX 3 renders typography directly in the scene, understands complex prompts, and uses world knowledge for documentaries. Dialogues are generated lip-synced in more than 14 languages. According to BFL's own evaluations, FLUX 3 achieves first place in ELO rankings for text-to-video (1135) and image-to-video (1051), ahead of Gemini Omni Flash, Minimax H3, and Seedance 2.0. Pricing is per second of video output. The draft mode is available only in HD and costs 0.06 dollars per second (text/image-to-video) or 0.12 dollars (video-to-video). Full quality costs 0.17 or 0.41 dollars per second in HD, and 0.29 or 0.53 dollars in Full HD. Audio is included in the price. More video examples are available in the BFL blog.
Abbreviations
API = Application Programming Interface — программный интерфейс приложения
HD = High Definition — высокая четкость
Source: The Decoder (DE) — original
Our earlier posts on this topic ↓
Fresh news