multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
Why it matters
FLUX 3 Video is generally available through the BFL API: up to 20-second HD clips with native audio, multi-language lip-synced dialogue, keyframe and four-second continuation inputs, plus a low-cost draft mode for iteration.