FLUX.2 [klein] 9B just got 2x faster at image editing, especially when you use multiple reference images. Same quality, no price increase.

Under the hood: KV-caching lets the model skip redundant computation on your reference images. The more references you use, the bigger the speedup. Inference is up to 2x+ faster for multi-reference editing.
We're also releasing FP8 quantized weights, built with @NVIDIA_AI_PC Run Klein 9B with less VRAM and faster inference for local and self-hosted deployments.
Already on FLUX.2 [klein] 9B via API? Free upgrade, faster, same price. On [klein] 4B and want better quality? 9B is now closer in speed. Docs → https://t.co/ZhlyBkTgG9 Weights → https://t.co/Aao6QYvTAt Try it → https://t.co/b7cnzrss3B
Multi-reference editing was the expensive path in FLUX pipelines. Caching the reference computation removes that penalty at no price change, and the speed gap to the smaller 4B model narrows.
Checking sign-in…
Loading comments…