Vibeleaderboard
← All Intel
Intel / blog

Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

Source
blog.cloudflare.com
Date
Why it matters

decision models now accept audio (wav, mp3) and video (mp4, webm) alongside text and images, so classification and routing over input no longer needs a separate transcription step. Pricing also dropped.

Key takeaways · AI-distilled
  • Cloudflare built Clef-omni on Qwen3-Omni-30B-A3B without its speech output. It scores options in one prefill pass with no token generation, so audio and video are not transcribed or captioned first. Training froze the backbone and used adapters.
  • Cloudflare reports median latency of about 130 ms for text, 150 ms for images, a few hundred ms for audio clips, and about 1.5 seconds for a 21-second video with sound.
  • Clef-flash fell from $0.09 to $0.038 per million input tokens, but its hosted drops from 64k to 24k. Cloudflare says only 0.24% of requests exceeded 24k, and the self-hosted weights still support 256k.
  • Hosted Clef got 1.7x to 2.0x faster at the median, mostly from serving changes including a move to SGLang, not new weights. On Cloudflare's own benchmarks Clef-omni trails Clef on several tasks, such as Home appliances (69.3 vs 82.95).
Terms in this piece · Glossary
  • open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
  • multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
  • LoRA — A cheap way to fine-tune a model by training a small add-on layer instead of changing all of the model's weights.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Read the source blog.cloudflare.com
Recommended reads
Comments

Checking sign-in…

Loading comments…