Introducing Clef: our open-source decision models, and new RL fine-tuning platform
Source
blog.cloudflare.com
Date
Key takeaways · AI-distilled
Cloudflare released Clef and Clef-flash, decision models that return typed answers with probabilities instead of free text. They run on Workers AI, are compatible with the Jev API, and ship as Apache 2.0 weights on Hugging Face.
Clef does a prefill-only pass on a frozen Qwen backbone (Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash), then scores the valid schema choices in parallel. Cloudflare says skipping token-by-token generation makes it much faster than autoregressive LLMs.
Training paired label-smoothed cross-entropy with a Brier loss to calibrate probabilities, plus an RL step Cloudflare calls RLCD that gives partial credit to adjacent ordinal choices and applies a reference penalty to prevent distribution shift.
In Cloudflare's own tests, median latency was 209 ms for Clef and 39 ms for Clef-flash versus 524 ms for Jev. Quality was mixed: Clef led on BANKING77 and CLINC150, while Jev led on When2Call and AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → trace observability.
Unlike Jev, Clef accepts images and has a 64k context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → versus Jev's 32k. Cloudflare's threat intel team used it to fetch, render and classify a domain in 2.2s, against 4.7s for gpt-oss-120b, which returned only two categories.
Terms in this piece · Glossary
open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters
Clef gives workflows a cheap, bounded-output decision step with open Apache 2.0 weights and a Jev-compatible API. Cloudflare's RL fine-tuningTaking a trained model and training it a bit more on your own examples so it gets better at one specific job.Full definition → lets you adapt it to your own categories instead of prompting a general LLM.