Cloudflare Releases Open-Weight Clef Decision Models
Cloudflare released two open-weight models that return calibrated decisions instead of text, targeting faster and cheaper control loops for AI agents.
What happened
Cloudflare released Clef and Clef-flash, two decision models trained by its Workers AI team. The models accept an input state and typed questions, then return probabilities for bounded answers such as whether a support request is urgent or which team should handle it. Unlike a conventional language model, they do not generate free-form text or require reasoning tokens that an application must parse.
Cloudflare says Clef is available through Workers AI and that both models are open-sourced on Hugging Face under the Apache 2.0 license. The company also introduced an reinforcement-learning fine-tuning service, initially aimed at design partners that want to adapt the models to internal decision rubrics.
Why developers may care
The release is aimed at the hot path of agent systems: routing tickets, checking whether a tool call should proceed, classifying content, or deciding when to escalate to a person. Cloudflare reports median latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, compared with 524.1 milliseconds for Typesafe AI’s Jev in its comparison. It also says a Clef model ranked first on seven of ten decision benchmarks.
Those figures are company-reported, and the benchmark advantage may not transfer to every domain. The more durable proposition is architectural. A small model that emits a constrained probability distribution can sit before a larger model, reducing unnecessary generation and making application behavior easier to audit. Cloudflare’s global network further lets it place these decisions near users and existing Workers workloads.
Why it matters
As agents gain permission to act, many production failures will involve choosing the wrong action rather than writing poor prose. Clef makes that intermediate control layer a directly deployable, open-weight product, potentially turning decision models into standard infrastructure for agent routing and guardrails.