Read as article
Cloudflare Clef: Open-Weight Models Return Probabilities
By @sharedot · · 8 pages
- AI Frontier
- Open Weight Models
- Decision Models
- Cloudflare
- Agents
Cloudflare released Clef and Clef-flash, Apache 2.0 decision models that return typed probabilities instead of text and beat TypeSafe's Jev on most cited benchmarks.
What happened
They are decision models, not chatbots: each reads an input state plus a schema of typed questions — yes/no, multiple choice, or rubric scores — and returns a probability for every allowed answer with no free-form text. Both are open-weight under Apache 2.0, posted to Hugging Face, run today on Workers AI, and are compatible with TypeSafe AI's Jev API, so switching means changing just the endpoint and model name.
Why it is surprising
The speed numbers are the shock. According to MarkTechPost, Clef's median latency is 209.3 ms and Clef-flash's is 38.8 ms, against 524.1 ms for Jev — Clef-flash runs roughly thirteen times faster. Startup Fortune reports Cloudflare's threat-intelligence workflow classified a domain in 2.2 seconds versus 4.7 seconds for gpt-oss-120b, while returning more labels. Unlike OpenAI's Decisions API or Amazon's Strands Decider 2B, Cloudflare gave away full weights under Apache 2.0 at both sizes.
The evidence
On Cloudflare's 10-benchmark shortlist from the Decision Index 0.2.1 suite, a Clef model scored highest on 7, per MarkTechPost: BANKING77 macro-F1 of 94.20 versus Jev's 79.74, and Clef-flash's 97.73 versus Jev's 52.27 on home-appliance classification. The flip side: Jev stays ahead on knowledge-heavy tests, leading GPQA Diamond 78.3 to 48.0, MMLU-Pro 82.7 to 65.9, and BBH 92.9 to 73.7. Pasquale Pillitteri adds that Jev keeps leads on When2Call (80.97 vs 72.37) and BRIGHT (47.52 vs 45.91), and Streamline notes all benchmark figures are vendor-reported and not yet independently reproduced.
Under the hood
Clef is post-trained from Qwen3.8-27B and Clef-flash from Qwen3.5-9B, both keeping the backbone's vision encoder and accepting up to 4 images per request within a 64K-token context. MarkTechPost describes two-stage inference: a prefill-only pass over state and questions, then a small joint-schema-head transformer that routes evidence, lets fields cross-attend, and scores all options jointly. Training froze the backbones, tuned rank-256 LoRA adapters with cross-entropy plus Brier loss for calibration, and added RLCD, an objective giving partial credit to adjacent ordinal choices.
A crowded field forms fast
The decision-model category went from one product to many in under two weeks. PR Newswire reports Fastino Labs also released GLiDE, a 'thinking' decision model scoring 64.81 on Decision Index 0.2.1, 6.90 points ahead of Jev's published 57.91, with adaptive reasoning reserving extra compute for the hardest third of requests.
Stakes and what's next
The real bet, as Startup Fortune frames it, is infrastructure: Cloudflare doesn't need Clef to be the smartest model, it needs agent pipelines running through its edge network. Alongside the models, MarkTechPost reports an RL fine-tuning service starting with forward-deployed engineers before going self-serve, built on AI Gateway, Workers AI, Containers and a new Trainer component. Developers can call the models via env.AI.run(), the REST API, or AI Gateway, or self-host on a single H200. Expect the contest to shift from benchmarks to calibration and provenance in real workflows.
Sources
- marktechpost.com › Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text
- startupfortune.com › Cloudflare Open Sources Clef to Challenge OpenAI and Amazon on AI Agent Decisions
- pasqualepillitteri.it › Cloudflare launches Clef, the model challenging Jev on AI decisions
- streamlinefeed.co.ke › Cloudflare Clef Brings Open Weights to AI Decisions
- prnewswire.com › Fastino Labs Releases GLiDE, the First Thinking Decision Model