ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Cloudflare Clef decision model released as open source: writes not a single word, delivers a judgment in 38 milliseconds

Cloudflare Clef decision model released as open source: writes not a single word, delivers a judgment in 38 milliseconds

AI information • Admin • • 5 views

Cloudflare Clef was officially released by Cloudflare through its official blog on October 1, 2026, as the first models trained in-house by the Workers AI team. The series includes Clef and the lighter Clef-flash. Both are now live on the Workers AI hosting service, and the weights have been open-sourced on Hugging Face under the Apache 2.0 license, so developers can either call the hosted API directly or download the weights and deploy them themselves.

It writes no text, only probabilistic judgments

Clef belongs to the "decision models" that have heated up quickly over the past two weeks. The concept was popularized by TypeSafe AI's Jev: its System One model, launched in mid-September, generates no free text, but instead gives typed answers with probabilities to predefined questions, such as yes or no, one choice among several, or a score. For background, see TypeSafe AI and Jev's decision model concept. Once a program receives such a structured result, it can directly classify, route, escalate, or allow something based on it. Clef is fully API-compatible with Jev, so teams already using Jev only need to change the endpoint and model name to switch.

In terms of specs, Clef was post-trained from Qwen3.8-27B, using a frozen backbone plus rank-256 LoRA, and it carries a vision encoder that can handle image inputs. Its context length is 64k, above Jev's 32k. Clef-flash is based on Qwen3.5-9B and is built for even lower latency.

Official benchmarks are fast, but not ahead on everything

According to Cloudflare's own test data, on the BANKING77 intent classification task Clef scored a macro-F1 of 94.20, compared with Jev's 79.74. For median latency, Clef took 209.3 milliseconds, Clef-flash just 38.8 milliseconds, and Jev 524.1 milliseconds. Cloudflare says Clef currently ranks first on the Jev Decision Index, but the boundary needs stating clearly: on individual items such as When2Call and BRIGHT, Jev still leads, and Clef does not win every single test.

Cloudflare is already using it internally. Its threat intelligence team uses Clef to classify website domains, and fetching the page, rendering it, and classifying it takes 2.2 seconds in total. The same workflow with the fastest general-purpose large model, gpt-oss-120b, took 4.7 seconds and returned only two categories, far less granular than a dedicated decision model.

Why it is fast, and the fine-tuning service that comes with it

Clef is fast because it infers differently: it performs a single prefill, does not generate word by word autoregressively, and instead scores every valid option in parallel in a non-autoregressive way, so one forward pass yields the probabilities for all options. During training, it used cross-entropy with label smoothing together with a Brier loss, and calibrated its probabilities through RLCD (Reinforcement Learning for Calibrated Decisions), so the confidence it outputs sits closer to its true accuracy.

Launched alongside the models is a reinforcement learning fine-tuning service. At first, Cloudflare's forward deployed engineers will work alongside customers, and later it will become a self-service platform. Its components include AI Gateway, which accumulates request data; Workers AI, which generates rollouts; Containers, which provide the RL sandbox; Trainer, which updates the weights; and BYO Model, the redeployment step built on technology from the Replicate acquisition.

Within two weeks, decision models have gone from a single product to a category. The contest is not about being smarter than general-purpose large models, but about splitting the high-frequency, low-risk judgments in agent workflows out of expensive large-model calls and handling them separately. The boundary is just as clear: a decision model only judges, it does not reason or generate, so in real deployments it has to be paired with a large model.

Recommended Tools

More