Introducing Clef: our open-source decision models, and new RL fine-tuning platform
- ID
- 30846
- Status
- summarized
- Published
- 01 Oct 2026, 11:34 PM
- Fetched
- 02 Oct 2026, 1:32 AM
- Provider
- Cloudflare Blog
- Category
- infrastructure
- Original URL
- https://blog.cloudflare.com/clef-decision-models/
- Source URL
- https://blog.cloudflare.com/rss/
Summary
- Score
- 7.0
- Created
- 02 Oct 2026, 1:32 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_usersvibe_coders
What happened
Cloudflare released two Cloudflare-trained "decision models" — Clef and Clef-flash — hosted on Workers AI, open-sourced on Hugging Face under Apache 2.0, and made Jev-API compatible with Typesafe AI's Jev System One. Decision models return bounded, typed outputs with probabilities (e.g. 95% fashion, 85% ecommerce, <1% phishing) instead of open-ended text, and Cloudflare says Clef currently leads the Jev Decision Index. Cloudflare also debuted an RL product for fine-tuning Clef, and reported its own Threat Intelligence workflow classified a domain in 2.2s with Clef versus 4.7s for gpt-oss-120b, which returned only two classifications.
Why it matters
If you are routing tickets, escalations, or domain/page categories inside an agent loop, a classifier that returns typed labels plus probabilities lets your code branch deterministically instead of parsing LLM prose — and since Clef is Apache 2.0 on Hugging Face you can self-host and test it without committing to Workers AI billing. Treat the 2.2s vs 4.7s figure as vendor-reported on Cloudflare's own Threat Intelligence workflow, so benchmark it on your own inputs before swapping out a prompt-based classifier. The new RL fine-tuning option is the piece to evaluate if your label set is domain-specific and you don't want to retrain a full classifier each time categories change.
Discussion angle
Is a "decision model" actually a different category from a classifier, or is the interesting part that Cloudflare is tying the format to a competitor's (Typesafe AI's Jev) API spec? Worth opening the benchmark demo site live and asking whether a vendor-hosted leaderboard on a spec owned by another vendor is a signal you can act on — then trying the Apache 2.0 weights locally.