Cloudflare has long been a compute partner for AI workloads through Workers AI, but on October 1, 2026, during Birthday Week, the company stepped forward into model production: it released Clef and the lighter Clef-flash, its first self-trained decision models, open-sourced under Apache 2.0 on Hugging Face. The decision-model category is itself only a few months old — Typesafe's Jev packaged the idea that classifiers shouldn't have to be retrained for every new class, letting agents pick from structured options without burning an LLM inference. Clef targets that category head-on, with numbers that are noticeably more aggressive than Jev's.

Frozen Qwen, non-autoregressive schema-bound scoring

Both Clef and Clef-flash are built on open-weight Qwen backbones. Clef uses Qwen3.8-27B; Clef-flash uses Qwen3.5-9B. Cloudflare freezes the Qwen weights entirely and trains only a small joint schema head (a lightweight transformer) plus rank-256 LoRA adapters. At inference, the model runs a single prefill pass through Qwen, then the schema head scores every legal option for every field in parallel — there is no autoregressive text generation and no post-hoc parsing of free-form output into structured answers. Clef essentially treats the schema as a skeleton and emits per-option probabilities in one forward pass.

The technical details sit in the Cloudflare Blog training section. The team used label-smoothed cross-entropy plus a Brier loss for probability calibration, and for the first time disclosed RLCD (Reinforcement Learning for Calibrated Decisions) — a custom RL objective designed for decision models, which gives partial credit to adjacent ordinal choices, full reward to exact outputs, and applies a reference penalty to prevent distribution shift. It is the first time Cloudflare has used RL for probability calibration rather than the usual RLHF style of optimizing answer quality.

Benchmarks and latency: Clef beats Jev across the board

Cloudflare ran the full 43-benchmark Jev Decision Index against Clef, Clef-flash, Jev, DiffusionGemma Jev, Kev-9B, and Laya. Clef takes the top score on roughly half the benchmarks, including ToolRet, BANKING77, CLINC150+OOS, and PhishNChips, and trades wins with Jev on the rest. Clef-flash, tuned harder for latency, wins more subtests — BFCL, API-Bank, MMLU, ARC-Challenge, HellaSwag, SATA-Bench and others. Clef-flash hits a 38.8 ms median latency, more than 13x faster than Jev's 524.1 ms; Clef itself is 209.3 ms, still about 2.5x faster than Jev.

Clef also inherits Qwen's native vision encoder, so it accepts image and video inputs out of the box; Jev today only supports text. The context window is 64k — double Jev's 32k. The full API is Jev/SystemOne-compatible, so existing Jev code can swap endpoints with zero changes.

A new RL fine-tuning platform ships alongside

The release is not just the model. Cloudflare also announced a new RL fine-tuning product: AI Gateway captures every request flowing through your account as training data, Workers AI runs rollouts against the base Clef model, Cloudflare Containers act as the RL sandbox for scoring and replay, a new Trainer component updates the weights, and BYO Model deploys the fine-tuned variant back to Workers AI. The whole loop is built on Cloudflare's existing primitives (AI Gateway, Containers, the Cog work from the Replicate acquisition), so customers do not need to assemble their own backend. The early phase is hand-on with Cloudflare's Forward Deployed Engineer team; a self-serve platform will follow.

Take: Cloudflare's agent-cloud strategy just landed at the model layer

Cloudflare CEO Matthew Prince has been saying for two years that Cloudflare wants to be the agent cloud — stringing Workers AI, AI Gateway, Containers, and Access into a coherent edge infrastructure that agents can lean on. Clef plus the new RL platform is that strategy reaching the model layer. Cloudflare is no longer just selling GPU time; it is positioning itself to own the full stack that lets enterprises fine-tune decision models on their own data. Combined with Cloudflare's accumulated network data (the team mentions 15 years of cross-domain labels for training Trust & Safety classifiers), that vertical integration is hard for incumbents or startups to match in one place.

A second signal: the decision-model category is barely two months old and already has five coexisting products — Typesafe Jev, DiffusionGemma, Cloudflare Clef / Clef-flash, Kev-9B, Laya — three of them from internet infrastructure companies. That tells you the appetite for small, fast, controllable, non-hallucinating models at the agent decision node is real, and general-purpose LLMs are losing ground there fast. The fact that Cloudflare built Clef on Qwen is also the latest data point on a trend we have been tracking: Chinese open-weight base models are now sturdy enough to anchor serious edge and small-model innovation in the rest of the world.

So: if you build agents, Clef's hard numbers today — 13x faster than Jev, broadly ahead on benchmarks, Apache 2.0, Jev-compatible API — mean you can treat it as a drop-in Jev replacement right now, without waiting for the ecosystem to grow. If you sit on the enterprise AI decision side, the new RL fine-tuning loop is probably the most credible 'vertical decision-model' infrastructure on offer this quarter — not because the underlying model is the strongest, but because it bundles private enterprise data, edge inference, and model fine-tuning into a single vendor.