[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cloudflare-clef-open-source-decision-model-jev":3,"topics-all":38,"news-related-7a2aa1aa-74ec-494a-a828-e6c1ad5b4cfe":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"7a2aa1aa-74ec-494a-a828-e6c1ad5b4cfe","Cloudflare Clef 决策模型开源,Jev 被压制","Cloudflare 发布首款自训练决策模型 Clef 和 Clef-flash,基于 Qwen 做非自回归 schema-bound 打分,Apache 2.0 开源。Clef-flash 中位延迟 38.8 毫秒,比 Jev 快 13 倍,在 43 项跑分上多数基准击败 Jev。配套推出 RL 微调平台。","Cloudflare 在 AI 圈是 Workers AI 的算力服务商,而 2026 年 10 月 1 日的 Birthday Week 上,他们突然把身份往前挪了一步:发布首款自训练决策模型 Clef 与轻量版 Clef-flash,Apache 2.0 开源到 Hugging Face。决策模型是过去几个月新冒出的一个细分赛道——Typesafe 的 Jev 把「分类器不再需要为每个新类别重训」做成了产品,让 AI agent 在需要做选择题时不用去调一个庞大的通用 LLM。Cloudflare 的 Clef 直接瞄准这个赛道,而且拿出了比 Jev 更激进的数字。\n\n## 从 Qwen 出发,做非自回归的 schema-bound 打分\n\nClef 和 Clef-flash 的基座都是开源 Qwen——Clef 用 Qwen3.8-27B,Clef-flash 用 Qwen3.5-9B。Cloudflare 把 Qwen 的权重全部冻结,只额外训练一个很小的「联合 schema head」(一个小型 transformer),同时配合 rank-256 的 LoRA 适配器。推理时,模型先做一次 prefill,然后用这个 schema head 在所有合法选项之间并行打分——而不是像普通 LLM 那样一个字一个字生成中间文字再 parse 成结构化结果。这意味着 Clef 没有「自由文本生成」,而是把 schema 当作骨架,直接把每个候选项的概率一次性输出。\n\n论文级的技术细节都藏在 Cloudflare Blog 的训练方法段:他们用了标签平滑交叉熵 + Brier 损失来校准概率,还在这篇博客里首次披露了一个叫 RLCD (Reinforcement Learning for Calibrated Decisions) 的自创方法——一种专门为决策模型设计的强化学习目标,给相邻的 ordinal 选项部分奖励,对完全精确的输出给满分,并用参考惩罚避免分布漂移。这是 Cloudflare 第一次把 RL 用在「概率校准」上,而不是常规的 RLHF 优化回答质量。\n\n## 跑分与延迟全面压制 Jev\n\nCloudflare 在 43 项 Jev Decision Index 跑分上,把 Clef、Clef-flash、Jev、DiffusionGemma Jev、Kev-9B 和 Laya 全部跑了一遍。Clef 在 ToolRet、BANKING77、CLINC150+OOS、PhishNChips 等近半数基准上拿到最高分,其余多与 Jev 互有胜负;Clef-flash 因为更激进地优化 latency,在 BFCL、API-Bank、MMLU、ARC-Challenge、HellaSwag、SATA-Bench 等更多子项上拿下第一。Clef-flash 的中位延迟只有 38.8 毫秒,比 Jev 的 524.1 毫秒快 13 倍以上;Clef 自己是 209.3 毫秒,仍然比 Jev 快一倍多。\n\nClef 还配了一个 Qwen 原生就有的视觉编码器,可以同时处理图像和视频输入,Jev 目前只支持文本。上下文窗口也给到了 64k,是 Jev 32k 的两倍。整套 API 兼容 Jev\u002FSystemOne,所以现有基于 Jev 的代码可以零修改切换到 Clef,只换 endpoint 即可。\n\n## 配套上新的 RL 微调平台\n\nClef 不只是模型。Cloudflare 同步发布了他们的新 RL 产品:先用 Cloudflare AI Gateway 抓取所有经过你账户的 AI 请求作为训练数据,Workers AI 用来跑 rollout,Cloudflare Containers 做 RL sandbox 评分,新的 Trainer 更新权重,最后用 BYO Model 部署回 Workers AI。这个闭环走的是 Cloudflare 自己的现有基建——AI Gateway、Containers(收购自 Replicate 之后的 Cog 工作)——所以客户不需要自己搭任何后端。起步阶段是 Cloudflare 的前向工程师(FDE)团队手把手合作,之后会开放自服务。\n\n## 个人判断:Cloudflare 的「agent cloud」战略在这里落地了\n\nCloudflare 的 CEO Matthew Prince 这两年反复说 Cloudflare 要做「agent cloud」,把 Workers AI、AI Gateway、Containers、Access 这些产品线串成一个 agent 友好的边缘基础设施。这次 Clef + RL 平台其实是把这条战略落到了模型层:Cloudflare 不再只是卖 GPU 算力,而是要拿到企业用自家数据 fine-tune 决策模型的全栈能力。配合他们自己多年沉淀的网络数据(Cloudflare 提到有 15 年跨域标签数据可用于训练 Trust & Safety 分类器),这种垂直整合的能力,大厂和初创公司都很难拼。\n\n更深一层的信号是「决策模型」这个赛道才刚冒头两个月,就已经同时有 Typesafe Jev、DiffusionGemma、Cloudflare Clef \u002F Clef-flash、Kev-9B、Laya 等五家产品共存,而且其中三家是互联网基础设施工厂——这说明市场对「小、快、可控、不抽风」的需求正在快速增长,LLM 这种大而通用、不可预测的形态在 agent 的「决策节点」上越来越不顺手。Clef 把 Qwen 当基座也再次印证了之前的观察:中国开源基础模型已经在支撑得起「边缘模型」和「小模型场景化创新」的格局——海外大厂和初创都在用 Qwen 当后厨的料。\n\n所以呢:如果你是 agent 开发者,Clef 现在的硬指标(速度 13 倍快、跑分基本压制、Apache 2.0 开源、API 兼容 Jev)意味着今天就能把它当作 Jev 的替代品试起来,而不必等生态慢慢长大;如果你是企业 AI 决策方,Cloudflare 这套 RL 微调闭环可能是接下来几个季度里最值得认真评估的「垂直决策模型」基础设施之一——不是因为它技术最强,而是因为它把「企业私有数据 + 边缘推理 + 模型微调」打包到了一家公司里。","https:\u002F\u002Fblog.cloudflare.com\u002Fclef-decision-models\u002F","016517e1-7a07-4d85-a702-cbf30d014b76",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"f4761185-4ebd-45a1-8490-f3078356c4d4","en","Cloudflare Clef: open-source decision model beats Jev","Cloudflare released its first self-trained decision models, Clef and Clef-flash, built by freezing Qwen backbones and training a joint schema head for non-autoregressive schema-bound scoring, released under Apache 2.0. Clef-flash hits a 38.8 ms median latency, roughly 13x faster than Jev's 524.1 ms, and beats Jev on most of the 43 Jev Decision Index benchmarks. A new RL fine-tuning platform ships alongside the models.","Cloudflare has long been a compute partner for AI workloads through Workers AI, but on October 1, 2026, during Birthday Week, the company stepped forward into model production: it released Clef and the lighter Clef-flash, its first self-trained decision models, open-sourced under Apache 2.0 on Hugging Face. The decision-model category is itself only a few months old — Typesafe's Jev packaged the idea that classifiers shouldn't have to be retrained for every new class, letting agents pick from structured options without burning an LLM inference. Clef targets that category head-on, with numbers that are noticeably more aggressive than Jev's.\n\n## Frozen Qwen, non-autoregressive schema-bound scoring\n\nBoth Clef and Clef-flash are built on open-weight Qwen backbones. Clef uses Qwen3.8-27B; Clef-flash uses Qwen3.5-9B. Cloudflare freezes the Qwen weights entirely and trains only a small joint schema head (a lightweight transformer) plus rank-256 LoRA adapters. At inference, the model runs a single prefill pass through Qwen, then the schema head scores every legal option for every field in parallel — there is no autoregressive text generation and no post-hoc parsing of free-form output into structured answers. Clef essentially treats the schema as a skeleton and emits per-option probabilities in one forward pass.\n\nThe technical details sit in the Cloudflare Blog training section. The team used label-smoothed cross-entropy plus a Brier loss for probability calibration, and for the first time disclosed RLCD (Reinforcement Learning for Calibrated Decisions) — a custom RL objective designed for decision models, which gives partial credit to adjacent ordinal choices, full reward to exact outputs, and applies a reference penalty to prevent distribution shift. It is the first time Cloudflare has used RL for probability calibration rather than the usual RLHF style of optimizing answer quality.\n\n## Benchmarks and latency: Clef beats Jev across the board\n\nCloudflare ran the full 43-benchmark Jev Decision Index against Clef, Clef-flash, Jev, DiffusionGemma Jev, Kev-9B, and Laya. Clef takes the top score on roughly half the benchmarks, including ToolRet, BANKING77, CLINC150+OOS, and PhishNChips, and trades wins with Jev on the rest. Clef-flash, tuned harder for latency, wins more subtests — BFCL, API-Bank, MMLU, ARC-Challenge, HellaSwag, SATA-Bench and others. Clef-flash hits a 38.8 ms median latency, more than 13x faster than Jev's 524.1 ms; Clef itself is 209.3 ms, still about 2.5x faster than Jev.\n\nClef also inherits Qwen's native vision encoder, so it accepts image and video inputs out of the box; Jev today only supports text. The context window is 64k — double Jev's 32k. The full API is Jev\u002FSystemOne-compatible, so existing Jev code can swap endpoints with zero changes.\n\n## A new RL fine-tuning platform ships alongside\n\nThe release is not just the model. Cloudflare also announced a new RL fine-tuning product: AI Gateway captures every request flowing through your account as training data, Workers AI runs rollouts against the base Clef model, Cloudflare Containers act as the RL sandbox for scoring and replay, a new Trainer component updates the weights, and BYO Model deploys the fine-tuned variant back to Workers AI. The whole loop is built on Cloudflare's existing primitives (AI Gateway, Containers, the Cog work from the Replicate acquisition), so customers do not need to assemble their own backend. The early phase is hand-on with Cloudflare's Forward Deployed Engineer team; a self-serve platform will follow.\n\n## Take: Cloudflare's agent-cloud strategy just landed at the model layer\n\nCloudflare CEO Matthew Prince has been saying for two years that Cloudflare wants to be the agent cloud — stringing Workers AI, AI Gateway, Containers, and Access into a coherent edge infrastructure that agents can lean on. Clef plus the new RL platform is that strategy reaching the model layer. Cloudflare is no longer just selling GPU time; it is positioning itself to own the full stack that lets enterprises fine-tune decision models on their own data. Combined with Cloudflare's accumulated network data (the team mentions 15 years of cross-domain labels for training Trust & Safety classifiers), that vertical integration is hard for incumbents or startups to match in one place.\n\nA second signal: the decision-model category is barely two months old and already has five coexisting products — Typesafe Jev, DiffusionGemma, Cloudflare Clef \u002F Clef-flash, Kev-9B, Laya — three of them from internet infrastructure companies. That tells you the appetite for small, fast, controllable, non-hallucinating models at the agent decision node is real, and general-purpose LLMs are losing ground there fast. The fact that Cloudflare built Clef on Qwen is also the latest data point on a trend we have been tracking: Chinese open-weight base models are now sturdy enough to anchor serious edge and small-model innovation in the rest of the world.\n\nSo: if you build agents, Clef's hard numbers today — 13x faster than Jev, broadly ahead on benchmarks, Apache 2.0, Jev-compatible API — mean you can treat it as a drop-in Jev replacement right now, without waiting for the ecosystem to grow. If you sit on the enterprise AI decision side, the new RL fine-tuning loop is probably the most credible 'vertical decision-model' infrastructure on offer this quarter — not because the underlying model is the strongest, but because it bundles private enterprise data, edge inference, and model fine-tuning into a single vendor.","cloudflare-clef-open-source-decision-model-jev","2026-10-02T09:00:00Z","2026-10-02T09:09:30.088207Z","2026-10-02T09:09:30.088216Z",true,"agent",75,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"71dde565-b87a-4470-83a5-0ad7a4ee1787","IBM Granite 4.2 开源:原生思维链做成开关,30B 拿下 SWE Bench Pro","ibm-granite-4-2-native-reasoning-agents","2026-08-28T14:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"b4754043-6b19-499f-8459-f8fc786f4d80","Pokee-Isaac 28B 把 10M 上下文塞进客户边界:28B 参数在 RULER 10M 上 93.3%","pokee-isaac-28b-10m-context","2026-08-20T14:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"79c1684f-f61d-4799-b3d0-6450c4ad10e8","Muse Glimmer:Meta 把 30B 「常驻本地的智能体」开源,把 Agent 拉到笔记本里 7×24 跑","muse-glimmer-30b-open-agentic-local","2026-08-10T00:00:00+00:00"]