[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-perplexity-decider-v1-27b-open-source":3,"topics-all":38,"news-related-983fb4d1-6c63-4828-adbf-a59c57e02b64":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"983fb4d1-6c63-4828-adbf-a59c57e02b64","Perplexity开源决策模型:总分微胜Jev","Perplexity开源首个决策模型：26B权重、Apache 2.0协议，基于Qwen3.8-27B微调，面向智能体的路由分类判断。模型卡显示11项基准总评85.71%略超Jev，检索真值一项领先逾11分，体现搜索数据基因。运行需约49GiB放权重，月下载量仅165次，目前是给有显卡的工程团队用的决策组件。","2026 年 10 月 1 日，Perplexity 把自己的第一个决策模型放上了 Hugging Face：pplx-decider-v1-27b，Apache 2.0 协议，权重 26B。一家以搜索起家的公司，入场位置不是又一个通用大模型，而是智能体工作流里最不起眼也最烧钱的一环：做决定。\n\n## 决策模型是什么\n\n按 Cloudflare 在同期发布里的解释，决策模型做的事是分类与判断：输入一段客服消息，输出带概率的类型化答案——是否紧急、该哪个团队处理，代码直接拿去路由、升级或转人工。它和 LLM 的分工在于：LLM 开放生成但非确定，决策模型输出有界、便宜、快、稳定。这个品类由 Typesafe AI 的 Jev 带火，Cloudflare 直言市场已经「越来越饱和」——一周之内，Fastino 的 GLiDE、Cloudflare 的 Clef、Perplexity 的 Decider 相继亮相，迭代速度以周计。\n\n## 成绩单：总评压 Jev，赢在检索真值\n\n官方模型卡给出了 11 项基准的对照：总评 85.71%，略超 Jev 的 84.51%，大幅超过底座 Qwen3.8-27B 的 74.76%。细看单项更有意思：Decider 赢在 RAGTruth（88.80% 对 77.27%，领先逾 11 个百分点）、FinancialPhraseBank、TabFact、Circa 这类「给定材料做判断」的题；Jev 则守住 WinoGrande、BBH、JudgeBench、TruthfulQA 和自家 JevBench。一家搜索引擎公司微调出的模型，最强的恰好是检索真值判断——数据基因藏不住。需要说明：这些数字由 Perplexity 通过自家 API 测得并公布，属自报成绩。\n\n## 开源，但有门槛\n\nDecider 支持 choice、noul 两类输出，返回校准过的概率，还能直接吃图片做视觉判断。但 26B BF16 权重就需要约 49 GiB 空间，还要另加工作内存，CUDA 环境、Python 3.12 起步；模型卡页面显示的月下载量还只有 165。换句话说，这是给有显卡的工程团队准备的组件，不是拿来即用的 API 替代品。\n\n## 所以呢\n\n决策模型正在变成智能体栈里的独立一层：路由、分类、护栏这类高频小判断，不必每次都请示前沿大模型。Perplexity 的入场说明两点：第一，搜索公司攒下的判断类数据能直接换成模型优势；第二，这个品类的窗口期不会太长，等 API 巨头把决策能力打包进平台，开源权重的差异化就是入场券。做智能体的团队现在可以问自己一个问题：你的工作流里，有多少次大模型调用其实只是在做一道选择题？模型卡见 [Hugging Face](https:\u002F\u002Fhuggingface.co\u002Fperplexity-ai\u002Fpplx-decider-v1-27b)。","https:\u002F\u002Fhuggingface.co\u002Fperplexity-ai\u002Fpplx-decider-v1-27b","24d5c6c5-6573-4180-a1fd-f1459842d1af",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"3f00b080-2ae5-465f-b27d-d502b48f1c6d","en","Perplexity Open-Sources Decider, Edging Jev Overall","Perplexity open-sources pplx-decider-v1-27b, 26B Apache 2.0 weights from Qwen3.8-27B: 85.71% on 11 benchmarks, edging Jev, 11 up on retrieval truth.","On October 1, 2026, Perplexity put its first decision model on Hugging Face: pplx-decider-v1-27b, Apache 2.0, 26B weights. A company known for search entered the arena not with another general LLM, but with the least glamorous and most expensive part of agent workflows: making decisions.\n\n## What is a decision model\n\nAs Cloudflare explained in its same-week release, a decision model classifies and judges: feed in a customer support message, get typed answers with probabilities — is it urgent, which team should handle it — and your code routes, escalates, or defers to a human directly. The division of labor with LLMs: LLMs generate openly but non-deterministically; decision models produce bounded, cheap, fast, consistent outputs. The category was ignited by the Jev model from Typesafe AI, and Cloudflare openly described the market as \"increasingly saturated\" — within one week, the GLiDE model from Fastino, the Clef model from Cloudflare, and the Decider model from Perplexity all shipped, with iteration measured in weeks.\n\n## The scorecard: edges Jev overall, wins on retrieval truth\n\nThe official model card benchmarks 11 tasks: overall 85.71%, slightly ahead of Jev at 84.51% and well above the Qwen3.8-27B base at 74.76%. The per-task pattern is more telling. Decider wins on RAGTruth (88.80% vs 77.27%, an 11-point lead), FinancialPhraseBank, TabFact, and Circa — tasks about judging given material. Jev holds WinoGrande, BBH, JudgeBench, TruthfulQA, and its own JevBench. A search company fine-tuned a model whose strongest suits are exactly retrieval-grounded judgment — data DNA shows. Caveat: these numbers were measured through the Perplexity API and published by Perplexity itself, so treat them as self-reported.\n\n## Open source, with a catch\n\nDecider supports choice and noul output types, returns calibrated probabilities, and can take images for visual decisions. But 26B BF16 weights need roughly 49 GiB plus working memory, with a CUDA GPU and Python 3.12 as the floor; the model card showed only 165 downloads in the last month. In other words, this is a component for engineering teams with GPUs, not a drop-in API replacement.\n\n## So what\n\nDecision models are becoming a distinct layer in the agent stack: routing, classification, guardrails — high-frequency small judgments that no longer need to consult a frontier LLM every time. The Perplexity entry says two things. First, judgment data accumulated by a search company converts directly into model advantage. Second, the window for this category will not stay open long; once API giants bundle decision-making into their platforms, open weights are the entry ticket. If you build agents, ask yourself: how many of your LLM calls are really just multiple-choice questions? Model card on [Hugging Face](https:\u002F\u002Fhuggingface.co\u002Fperplexity-ai\u002Fpplx-decider-v1-27b).","perplexity-decider-v1-27b-open-source","2026-10-02T21:20:00Z","2026-10-02T21:15:16.745603Z","2026-10-02T21:15:16.745613Z",true,"agent",12,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"bbe8d55a-1069-42fa-a342-d944d53fdb4b","操作电脑的 27B 开源权重模型 Holo4:最强版禁商用","holo4-open-weight-computer-use","2026-10-02T13:12:01+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"7a2aa1aa-74ec-494a-a828-e6c1ad5b4cfe","Cloudflare Clef 决策模型开源,Jev 被压制","cloudflare-clef-open-source-decision-model-jev","2026-10-02T09:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"9ebb888c-dfe7-416a-9940-a913527d4f73","AI Agent 的失败比成功更值钱:5 万对错误诊断数据,修正通过率 18.4%→51.1%","agent-error-dataset","2026-10-01T15:11:08+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"dc61debd-1d77-4d5a-9d59-5b23c3da07de","蚂蚁开源Realtime-Venus：9B全双工模型边说边干活，三项续聊指标超GPT-4o","ant-realtime-venus-full-duplex-delegation","2026-09-30T23:10:53+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"e98cf9a2-8348-4142-9a56-c11166774798","SpeakerMem-R1:多方对话记忆,分清谁说了什么","speakermem-r1-multi-party-memory","2026-09-24T19:05:00+00:00"]