[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-microsoft-decision-1-fast-decision-model":3,"topics-all":38,"news-related-40fc3d8d-da30-43cb-a945-79b9c34fbb09":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"40fc3d8d-da30-43cb-a945-79b9c34fbb09","微软决策模型Decision-1:比Sol快35倍,输出免费","微软发布 Microsoft-Decision-1，基于 Qwen3.5-9B 后训练的单遍决策打分模型，对固定选项返回校准概率，用于路由、护栏、AI 评判等 agent 场景。官方 36 项盲测基准称准确率居首，比 GPT-6 Sol 快 35 倍；输入每百万 token 0.042 美元，输出免费。","Agent 工作流里塞满了不起眼的小决策：这条请求该路由给哪个模型、上一步结果要不要重试、异常该不该转人工。过去这些决策要么写成硬编码规则，要么直接丢给大模型——前者脆弱，后者又慢又贵。微软在 10 月 9 日交出了自己的答案：Microsoft-Decision-1，一个不生成文本、专门做判断的小模型。\n\n## 不写文章，只做判断\n\n与 LLM 不同，决策模型（decision model）为结构化输出而生：给定一组固定选项，模型对每个选项返回一个校准过的概率分数，软件可以直接拿这个分数行动。官方列出的适用场景包括模型路由、分类、优先级排序、验证、工作流控制、agent 护栏和 AI 评判——都是 agent 基础设施里高频、低复杂度但量极大的环节。底座是 Qwen3.5-9B，经过后训练做单遍打分；微软表示后续会换用包括 MAI 和 OpenAI 在内的其他底座。模型已在 Microsoft Foundry 和 OpenRouter 上线。\n\n值得玩味的是：微软手里有 Phi 系列，与 OpenAI 的合作也在加深，但第一版选了阿里的 Qwen3.5-9B 当底座——9B 这个量级在延迟和成本上刚好卡在「决策」这个位置。\n\n## 数字都在官方评测里\n\n微软的评测覆盖 36 项对训练保密的基准、近 15 万道题，官方称准确率居参测模型之首；速度上比第二名 H2O-Lightning-4B v1.1 快 2.5 倍，比 GPT-6 Sol 快 35 倍。OpenRouter 独立页面给出参考：P50 延迟 0.27 秒，上下文 32,768 token。\n\n官方用 8 种方式扰动同一请求（改写、重排、格式噪声），平均决策翻转率 1.3%，选项改写或重排时零翻转；安全测试覆盖 5,250 条请求、11 项基准。定价是这类模型的杀手锏：输入每百万 token 0.042 美元，输出免费。\n\n## 微软自己就是第一个客户\n\n最有说服力的是内部案例。Xbox Research 用它处理超过 1 万条开放反馈和评测，质量与 GPT-6 Sol 相当，速度快 14 倍、成本低 200 倍；Copilot 团队用它评估对话和 agent 回复质量，与 GPT5.6 Luna 质量相当、快 100 倍；Microsoft Discovery 的自适应重规划用它打分，一致性比 LLM 评分高 46 倍。微软是先算清了自家的账，才把这个模型拿出来卖。\n\n## 两周四家入场，品类成立\n\n过去两周 Cloudflare 开源了 Clef，亚马逊放出了 Strands Decider，Perplexity 发布 pplx-decider，加上 9 月底的 Firelex Jeff 和今天的微软，「决策模型」正在从论文概念变成 agent 栈的独立层。逻辑并不复杂——LLM-as-judge 的成本和延迟痛点是真实存在的，当 agent 每天要做上百万次小判断，「用 GPT-6 判断要不要调用 GPT-6」就成了笑话。生成和判断正在分工：大模型负责难的，小模型负责快的。\n\n真正的看点：当判断本身便宜到可以随处嵌入，agent 的护栏、路由和质检会从「事后补丁」变成「出厂内置」。（官方原文地址：https:\u002F\u002Fcommandline.microsoft.com\u002Fmicrosoft-decision-1-model-foundry\u002F ；OpenRouter 模型页：https:\u002F\u002Fopenrouter.ai\u002Fmicrosoft\u002Fmicrosoft-decision-1 ）\n","https:\u002F\u002Fcommandline.microsoft.com\u002Fmicrosoft-decision-1-model-foundry\u002F","144f3573-80d9-4d33-b645-ea97eabc398b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"2a312f74-acb9-48a0-8a01-9dc5955e3c44","en","Microsoft Decision-1: 35x faster than Sol, output free","Microsoft's Decision-1: calibrated decision scoring for agent routing and guardrails. Qwen3.5-9B, topped 36 benchmarks, 35x faster vs Sol at $0.042\u002FM.","Agent workflows are full of small, thankless decisions: which model should handle this request, should the previous step be retried, and whether an anomaly deserves a human. Microsoft's answer, published October 9, is Microsoft-Decision-1 — a small model that generates no text and exists to judge.\n\n## It doesn't write, it decides\n\nUnlike LLMs, decision models are purpose-built for structured output: given a fixed set of answer options, the model returns a calibrated probability for each, and software can act on that number directly. Microsoft's official use-case list covers model routing, classification, prioritization, verification, workflow control, agent guardrails and AI judging — the high-frequency, low-glamour plumbing of agent infrastructure. The base is Qwen3.5-9B, post-trained for single-pass scoring; Microsoft says it will rebase the model on others, including MAI and OpenAI models. It is available now on Microsoft Foundry and OpenRouter.\n\nOne detail worth savoring: Microsoft owns the Phi family and partners closely with OpenAI, yet picked Alibaba's Qwen3.5-9B as the first base — the 9B class happens to sit right at the latency and cost sweet spot for decisions.\n\n## The numbers, as officially measured\n\nMicrosoft evaluated the model across 36 benchmarks kept blind from training, nearly 150,000 questions in total, and reports the highest accuracy among the models it tested. On speed it measured 2.5x quicker than H2O-Lightning-4B v1.1 (the runner-up) and 35x quicker than GPT-6 Sol at P50 latency. OpenRouter's independent model page lists a 0.27s best-provider P50 and a 32,768-token context window.\n\nOn robustness, the team perturbed identical requests in eight ways; decisions flipped on 1.3% of perturbations on average, with zero flips when option descriptions were paraphrased or options reordered. Safety testing covered 5,250 requests across 11 benchmarks. Pricing is the category's signature weapon: $0.042 per million input tokens, output free.\n\n## Microsoft is its own first customer\n\nThe most persuasive part of the announcement is the internal case work. Xbox Research used the model to process more than 10,000 pieces of open-ended feedback and reviews — competitive in quality with GPT-6 Sol while running over 14x faster and 200x cheaper. The Copilot team uses it to grade chat and agentic responses, competitive with GPT5.6 Luna at 100x the speed. Microsoft Discovery's adaptive replanning scored it 46x more consistent than the LLM-based score. In other words, Microsoft did its own accounting first, then took the model to market.\n\n## Four vendors in two weeks: a category forms\n\nZoom out: in the past two weeks Cloudflare open-sourced Clef, Amazon released Strands Decider, and Perplexity shipped pplx-decider — add Firelex's Jeff from late September and now Microsoft, and \"decision models\" are graduating from paper concept to a distinct layer of the agent stack. The logic is straightforward: the cost and latency pain of LLM-as-judge is real, and once an agent makes a million small calls a day, \"using GPT-6 to decide whether to call GPT-6\" becomes a joke. Generation and judgment are splitting up: big models handle the hard, small models handle the fast.\n\nThe real thing to watch is what comes next: when judgment becomes cheap enough to embed everywhere, guardrails, routing and quality control shift from afterthoughts to factory defaults. (Official announcement: https:\u002F\u002Fcommandline.microsoft.com\u002Fmicrosoft-decision-1-model-foundry\u002F ; OpenRouter model page: https:\u002F\u002Fopenrouter.ai\u002Fmicrosoft\u002Fmicrosoft-decision-1 )\n","microsoft-decision-1-fast-decision-model","2026-10-10T15:10:00Z","2026-10-10T15:11:22.812991Z","2026-10-10T15:11:22.813002Z",true,"agent",41,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"206ea36a-1eea-463f-a241-1e3b32f5ec2d","USTC GraphForge:证据图把任务和 rubric 钉在一起,Qwen3.6-27B 涨三基准","graphforge-ustc-qwen36-27b-evidence-graph","2026-10-04T03:05:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"983fb4d1-6c63-4828-adbf-a59c57e02b64","Perplexity开源决策模型:总分微胜Jev","perplexity-decider-v1-27b-open-source","2026-10-02T21:20:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"c4ec4625-4a84-4c24-88f0-0ef1beb4f19e","Grok 4.6 发布:61 分追平 GPT-5.6 Sol,把长程 Agent 的 token 账单砍到四分之一","grok-4-6-agentic-cost-frontier","2026-08-14T19:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"40210e0d-84e3-460b-bde2-295b77573ab8","Qwen3.7-Text-Embedding 上线:20% 检索增益、256-2560 可变维度,阿里把 RAG 的地基悄悄重浇了一遍","qwen3-7-text-embedding-launch","2026-08-14T13:10:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"3cc63477-1334-497d-80cb-90850c019101","DeepSeek-V4-Flash 转正:不靠换架构,只做后训练重新发力 Agent","deepseek-v4-flash-official-post-training-agent-0731","2026-07-31T08:00:00+00:00"]