[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-taste-bench-agent-decision-forks":3,"topics-all":35,"news-related-832948d8-aa11-4589-b406-07ffc09eccd0":54},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"832948d8-aa11-4589-b406-07ffc09eccd0","微软Taste-Bench:502个决策岔口,最强模型也只答对59.7%","微软团队从智能体真实任务轨迹里自动挖出502个决策岔口,建成Taste-Bench基准:模型在看不到后续结果时选方向,最强模型正确率只有59.7%,加大推理预算也无济于事;而蒸馏事后判断能把27B学生模型从30%拉到47.9%。","长程任务里,智能体每天都要做同一个动作:岔口选路。测哪个假设、沿用哪份实现、走哪条修复路径——这些中途决策决定整次运行的成败。微软团队把这种「在结果揭晓前挑对方向」的能力叫作品味(taste),并在论文《The Tasteful Agent》([arXiv:2609.25804](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25804))里发布了专门测量它的基准 Taste-Bench,目前已登上 Hugging Face 每日论文榜第一,61 个赞。\n\n## 品味测的是什么\n\nTaste-Bench 的题目形态很克制:给模型任务描述、岔口之前的完整轨迹、两个候选下一步,让它选哪个更好。关键在于「隐藏的另一半轨迹」会证明哪条路对——选错的那个方向在当时往往看起来也合理,代价要花掉后面大半预算才显现。论文评了 14 个前沿模型,最好的 GPT-5.6 Sol 正确率只有 59.7%,而随机猜测是 25%。旗舰模型在真岔口上的判断力,离靠谱还差得远。\n\n更扎心的是两个反直觉发现。其一,决定性证据在轨迹里出现得越晚,模型越抓瞎:准确率从 62.3% 一路掉到 21.0%。其二,更大的推理预算救不回来——给模型更多思考 token,选路准确率没有改善。品味似乎不是「想得更久」能换来的东西。\n\n## 502 道题是怎么来的\n\n这个基准最漂亮的地方是不需要人工标注。作者从 SWE-bench、SWE-bench Pro 和 METR 的 MALT 公开轨迹(RE-Bench 与 HCAST)里挖岔口:同一任务的多次独立尝试在同一个点分叉、结局不同,这是平行岔口;单条轨迹里智能体走错方向、撞墙后折返恢复,这是弯路岔口。轨迹的后半段天然为前半段的决策提供事后标签。\n\n过滤也够狠:所有裁判模型仅凭选项措辞就能答对的题,当作平凡题剔除;任何一个裁判读完全记录后不同意标签的题,当作不可判定剔除。4,657 个挖出的岔口最后只活下来 502 个:工程域 390 题,研究域 112 题。评测协议还防位置偏好:每道题按原始选项顺序和完全反转的顺序各问一次,两次都答对才算对,恒选同一个位置的模型直接得零分。作者在社区帖里给出标签与人工审查的一致率:98.8%。\n\n## 品味可以练出来\n\n榜单本身信息量不小:GPT-5.6 Sol(59.7)与 GPT-5.5(59.5)并列头部,Claude Opus 5 是 55.5,GLM-5.2 拿到 53.9,DeepSeek V4 Flash 43.3,垫底的 Grok 4.20 Reasoning 只有 15.7——它的 1,004 次请求里有 459 次输出无法解析,按协议全算错。\n\n但论文真正的钩子是:品味可训练。作者把「看过结局的教师」的判断蒸馏给学生模型,Qwen3.6-27B 在未见过的任务上从 30.0% 提到 47.9%;再进一步,让这个学生给智能体当顾问,SWE-bench Pro 的端到端成功率从 14.6% 翻到 33.7%。判断力可以从 hindsight 里蒸馏出来,再反哺给运行中的智能体。\n\n## 所以呢\n\n过去两年大家卷的是智能体能不能跑完全程,基准也都盯着端到端成功率。Taste-Bench 把镜头对准了中途的每一次转向,而数据说明这恰恰是当前最薄的一块短板——最强模型正确率六成都不到,且砸推理预算没用。对做智能体工程的人来说,「在岔口上停下来,问问一个蒸馏过 hindsight 的顾问」可能比换更大的模型更划算。基准代码 MIT、数据 CC BY 4.0,数据集做了门控以防训练污染,完整跑一次约 1,004 次请求、800 万输入 token。值得跑一遍自家模型,看看它的品味值几分。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25804","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"7d6f1559-28dc-4383-9150-daecd4e774c9","en","Taste-Bench: Top Model Gets Only 59.7% at Agent Decision Forks","Taste-Bench tests if models pick the right path at 502 real agent forks. Best score 59.7%; distilling hindsight lifts a 27B student to 47.9%.","Long-horizon agents keep making the same move: picking a path at a fork. Which hypothesis to test, which implementation to build on, which fix to attempt — these mid-run decisions decide the whole trajectory. A Microsoft team calls this ability to pick the better direction before the outcome shows \"taste,\" and they just released a benchmark that measures it: Taste-Bench, in the paper \"The Tasteful Agent\" ([arxiv.org\u002Fabs\u002F2609.25804](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25804)), now #1 on Hugging Face Daily Papers with 61 upvotes.\n\n## What taste means here\n\nEach question is deliberately spare: the model gets the task, the full trajectory up to a decision fork, and two candidate next steps. The hidden rest of the trajectory proves which one is right. A wrong choice often looks reasonable in the moment and burns most of the budget later. The authors evaluated 14 frontier models; the best, GPT-5.6 Sol, answers only 59.7% correctly, against 25% for random guessing. Flagship judgment at real forks is nowhere near reliable.\n\nTwo counterintuitive findings stand out. First, the later the deciding evidence appears in the trajectory, the worse models do: accuracy falls from 62.3% to 21.0%. Second, a larger reasoning budget does not rescue it — more thinking tokens do not improve choice accuracy. Taste is not something a model earns by thinking longer.\n\n## How the 502 questions were built\n\nThe benchmark needs no human annotation. The authors mined forks from SWE-bench, SWE-bench Pro, and METR's MALT release of RE-Bench and HCAST trajectories: parallel attempts at the same task that diverge at the same point with different recorded outcomes, and detours inside a single run where the agent abandons a direction after an observed failure and recovers. The later part of a trajectory labels the earlier decision for free.\n\nFiltering is strict: a question is dropped as trivial when every judge model answers it from the candidate wording alone, and as undecidable when any judge disagrees with its label after reading the full record. Of 4,657 mined forks, 502 survive — 390 engineering and 112 research questions. The protocol also blocks position bias: every question is asked in the published order and in its exact reverse, and only double-correct answers count; a model that always picks the same position scores 0. The authors report 98.8% agreement with human review.\n\n## Taste is trainable\n\nThe leaderboard itself is informative: GPT-5.6 Sol (59.7) and GPT-5.5 (59.5) lead, Claude Opus 5 sits at 55.5, GLM-5.2 at 53.9, DeepSeek V4 Flash at 43.3, and Grok 4.20 Reasoning bottoms out at 15.7 — 459 of its 1,004 responses were unparseable and count as wrong.\n\nThe real hook: taste can be trained. The authors distilled the judgment of a teacher that had seen the outcome into a student; Qwen3.6-27B goes from 30.0% to 47.9% on unseen tasks. Let that student advise an agent, and end-to-end success on SWE-bench Pro jumps from 14.6% to 33.7%. Judgment can be distilled from hindsight and fed back into running agents.\n\n## So what\n\nFor two years the field has optimized whether agents finish the course, and benchmarks have watched end-to-end success. Taste-Bench points the camera at every mid-run turn, and the data says that is exactly today's thinnest weak spot — the best model is under 60%, and brute-force reasoning budgets do not help. For agent engineers, pausing at a fork to ask a hindsight-distilled advisor may beat swapping in a bigger model. Code is MIT, data CC BY 4.0, with a gated dataset release to limit training contamination; a full run is 1,004 requests and about 8M input tokens. Worth running against your own model to see how much taste it has.","taste-bench-agent-decision-forks","2026-09-23T13:10:11Z","2026-09-23T13:10:14.998020Z","2026-09-23T13:10:14.998030Z",true,"agent",330,[36,45],{"slug":37,"tag_slug":37,"title_zh":38,"title_en":39,"intro_zh":40,"intro_en":41,"id":42,"is_active":32,"created_at":43,"modified_at":44},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":32,"created_at":52,"modified_at":53},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":55},[56,61,66,71,76,81],{"id":57,"title":58,"news_slug":59,"published_at":60},"4f5c4072-660c-47e7-9d9a-9782eb5591bf","Pistis 报告:IDRL 让蒸馏和 RL 交替上岗","pistis-idrl-interleaved-distillation-rl","2026-09-25T19:11:33+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"e98cf9a2-8348-4142-9a56-c11166774798","SpeakerMem-R1:多方对话记忆,分清谁说了什么","speakermem-r1-multi-party-memory","2026-09-24T19:05:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"3bcb0e1d-99ea-4fae-9bdd-b6b625aabf10","代码 agent 8 成都在骗你:12 模型实测揭晓","overclaimbench-llm-agents","2026-09-21T07:00:00+00:00",{"id":72,"title":73,"news_slug":74,"published_at":75},"9822a1a7-0014-4bd5-bbe0-492401fe6b96","AllSpark 把搜索 Agent 推到 BrowseComp 88.6:SFT-RL Climbing 与推理时上下文管理","allspark-iris-search-agent-sft-rl-climbing","2026-09-07T07:11:17+00:00",{"id":77,"title":78,"news_slug":79,"published_at":80},"00346b75-f071-42fd-ae16-db4c5569f01a","EarlyEval 提前叫停注定失败的 Agent:近半 token 省下,分辨率只动一两个点","earlyeval-early-stop-agent-eval","2026-09-03T21:04:52+00:00",{"id":82,"title":83,"news_slug":84,"published_at":85},"14a7f5ab-e270-461c-b862-4bde139e463f","HarnessDev 基准:让 LLM 自建 Agent Harness,代码领域仍输人类工程师","harnessdev-llm-selfbuilt-agent-harness","2026-09-03T19:10:00+00:00"]