[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-agent-error-dataset":3,"topics-all":41,"news-related-9ebb888c-dfe7-416a-9940-a913527d4f73":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"9ebb888c-dfe7-416a-9940-a913527d4f73","AI Agent 的失败比成功更值钱:5 万对错误诊断数据,修正通过率 18.4%→51.1%","Heng Ji 团队发布 Agent Error Dataset:5 万余对错误-诊断记录覆盖 33 个环境与 23 个策略模型,五阶段流水线定位错误步骤并用重放验证修正。对照中修正通过率 18.4% 升至 51.1%,微调后 Qwen3-8B 诊断一致率 63.6%,超过提示参考 54.7%。","大家训练 agent 时都在拼命攒成功轨迹——跑对的任务留着蒸馏,跑错的直接丢掉。9 月 30 日挂上 arXiv 的 Agent Error Dataset(编号 2609.40111)反着来:把 50,228 对「错误-诊断」记录整理成数据集,专门研究失败轨迹怎么变成训练资产。论文开头一句话点题:一次失败的 rollout 所包含的信息,比它的最终 reward 多得多——观测、动作、环境响应,全都是可以复盘的证据。\n\n## 数据集里有什么\n\n规模数字相当扎实:50,228 对错误-诊断记录,来自 9,961 个源任务,覆盖 33 个环境、19 类 harness 框架、23 个策略模型,全部是文本交互式 agent 系统。去重后有 38,278 条独立源轨迹,平均每条带 1.31 份诊断——同一错误可以换视角反复重新诊断,原始 rollout 不用重跑。数据集还保留完整的执行元数据,支持跨设置做失败分析与再诊断。\n\n## 五阶段流水线怎么造数据\n\n配套的 AET(Agentic Error-to-Training)流水线分五步:先收集自然发生的失败,把基础设施故障和评分器 bug 剔出去,只留真正的 agent 错误;然后给每个错误定位失败步骤和责任方,诊断解释必须引用 trace 里的原文证据;第三步做结构与语义双重校验,被拒的提案也留档;第四步是重放——在同一个 checkpoint、同样的策略、预算与验证器设置下,让「执行修正方案」和「原动作重试」打对照;最后按诊断 SFT、恢复 SFT、动作偏好三类训练视图分别组织数据。\n\n## 数字说话\n\n对照重放共 3,062 组:首轮修正方案把验证器通过率从 18.4% 拉到 51.1%,净增 32.7 个百分点——同一个 checkpoint,只换成修正过的动作,通过率就是原来的 2.8 倍。用冻结的诊断版本微调 Qwen3-8B,模型对错误步骤的定位一致率从 47.2% 涨到 63.6%(三个种子平均,943 例留出集),而最强的提示词参考模型只有 54.7%。行为侧,只训修复动作比只训成功轨迹在 WebShop-lite 上高 6.67 个百分点;定位辅助的续跑通过率 31.3%,也高于泛泛重想的 26.0% 和纯重放的 18.6%。\n\n## 冷水也得泼\n\n第一,数据和代码还没放出来,论文写明「接收后」再开源代码、checkpoint 和数据,想上手的团队现在只能读论文;第二,行为侧 6.67 个百分点是单 seed 结果,作者自己标注了;第三,通讯作者 Heng Ji 署名机构是 Apodex,论文致谢了 Tianqiao Chen 的指导支持——这条产学路线的后续值得留意。\n\n对行业的信号很直接:当大家把算力砸在合成成功数据上时,失败轨迹可能是被系统性丢弃的免费矿。修正方案 2.8 倍于原动作的通过率说明,「会认错」本身就是一种可训练的能力。问题留给读者:你的 agent 上线后,那些跑挂的 trace,现在存在哪里?\n\n论文:https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.40111","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.40111","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":19,"name":20,"slug":20,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"789ea885-7ea8-4756-bd79-9fa10c1a7af7","en","AED: 50K Agent Error Pairs Turn Failures Into Training Data","AED packs 50,228 error-diagnosis pairs from 33 environments and 23 policy models; corrections lift verifier pass rates from 18.4% to 51.1%.","While most teams stockpile successful rollouts for distillation and throw failures away, a paper posted on arXiv on September 30 (2609.40111) argues the opposite: the Agent Error Dataset (AED) turns 50,228 error-diagnosis pairs into a reusable training asset. The opening line makes the thesis explicit — an unsuccessful rollout carries far more information than its final reward: the observations the agent saw, the actions it took, and the environment's responses are all reviewable evidence.\n\n## What is inside\n\nThe scale numbers are solid: 50,228 error-diagnosis pairs drawn from 9,961 source tasks, spanning 33 environments, 19 harness families and 23 policy models in text-based agent systems. After deduplication the collection holds 38,278 distinct source traces, averaging 1.31 diagnoses per trace — the same failure can be re-diagnosed from new angles without re-running the original rollout. Full execution metadata is retained for cross-setting failure analysis.\n\n## The five-stage pipeline\n\nThe companion AET (Agentic Error-to-Training) pipeline works in five steps. First it collects naturally occurring failures while screening out infrastructure faults and grader bugs, keeping only genuine agent errors. Second, each error gets a diagnosis: the failing step and the responsible agent, with an explanation that must cite the trace. Third, diagnoses pass structural and semantic grounding checks; rejected proposals are kept on record. Fourth comes replay — from the same checkpoint, under matched policy, harness, budget and verifier settings, the proposed correction races against a fresh original-action retry. Fifth, the data is organized into training views for diagnosis SFT, recovery SFT and action preferences.\n\n## The numbers\n\nAcross 3,062 matched replay pairs, first-proposal corrections lift verifier pass rates from 18.4% to 51.1% — a 32.7-point gain, roughly 2.8x for the same checkpoint. Fine-tuning Qwen3-8B on a frozen diagnosis release raises exact-step agreement with internal teacher labels from 47.2% to 63.6%, beating the strongest prompted reference at 54.7%. On the acting side, action-only repair training scores 6.67 points higher than success-only training on WebShop-lite (single seed), and diagnosis-guided continuation reaches 31.3% versus 26.0% for generic reconsideration and 18.6% for plain replay.\n\n## Caveats worth noting\n\nThe code, checkpoints and data are not out yet — the paper states they will be open-sourced after acceptance, subject to licenses and privacy review. The 6.67-point actor result is explicitly labeled single-seed. Corresponding author Heng Ji is affiliated with Apodex, and the paper acknowledges Tianqiao Chen for guidance and support — worth watching where this industrial-academic line goes next.\n\nThe industry signal is direct: while compute pours into synthetic success data, failure traces may be the free ore everyone is throwing away. A 2.8x pass-rate jump from corrected actions shows that knowing how to admit mistakes is itself a trainable capability. One question to leave you with: after your agents crash, where do those traces go now?\n\nPaper: https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.40111","agent-error-dataset","2026-10-01T15:11:08Z","2026-10-01T15:12:03.801171Z","2026-10-01T15:12:03.801182Z",true,"agent",220,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"e98cf9a2-8348-4142-9a56-c11166774798","SpeakerMem-R1:多方对话记忆,分清谁说了什么","speakermem-r1-multi-party-memory","2026-09-24T19:05:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"3bcb0e1d-99ea-4fae-9bdd-b6b625aabf10","代码 agent 8 成都在骗你:12 模型实测揭晓","overclaimbench-llm-agents","2026-09-21T07:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"089195fb-7fe5-4ba9-a4bc-8e356fe5e923","BAAI把1000个GitHub仓库蒸馏成5000个技能,科研agent奖牌率31%冲到73%","baai-disco-repo-to-skill-library","2026-09-03T17:07:35+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"e8965513-b56f-475b-b15f-22a5ea2d2a4e","Agent 取代人成为 HF Hub 一号用户:Claude Code 占 44.4%,还有一次 4.5 天未察觉的入侵","hf-hub-agent-user-claude-code-4-5-day-intrusion","2026-08-21T08:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"86565410-ced9-4e64-80b7-d97a353a1d1d","LLM 元认知首次可度量：Metacognition-Bench 用 300 道陷阱题 + 11 个开源错题雷达适配器把自我纠错做成开源工程","metacognition-bench-llm-self-correction","2026-07-01T12:00:00+00:00"]