[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-paperpilot-workflow-induction":3,"news-related-1316635c-88e1-41b6-a45c-df8ef217cf3f":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1316635c-88e1-41b6-a45c-df8ef217cf3f","PaperPilot 把文献搜索改写成「工作流归纳」：可编辑 DAG 把多轮检索错误率干到 0%","多轮文献检索的痛点在于「用户在迭代、agent 在即兴」。既有系统要么把流程藏进 chain-of-thought,要么套用固定 pipeline,既难调试也难对齐用户偏好。UIUC、宾大、斯坦福、Together AI 联合提出的 PaperPilot(arXiv:2607.00597)把这事儿重新定义为「工作流归纳」:给定锚点论文与查询,模型自动构造一个可执行 DAG——把关键词搜索、引文扩展、过滤、评分、重排、证据抽取当成节点化算子串起来;用户反馈不是「再跑一次」,而是直接对工作流做局部修正,查询和工作流一起迭代。\n\n训练上分两步:先用监督工作流模仿学高质量轨迹,再叠一层「受控工作流腐蚀」的偏好优化,让模型主动避开会跑崩的分支。在 Qwen3.5-9B 基础上训出的 PaperPilot-9B,多轮交互下 Hit@5 从 58.0 提到 77.0(+19pp),MRR 从 47.5 提到 59.4,nDCG@10 从 26.8 提到 32.5,最关键的一项是工作流执行错误率从 9.5% 直接降到 0%——对每天要做几十轮检索的研究者来说,工作流跑崩一次等于一次返工,这是数量级的可用性提升。\n\n更深远的意义在于「工作流即接口」。它把 agent 的能力边界从「自然语言 prompt」扩展到「可审计、可版本控制、可人工干预的检索流水线」,跟 agentic RL、tool-use 标准化的大趋势一脉相承。对企业知识库、法务检索、医学综述等需要可追溯检索路径的场景,PaperPilot 给出了一条工程化的范式参考。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.00597","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"01377e6f-8e0b-4eef-a85c-b64d42a8ec29","en","PaperPilot: editable DAGs take retrieval errors to zero","The pain point of multi-turn literature search is \"the user is iterating, the agent is improvising\". Existing systems either hide the flow in chain-of-thought, or apply a fixed pipeline, both hard to debug and hard to align with user preferences. PaperPilot (arXiv:2607.00597) jointly proposed by UIUC, Penn, Stanford, and Together AI redefines this as \"workflow induction\": given an anchor paper and query, the model automatically constructs an executable DAG — stringing keyword search, citation expansion, filtering, scoring, reranking, evidence extraction as node-level operators; user feedback isn't \"run again\", but directly locally editing the workflow, with the query and workflow iterating together. Training is in two steps: first use supervised workflow imitation to learn high-quality trajectories, then layer on a \"controlled workflow corruption\" preference optimization, letting the model actively avoid branches that would crash. PaperPilot-9B trained on Qwen3.5-9B improves Hit@5 from 58.0 to 77.0 (+19pp), MRR from 47.5 to 59.4, nDCG@10 from 26.8 to 32.5, with the most critical being workflow execution error rate dropping directly from 9.5% to 0% — for researchers doing dozens of rounds of retrieval a day, a workflow crash means a re-do, this is an order-of-magnitude usability improvement. A deeper significance lies in \"workflow as interface\". It extends the Agent's capability boundary from \"natural language prompt\" to \"auditable, version-controllable, human-intervenable retrieval pipeline\", in line with the broader trends of agentic RL and tool-use standardization. For scenarios like enterprise knowledge bases, legal retrieval, and medical reviews that need traceable retrieval paths, PaperPilot gives an engineering paradigm reference.","paperpilot-workflow-induction","2026-07-01T08:21:23Z","2026-07-05T08:09:29.557545Z","2026-08-19T02:08:40.142862Z",true,"agent",97,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"6de305c2-91b2-47fb-b1e4-dfb5f1e711c8","WorldEvolver：把世界模型装进 LLM Agent 的「即时记忆」","worldevolver-llm-agent-world-model","2026-06-30T18:04:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"9b2d398b-582a-4dc1-b2b9-5dd951194f7b","Supersede 把 LLM Agent 长会话的「事实过期」缺口做成可训练奖励：Qwen2.5-3B 上 GRPO 让准确率近翻倍","supersede-fact-staleness-rl","2026-06-29T22:01:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"5b909019-b85f-4ec2-9d7a-9b8808db49e1","MRAgent：NUS 把 LLM Agent 记忆从「查字典」改成「拼拼图」，单查询 token 直降 27 倍","mragent-nus-memory-jigsaw-27x","2026-06-28T10:09:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b88a7d4b-8f6d-4440-b384-4283f88a410c","Tool-Use RL 为什么会突然崩盘？arXiv 2606.26027 戳破 Agent 训练的'概率尖峰'陷阱","tool-use-rl-collapse-probability-spike","2026-06-25T20:25:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"f96a7b02-7bad-4ef2-98c4-a6b6aedefd0c","Constraint Tax：Tool Calling 遇 JSON Schema 悄悄失灵","constraint-tax-tool-calling-silent-disable","2026-06-25T14:00:00+00:00"]