[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tsinghua-agenticdatabench":3,"topics-all":36,"news-related-eac81652-8158-4b38-a1f5-7ba9420fb74e":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"eac81652-8158-4b38-a1f5-7ba9420fb74e","清华 AgenticDataBench：把 LLM 数据智能体拉进「真实业务」的统考卷","数据科学自动化的诱惑讲了十年,但真正能替数据科学家\"读脏数据、做特征、出报告\"的 LLM Agent,仍缺少一把公开的尺子。清华大学等机构刚发布的 AgenticDataBench (arXiv:2607.01647),试图补齐这把尺子。\n\n它和传统评测最大的区别,是引入了**\"数据科学技能(skill)\"**作为中间粒度:从 Stack Overflow 大规模任务解法里抽取 433 种操作模式——缺失值插补、时序重采样、异常校验等——再用技能对齐的层次聚类去冗余,最终组成 344 个任务、97 个数据集、27.3 GB 数据,覆盖 15 个垂直领域,包含一家头部金融科技公司的 5 个真实 B2B 业务流。\n\n另一亮点是**任务合成管线**——对缺乏真实数据的领域,作者用 LLM 围绕\"技能组合\"反向合成任务与标准答案,避免 benchmark 过度偏向金融、电商等常见热点。\n\n评测结果未在摘要中披露,但作者开源了测试台与 GitHub repo,给社区一个可复现入口。这条路线对国产 Agent 框架尤其关键——以前大家都只能在自家准备的几道示例题上自吹自擂,现在终于有第三方\"统考卷\"可以上分。\n\n**【观点】**AgenticDataBench 的真正价值或许不在\"哪家模型跑分第一\",而在于把\"数据科学技能\"这个抽象词变成**可枚举、可测试、可教学**的对象。当技能库公开后,做垂域 Agent 的团队可以反向挑选训练数据、针对性补齐短板——这或许是 LLM for Data Science 走向工程化的第一块拼图。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.01647","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d8b8ad76-6f47-41d8-93a0-32be80bb0264","en","AgenticDataBench: LLM data agents meet real business tasks","The temptation of data-science automation has been around for ten years, but LLM Agents that can truly replace data scientists to \"read dirty data, build features, produce reports\" still lack a public ruler. AgenticDataBench (arXiv:2607.01647), just released by Tsinghua University and other institutions, tries to fill this ruler. The biggest difference from traditional evaluation is the introduction of **\"data-science skills\"** as the intermediate granularity: 433 operational patterns are extracted from Stack Overflow's large-scale task solutions — missing-value imputation, time-series resampling, anomaly validation, etc. — and then skill-aligned hierarchical clustering removes redundancy, finally composing 344 tasks, 97 datasets, 27.3 GB of data, covering 15 vertical domains, including 5 real B2B business flows from a top fintech company. Another highlight is the **task synthesis pipeline** — for domains lacking real data, the authors use LLM to reverse-synthesize tasks and standard answers around \"skill combinations\", avoiding the benchmark being overly biased toward hot domains like finance and e-commerce. The evaluation results aren't disclosed in the abstract, but the authors have open-sourced the testing platform and GitHub repo, giving the community a reproducible entry point. This path is especially critical for domestic Agent frameworks — previously everyone could only self-promote on the few example questions they prepared at home, now there's finally a third-party \"exam\" to compete on. **[Opinion]** AgenticDataBench's real value may not lie in \"which model ranks first\", but in making the abstract phrase \"data-science skills\" into an **enumerable, testable, teachable** object. Once the skill library is public, teams building vertical Agents can reverse-select training data and target weak spots — this may be the first piece of the LLM for Data Science puzzle heading toward engineering.","tsinghua-agenticdatabench","2026-07-03T08:00:00Z","2026-07-06T00:09:25.152409Z","2026-08-19T02:08:40.142862Z",true,"agent",242,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"9822a1a7-0014-4bd5-bbe0-492401fe6b96","AllSpark 把搜索 Agent 推到 BrowseComp 88.6:SFT-RL Climbing 与推理时上下文管理","allspark-iris-search-agent-sft-rl-climbing","2026-09-07T07:11:17+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"00346b75-f071-42fd-ae16-db4c5569f01a","EarlyEval 提前叫停注定失败的 Agent:近半 token 省下,分辨率只动一两个点","earlyeval-early-stop-agent-eval","2026-09-03T21:04:52+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"14a7f5ab-e270-461c-b862-4bde139e463f","HarnessDev 基准:让 LLM 自建 Agent Harness,代码领域仍输人类工程师","harnessdev-llm-selfbuilt-agent-harness","2026-09-03T19:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"089195fb-7fe5-4ba9-a4bc-8e356fe5e923","BAAI把1000个GitHub仓库蒸馏成5000个技能,科研agent奖牌率31%冲到73%","baai-disco-repo-to-skill-library","2026-09-03T17:07:35+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"d1e8997e-bb60-453d-9ef8-71b8bdde5386","Harvey 首个自研法律模型 Tenet 曝光:底座没选 GPT 和 Claude,选了 Kimi K3","harvey-tenet-kimi-k3-legal-model","2026-08-18T17:30:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"c4ec4625-4a84-4c24-88f0-0ef1beb4f19e","Grok 4.6 发布:61 分追平 GPT-5.6 Sol,把长程 Agent 的 token 账单砍到四分之一","grok-4-6-agentic-cost-frontier","2026-08-14T19:00:00+00:00"]