[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-pku-dataprep-bench-das":3,"news-related-38fe9093-827f-43d1-8350-7cdd391cf1e3":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"38fe9093-827f-43d1-8350-7cdd391cf1e3","北大 DataPrep-Bench 把 LLM 当数据准备工来打分：DAS 评估器把「训练价值」算成分布距离","训练数据决定模型上限,但怎么测「LLM 准备数据」的能力一直是空缺。北京大学联合港中文、浙大等机构发布的 DataPrep-Bench 第一次给这件事立了标杆:把数据准备拆成两件并行的事——用 LLM\u002FAgent 从原始语料构造监督数据,以及用一个评估器预测候选数据集的「下游训练价值」,并在 6 个领域、多个基座模型上做端到端联合打分。\n\n论文同时开源两件工具:**Data-Construction-Skill** 技能导向 Agent,在 Llama-3.1-8B Finance 任务上比仅用 Dolly-15k 的基线高近 20 个绝对点,在知识抽取密集型领域也能和最强 Agent \u002F DataFlow 类方法打平;**DAS(Distributional Alignment Score)** 用候选集与领域代理之间的 MMD 距离衡量「训练价值」,在 6 个领域中的 4 个拿到跨模型最强相关,并且是唯一同时在 Math、Science、Medical 三个领域都把 r 拉到 0.7 以上的指标,把既有 quality \u002F diversity \u002F heuristic 评估器全部压在身后。\n\n真正的「so what」在于:这条赛道过去拼的是「数据多」,DataPrep-Bench 把评价尺度从「句子像不像」换成了「下游训练涨不涨分」,DAS 把这件事从专家经验变成可计算的分布对齐指标,意味着数据准备第一次进入了「可被系统性比较、可被规模化替代」的状态——对 LLM 工厂而言,数据团队的工程化拐点已经发生。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.20465","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"cc8b4d64-2538-49d0-8627-4682bfd55457","en","DataPrep-Bench: LLMs graded as data-prep workers","Training data sets the ceiling for a model, but measuring \"an LLM's data-preparation ability\" has long been a gap. A joint release by Peking University, CUHK, Zhejiang University and others has now set a benchmark with DataPrep-Bench: it splits data preparation into two parallel tracks — using an LLM\u002FAgent to construct supervised data from raw corpora, and using an evaluator to predict the \"downstream training value\" of a candidate dataset — and runs end-to-end joint scoring across 6 domains and multiple base models. The paper also open-sources two tools: Data-Construction-Skill, a skill-oriented Agent that improves over a Dolly-15k-only baseline by nearly 20 absolute points on the Llama-3.1-8B Finance task and matches the strongest Agent \u002F DataFlow-style methods on knowledge-extraction-heavy domains; and DAS (Distributional Alignment Score), which measures \"training value\" via the MMD distance between a candidate set and a domain proxy, taking the strongest cross-model correlation on 4 of 6 domains and being the only metric that simultaneously pulls r above 0.7 on Math, Science, and Medical — leaving all existing quality \u002F diversity \u002F heuristic evaluators behind. The real \"so what\": this track has always been about \"more data\", and DataPrep-Bench shifts the evaluation axis from \"do sentences look right\" to \"does downstream training gain\", turning the question from expert gut-feel into a computable distribution-alignment metric. For the first time, data preparation has entered a \"systematically comparable, scalable-substitutable\" state — and for LLM factories, the engineering inflection point in the data team has already happened.","pku-dataprep-bench-das","2026-07-27T22:00:00Z","2026-07-28T02:06:37.907238Z","2026-08-19T02:08:40.142862Z",true,"agent",107,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"5a274662-0c3f-492c-a0e8-a46c5a0be783","别再自己给自己打分了:Co-RL 让模型互相判卷,无标签 RL 追平有监督","co-rl-peer-reward-label-free-rl","2026-08-20T19:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5b76d928-8896-4bf8-8edd-e195ecf0094a","Ornith-1.5自报跑分赢了Claude,独立复测翻车了","ornith-1-5-benchmark-reality-check","2026-08-20T17:10:00+00:00"]