[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-graphforge-ustc-qwen36-27b-evidence-graph":3,"topics-all":44,"news-related-206ea36a-1eea-463f-a241-1e3b32f5ec2d":63},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":30,"news_slug":37,"published_at":38,"created_at":39,"modified_at":40,"is_published":41,"publish_type":42,"image_url":14,"view_count":43},"206ea36a-1eea-463f-a241-1e3b32f5ec2d","USTC GraphForge:证据图把任务和 rubric 钉在一起,Qwen3.6-27B 涨三基准","USTC 团队的 GraphForge 用证据图把任务需求和验证 rubric 锚在同一份工作区文件上,Qwen3.6-27B 微调后三基准分别涨 65.7\u002F7.7\u002F13.7 分。","AI 智能体要在真实工作流里交付成果,需要的是能扎根在真实文件、又有可验证结果的训练任务。USTC 团队 9 月 30 日放出的 GraphForge 框架,把这两件事缝进了同一个证据图(evidence graph):既负责拼出任务的工作区,也负责为每条评分项(rubric)锚定能验证它需要的文件。\n\n## 证据图同时管任务与评分\n\n具体做法分四步。第一步是 occupation-grounded 种子,从一份职业画像出发(行政、会计、销售、运营等)控制任务的多样性;第二步为每个种子把真实文件组成工作区,文件之间的关系被建模为一张证据图,任务的文字要求和每条 rubric 都从图中生成,这样任务需求天然由工作区文件支撑,评分项也天然锚定到验证所需的文件;第三步先跑一次 rollout,目的是测试任务是否真能跑通,跑不通的就让 revision agent 对照原文件改任务和 rubric;第四步才去采轨迹做微调。\n\n## 2169 条轨迹换来三基准齐涨\n\n数据规模不大但效果很硬。GraphForge 论文给出的数字是:把 Qwen3.6-27B 在 2,169 条 GraphForge 轨迹上微调后,OpenHands 框架下 GDPVal 跑到了 1445.7,比基线高 65.7;Claude Code 框架下 Workspace-Bench-Lite 涨到 63.7 (+7.7),SpreadsheetBench II 涨到 24.0 (+13.7)。后续再做一轮 rejection fine-tuning,候选轨迹按证据图锚定的 rubric 筛选,三个基准还能继续涨。\n\n这次结果指向的不是某个新模型本身,而是一类新数据\u002F训练范式:训练工作智能体(working agents)的关键卡点是「任务真实 + 验证可信」,过去要么合成在 NLM 生成的假文件(缺真实性)、要么拼在真文件上但没有针对任务的验证器(结果质量无法保证),GraphForge 把这两面用同一张证据图钉死。这种「rubric 与文件同源」的思路,也呼应了最近一连串 agent 训练论文(Terminal-Universe、SkillGym、CompoWorld 等)在做的事:不再追求环境规模本身,而是追求环境与验证之间的可追溯关系。原始论文见 [arXiv:2609.38923](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38923)。\n\n## 数据与权重都已开源\n\n可公开的部分很全。USTC 团队在 Hugging Face 上放了 GraphForge 集合,里面包括 2,169 条 SFT 轨迹数据集 GraphForge-SFT-2169,以及两个 SFT 模型权重:GraphForge-Qwen3.6-27B-SFT 和更大一档的 GraphForge-Qwen3.6-35B-A3B-SFT(35B 总参\u002F3B 激活的 MoE 变体)。作者署名里既有 USTC 体系下的研究组,也有原作者 Qisheng Su(groundhogLLM)署名的 Hugging Face 提交,加上 Tao Gui、X.F. Zhao 等知名 LLM 训练研究者。论文挂在 arXiv:2609.38923,HF papers #3 论文榜、144 upvotes,目前仍处于「学术 + 模型开源」双发布的早期阶段。\n\n## 为什么这条工作值得产业读者关注\n\n放到更大的语境里看,GraphForge 的实际意义是:在 GPT\u002FClaude\u002FGemini 主导的「能聊能写」能力红利期过去之后,2026 下半年的 LLM 增量战场已经清楚地切到了「真能跑在能干工作里」的 working agents 这边 —— 而要让 working agent 真的工作起来,核心瓶颈不是更大的模型,是任务-验证闭环的真实度。GraphForge 给出的解法是证据图,与 Terminal-Universe 的「反向重建轨迹」、SkillGym 的「自动可验证环境」是同一条赛道上的不同切入点。对关注 agent infra 与训练数据的人,GraphForge 这一波是把「训练数据合成」从纯研究推到了可下载、可在 Qwen3.6-27B 上复现的实用阶段。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38923","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21,24,27],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":28,"name":29,"slug":29,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[31],{"id":32,"lang":33,"title":34,"summary":35,"content":36},"e318b139-8cef-4c99-acd2-df94a8c723d6","en","GraphForge: USTC ties tasks and rubrics together, lifts Qwen3.6-27B","USTC's GraphForge anchors task statements and verifiers in one evidence graph over real files. Qwen3.6-27B gains 65.7\u002F7.7\u002F13.7 across three benchmarks.","Working agents have to read real files, coordinate tools, and ship real deliverables — but the training data behind them is mostly synthetic, and the verifiers are usually decoupled from the task itself. The 30 September paper from USTC, GraphForge, attacks that gap with one move: anchor both the task statement and each rubric criterion to the same evidence graph over real workspace files.\n\n## Evidence graph as task + verifier anchor\n\nThe pipeline runs in four steps. A seed is sampled from an occupation-grounded distribution to control task diversity; for each seed, GraphForge assembles a workspace of real files and builds an evidence graph over their relations; task requirements and rubrics are derived from that graph, so every criterion is naturally traceable to the files needed to verify it; an initial rollout tests executability and a revision agent repairs the task and rubrics against the original files before trajectories are collected. The paper lives at [arXiv:2609.38923](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38923).\n\n## 2,169 trajectories, three benchmarks move\n\nThe numbers are small-scale but hard. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories under OpenHands takes GDPVal to 1445.7 (+65.7). Under Claude Code, Workspace-Bench-Lite reaches 63.7 (+7.7) and SpreadsheetBench II reaches 24.0 (+13.7). A rejection fine-tuning pass on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks — suggesting the rubrics themselves are the useful artefact. This is consistent with the wave of working-agent papers (Terminal-Universe, SkillGym, CompoWorld) that all pivot the same way: away from environment scale, toward environment-verifier traceability.\n\n## Data and weights are open\n\nThe artefacts are fully open. The team has released a Hugging Face collection at `groundhogLLM\u002Fgraphforge` containing the 2,169-trajectory SFT dataset (`GraphForge-SFT-2169`), a 27B SFT model (`GraphForge-Qwen3.6-27B-SFT`), and a 35B-A3B MoE variant SFT (`GraphForge-Qwen3.6-35B-A3B-SFT`). The HF papers page currently lists the work at rank #3 with 144 upvotes. Authorship spans the USTC system (Tao Gui, X.F. Zhao are recognisable names in LLM training) and groundhogLLM, the Hugging Face submitter (paper author Qisheng Su).\n\n## Why this matters for industry readers\n\nThe broader takeaway: as the chat-and-write capability race across GPT, Claude, and Gemini matures, the next real LLM battleground is whether models can actually do real work. The bottleneck is no longer model size — it is the realism of the task-verifier loop. GraphForge's evidence-graph approach is one specific answer; Terminal-Universe's reverse-engineered trajectories and SkillGym's auto-verifiable environments are different answers to the same problem. For anyone tracking agent infrastructure and training data, GraphForge is a clean example of that \"rubric-traceable task synthesis\" idea moving from a paper concept to a downloadable 27B checkpoint you can fine-tune on.","graphforge-ustc-qwen36-27b-evidence-graph","2026-10-04T03:05:00Z","2026-10-04T03:07:11.497022Z","2026-10-04T03:07:11.497033Z",true,"agent",2441,[45,54],{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":41,"created_at":52,"modified_at":53},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":55,"tag_slug":55,"title_zh":56,"title_en":57,"intro_zh":58,"intro_en":59,"id":60,"is_active":41,"created_at":61,"modified_at":62},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":64},[65,70,75,80,85,90],{"id":66,"title":67,"news_slug":68,"published_at":69},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35+00:00",{"id":71,"title":72,"news_slug":73,"published_at":74},"e40de2f9-9ee4-47f3-b1f1-7bf93a4870a9","一个动词翻转工具调用决策,LLM 内部向量现形","llm-tool-call-decision-vector","2026-10-09T13:10:00+00:00",{"id":76,"title":77,"news_slug":78,"published_at":79},"c368ad9f-9308-4a0b-8f5c-3ae4601b48b9","D2K-Bench: 专家设计把 LLM 写 GPU 核提速 33.9%","d2k-bench-llm-gpu-kernel-design-guidance","2026-10-07T03:00:00+00:00",{"id":81,"title":82,"news_slug":83,"published_at":84},"85f7f1c2-d896-436b-a915-37faed8776eb","你改主意了,模型没改:被拒需求也会带偏大模型","intent-eval-rejected-change-confusion","2026-10-06T17:15:00+00:00",{"id":86,"title":87,"news_slug":88,"published_at":89},"4c4e444f-9614-42fe-b8f4-743204f854fa","RL 后训练收「锐化税」:base 模型配轻 harness,pass@K 反超官方版","sharpening-tax-rl-post-training","2026-10-03T15:09:45+00:00",{"id":91,"title":92,"news_slug":93,"published_at":94},"4977d1f4-c8f7-480c-aaa1-ec01d69f44d8","8B 拿 SFT+GRPO 打 685B MoE:KaliBench 把 LLM 网络安全工具调用拆到命令行级","kalibench-cybersecurity-cli-runtime-rewards","2026-10-03T05:00:00+00:00"]