[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-d2k-bench-llm-gpu-kernel-design-guidance":3,"topics-all":41,"news-related-c368ad9f-9308-4a0b-8f5c-3ae4601b48b9":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"c368ad9f-9308-4a0b-8f5c-3ae4601b48b9","D2K-Bench: 专家设计把 LLM 写 GPU 核提速 33.9%","D2K-Bench 把专家设计拆成 L1 算法\u002FL2 数据流\u002FL3 优化三级指引,让 GPT-6-Astra 等五个前沿 LLM 写出的 Triton 内核在 130 对模型-任务上正确率从 93.1% 升到 98.5%,综合性能分从 1.46 升到 1.95,几何加速比从 1.69× 升到 2.49×。","LLM 智能体写 GPU 内核已经不是新鲜事——KernelBench、MultiKernelBench、FastKernels 等基准已经能比正确率、比加速比。但 D2K-Bench 这套新基准指向了一个更精细的问题:智能体跟专家核的差距,究竟差在\"没找到好的设计\",还是\"找到了但写不出来\"?阿里、HKUST、中科大联合团队的解法是,把专家设计拆成 L1 \u002F L2 \u002F L3 三层指引喂给模型,然后看性能差消失多少。\n\n## 26 个任务、130 对模型-任务配对\n\nD2K-Bench 选了 26 个 GPU 内核任务、85 个 workload,涵盖 attention、mixture-of-experts、量化等 LLM 训练推理常见算子,实现语言统一用 Triton。任务是从 vLLM、SGLang、FlashAttention 这类系统的真实代码里抠出来的,每个任务配一份专家级 Triton\u002FCUDA\u002FCUTLASS 参考实现。基准跑在 NVIDIA B200 上,每个任务给智能体 350 个 turn 的预算,允许反复编译、profile、迭代。\n\n五款参评模型:GPT-6-Astra、Claude-Opus-4.8、GPT-5.6-Sol、GLM-5.3、Kimi-K3。每个模型跑两轮:一轮只给任务描述,一轮额外加三层专家设计指引,其它条件完全一致——同任务、同 workload、同工具、同硬件、同预算。这样的配对设计,就是为了把\"设计发现\"和\"代码实现\"两个变量干净地切开。\n\n## 三层指引从算法一直拆到 warp 调度\n\n指引按依赖关系分三层。L1 是高层算法洞察,告诉模型用什么算法变换、为什么能算得更省;L2 是数据流设计,说清楚每个变量放哪、生命周期多久、哪些中间量能复用;L3 是低层优化技巧,讲 GPU 执行机制怎么落地——流水线怎么叠、warp 怎么分工、哪些同步必须先做。指引**只讲设计思路**,不附源代码,也不给具体调参常数。\n\n论文以 Muon 正交化核作为示例:L1 提示可以把多次\"大方阵更新\"压缩成\"小方阵组合再做大乘法\",L2 提示让方阵状态在段内常驻、矩形矩阵只在边界落地,L3 提示要求对称乘积只算一次三角块、对角块只写一次。配对实验下,GPT-6-Astra 的 Performance Score 从无指引时的 2.68 升到 2.33 的专家基线还差一截的 3.33(满 26 题都正确)。\n\n## 几何加速比 1.69× → 2.49×\n\n结果分四块。\n\n**正确率**:130 对模型-任务里,引入指引后整体正确率从 93.1% 升到 98.5%,5 个模型里 GLM-5.3、Kimi-K3 在 26 题里多对了 4 道;GPT-6-Astra、Claude-Opus-4.8、GPT-5.6-Sol 在无指引下就已经 26 题全对,指引主要推的是性能不是正确性。\n\n**性能分**:五模型 Performance Score(几何平均,基线 PyTorch = 0.056)从无指引 1.46 升到有指引 1.95,相对涨 33.9%。三款 26 题全对的模型,几何平均加速比从 1.69× 升到 2.49×,单模型最大提升是 GPT-6-Astra 从 2.68× 升到 3.33×,逼近专家核自身的 2.33×。看起来指引把模型从\"接近基线\"推到了\"接近专家\"区间。\n\n**设计与实现的拆解**:论文还做了一个有意思的实验——让模型在另一轮里只写设计、不写代码,跑 LLM-as-a-judge 看它能不能自己写出合格的 L1\u002FL2\u002FL3 思路。GPT-6-Astra、Claude-Opus-4.8、GPT-5.6-Sol 三款拿到 70 分上下,GLM-5.3、Kimi-K3 落在 60 分区间。配合上\"代码实现的判官分\",两者的排序基本一致,但分差比性能分更大——说明专家指引的真正价值,不在于让模型\"想到\"更好的设计,而在于让它\"写对\"原本想不到的复杂实现。\n\n**层级叠加效应**:对三款全对模型拆 L1 \u002F L1+L2 \u002F L1+L2+L3 三档指引累加,每一档都涨性能。GPT-6-Astra 的大头收益来自 L1,说明它本身算法设计能力不弱,缺的是顶层算法选择;Claude-Opus-4.8 和 GPT-5.6-Sol 则是 L3 收尾更多,说明它们更缺 warp \u002F 流水线层面的执行机制知识。\n\n## 还差多少\n\n指引把分推到了 1.95,但专家基线是 2.33。三款前沿模型即便在指引下,代码实现分也只到 70 出头(满分 100)。也就是说,即便你把设计思路明明白白告诉它,L1\u002FL2\u002FL3 三个层面的若干属性仍然没被写进代码——边界同步、状态常驻、专用化 warp 这些细节,光看指引还摸不到手。\n\n另一层限制是指引本身来自专家代码的\"反向解读\",而实际工程里能掏到的指引颗粒度未必有这么细。但 D2K-Bench 给出的结论是清晰的:GPU 内核 agent 的瓶颈不在\"想法\",而在\"动手\"。把设计文档变成可执行 kernel 这段,大模型还远不如一个会读文档的人类工程师。\n\n代码与数据已开源,GitHub: QwenLM\u002FD2K-Bench。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.03226","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":22,"name":23,"slug":23,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":25,"name":26,"slug":26,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"17ffde8c-33ef-47ed-be64-3e549f646ba3","en","D2K-Bench: Expert Design Lifts LLM GPU Kernel Speed 33.9%","D2K-Bench: expert design guidance lifts LLM Triton-kernel speedup from 1.69x to 2.49x; correctness rises 93.1% to 98.5% on 26 tasks.","Writing GPU kernels via LLM agents isn't new — KernelBench, MultiKernelBench, and FastKernels already grade correctness and speedup. But D2K-Bench, a new benchmark from Alibaba, HKUST, and USTC, asks a sharper question: when agent-generated kernels lag expert code, is the gap a \"design discovery\" failure or an \"implementation\" failure? Their approach is to split expert design into three layers (L1 algorithmic, L2 dataflow, L3 execution tricks) and measure how much of the gap disappears when each layer is given.\n\n## 26 tasks, 130 model-task pairs\n\nD2K-Bench selects 26 GPU kernel tasks and 85 workloads covering attention, mixture-of-experts, and quantization — operators common in LLM training and inference. The implementation language is unified to Triton. Tasks are extracted from real systems code in vLLM, SGLang, and FlashAttention, and each task ships with an expert Triton \u002F CUDA \u002F CUTLASS reference. The benchmark runs on NVIDIA B200 with a 350-turn budget per task, giving agents room to compile, profile, and iterate.\n\nFive frontier LLMs are evaluated: GPT-6-Astra, Claude-Opus-4.8, GPT-5.6-Sol, GLM-5.3, and Kimi-K3. Each model runs twice — once with task descriptions only, once with the three layers of expert design guidance added. Everything else is held identical: same task, same workload, same tools, same hardware, same turn budget. The paired design cleanly separates \"did the model find a good design\" from \"did it actually implement what was given.\"\n\n## Three-layer guidance, from algorithm down to warp scheduling\n\nGuidance is structured by dependency. L1 covers high-level algorithmic insights — what transformation to use, why it preserves correctness, what compute or memory traffic it could save. L2 covers dataflow design — where each variable lives, how long it persists, which intermediates can be reused or skipped. L3 covers low-level optimization tricks — how to implement the design on real GPU execution units, how to pipeline data movement against compute, how to specialize warps, what synchronization ordering is required. Guidance offers **design rationale only**; no source code or tuning constants leak through.\n\nThe Muon orthogonalization kernel illustrates the structure. L1 hints that multiple updates to a large rectangular matrix can be compressed into composing small square matrices followed by a single large multiplication. L2 keeps the square state resident within each segment and materializes the rectangular matrix only at boundaries. L3 demands that symmetric products compute only one triangle and write diagonal tiles once. With guidance, GPT-6-Astra's Performance Score moves from 2.68 (no guidance) to 3.33 — still short of the 2.33 expert baseline, but a substantial step.\n\n## Geometric speedup: 1.69× → 2.49×\n\nThe findings break down four ways.\n\n**Correctness.** Across the 130 model-task pairs, guidance lifts overall correctness from 93.1% to 98.5%. For GLM-5.3 and Kimi-K3, that's 4 additional correct tasks out of 26. GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol already solved all 26 in the unguided run, so for them guidance mainly moves performance.\n\n**Performance Score.** The geometric-mean Performance Score (PyTorch baseline = 0.056) climbs from 1.46 to 1.95 across the five models — a 33.9% relative gain. For the three fully-correct models, geometric mean speedup goes from 1.69× to 2.49×. The biggest individual jump is GPT-6-Astra, from 2.68× to 3.33×, approaching the expert's 2.33× reference. Guidance seems to push models from \"near baseline\" toward \"near expert.\"\n\n**Design vs. implementation decomposition.** The team ran an extra experiment where each model produces only a written design (no code), then is judged by an LLM judge against L1\u002FL2\u002FL3 criteria. GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol cluster around 70\u002F100 on design; GLM-5.3 and Kimi-K3 hover near 60. Paired with the implementation judge score, the rankings roughly agree with measured performance — but the judge gaps are wider than the runtime gaps. The takeaway: expert guidance's real value isn't teaching models new designs, it's teaching them to correctly implement designs they didn't invent.\n\n**Layer-stacking effects.** Running three cumulative guidance configurations (L1 only, L1+L2, L1+L2+L3) on the fully-correct models, every step adds performance. GPT-6-Astra's biggest gain comes from L1, suggesting it already has decent algorithm intuition and mostly needs top-level algorithmic hints. Claude-Opus-4.8 and GPT-5.6-Sol gain more from L3, meaning they need warp-level execution knowledge more than dataflow advice.\n\n## How much is still missing\n\nGuidance pushes the score to 1.95, but the expert baseline sits at 2.33. Even with all three guidance layers, the three frontier models' implementation judge scores land in the low 70s out of 100. In other words, even when the design rationale is spelled out, several L1\u002FL2\u002FL3 properties still don't make it into the kernel — boundary synchronization, state residency, and warp specialization are details the model can't fully absorb from prose.\n\nAnother caveat: the guidance itself is reverse-engineered from expert code, and real engineering rarely offers design docs at that granularity. But D2K-Bench's verdict is clear. The bottleneck for GPU kernel agents isn't ideation; it's implementation. Converting a design doc into an executable kernel, the model is still well short of a competent human engineer reading the same doc.\n\nCode and data are open source on GitHub: QwenLM\u002FD2K-Bench.","d2k-bench-llm-gpu-kernel-design-guidance","2026-10-07T03:00:00Z","2026-10-07T05:14:00.720586Z","2026-10-07T05:14:00.720597Z",true,"agent",770,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"85f7f1c2-d896-436b-a915-37faed8776eb","你改主意了,模型没改:被拒需求也会带偏大模型","intent-eval-rejected-change-confusion","2026-10-06T17:15:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"206ea36a-1eea-463f-a241-1e3b32f5ec2d","USTC GraphForge:证据图把任务和 rubric 钉在一起,Qwen3.6-27B 涨三基准","graphforge-ustc-qwen36-27b-evidence-graph","2026-10-04T03:05:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"4c4e444f-9614-42fe-b8f4-743204f854fa","RL 后训练收「锐化税」:base 模型配轻 harness,pass@K 反超官方版","sharpening-tax-rl-post-training","2026-10-03T15:09:45+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"fe5546f3-09e4-42a3-bfde-6bbd0e8d474c","外星世界实测:探索 4 轮 87.6%,垫底 12.9%","explorationbench-alien-worlds","2026-09-28T21:07:30+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00"]