Writing GPU kernels via LLM agents isn't new — KernelBench, MultiKernelBench, and FastKernels already grade correctness and speedup. But D2K-Bench, a new benchmark from Alibaba, HKUST, and USTC, asks a sharper question: when agent-generated kernels lag expert code, is the gap a "design discovery" failure or an "implementation" failure? Their approach is to split expert design into three layers (L1 algorithmic, L2 dataflow, L3 execution tricks) and measure how much of the gap disappears when each layer is given.

26 tasks, 130 model-task pairs

D2K-Bench selects 26 GPU kernel tasks and 85 workloads covering attention, mixture-of-experts, and quantization — operators common in LLM training and inference. The implementation language is unified to Triton. Tasks are extracted from real systems code in vLLM, SGLang, and FlashAttention, and each task ships with an expert Triton / CUDA / CUTLASS reference. The benchmark runs on NVIDIA B200 with a 350-turn budget per task, giving agents room to compile, profile, and iterate.

Five frontier LLMs are evaluated: GPT-6-Astra, Claude-Opus-4.8, GPT-5.6-Sol, GLM-5.3, and Kimi-K3. Each model runs twice — once with task descriptions only, once with the three layers of expert design guidance added. Everything else is held identical: same task, same workload, same tools, same hardware, same turn budget. The paired design cleanly separates "did the model find a good design" from "did it actually implement what was given."

Three-layer guidance, from algorithm down to warp scheduling

Guidance is structured by dependency. L1 covers high-level algorithmic insights — what transformation to use, why it preserves correctness, what compute or memory traffic it could save. L2 covers dataflow design — where each variable lives, how long it persists, which intermediates can be reused or skipped. L3 covers low-level optimization tricks — how to implement the design on real GPU execution units, how to pipeline data movement against compute, how to specialize warps, what synchronization ordering is required. Guidance offers design rationale only; no source code or tuning constants leak through.

The Muon orthogonalization kernel illustrates the structure. L1 hints that multiple updates to a large rectangular matrix can be compressed into composing small square matrices followed by a single large multiplication. L2 keeps the square state resident within each segment and materializes the rectangular matrix only at boundaries. L3 demands that symmetric products compute only one triangle and write diagonal tiles once. With guidance, GPT-6-Astra's Performance Score moves from 2.68 (no guidance) to 3.33 — still short of the 2.33 expert baseline, but a substantial step.

Geometric speedup: 1.69× → 2.49×

The findings break down four ways.

Correctness. Across the 130 model-task pairs, guidance lifts overall correctness from 93.1% to 98.5%. For GLM-5.3 and Kimi-K3, that's 4 additional correct tasks out of 26. GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol already solved all 26 in the unguided run, so for them guidance mainly moves performance.

Performance Score. The geometric-mean Performance Score (PyTorch baseline = 0.056) climbs from 1.46 to 1.95 across the five models — a 33.9% relative gain. For the three fully-correct models, geometric mean speedup goes from 1.69× to 2.49×. The biggest individual jump is GPT-6-Astra, from 2.68× to 3.33×, approaching the expert's 2.33× reference. Guidance seems to push models from "near baseline" toward "near expert."

Design vs. implementation decomposition. The team ran an extra experiment where each model produces only a written design (no code), then is judged by an LLM judge against L1/L2/L3 criteria. GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol cluster around 70/100 on design; GLM-5.3 and Kimi-K3 hover near 60. Paired with the implementation judge score, the rankings roughly agree with measured performance — but the judge gaps are wider than the runtime gaps. The takeaway: expert guidance's real value isn't teaching models new designs, it's teaching them to correctly implement designs they didn't invent.

Layer-stacking effects. Running three cumulative guidance configurations (L1 only, L1+L2, L1+L2+L3) on the fully-correct models, every step adds performance. GPT-6-Astra's biggest gain comes from L1, suggesting it already has decent algorithm intuition and mostly needs top-level algorithmic hints. Claude-Opus-4.8 and GPT-5.6-Sol gain more from L3, meaning they need warp-level execution knowledge more than dataflow advice.

How much is still missing

Guidance pushes the score to 1.95, but the expert baseline sits at 2.33. Even with all three guidance layers, the three frontier models' implementation judge scores land in the low 70s out of 100. In other words, even when the design rationale is spelled out, several L1/L2/L3 properties still don't make it into the kernel — boundary synchronization, state residency, and warp specialization are details the model can't fully absorb from prose.

Another caveat: the guidance itself is reverse-engineered from expert code, and real engineering rarely offers design docs at that granularity. But D2K-Bench's verdict is clear. The bottleneck for GPU kernel agents isn't ideation; it's implementation. Converting a design doc into an executable kernel, the model is still well short of a competent human engineer reading the same doc.

Code and data are open source on GitHub: QwenLM/D2K-Bench.