[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-naver-on-policy-delta-distillation":3,"news-related-4d92e0b1-04a3-4524-9ae3-b8456aa74f2a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"4d92e0b1-04a3-4524-9ae3-b8456aa74f2a","NAVER 提出 On-Policy Delta Distillation:用「差分信号」重新定义推理蒸馏","NAVER AI Lab 于 2026 年 7 月 16 日在 arXiv 发布《On-Policy Delta Distillation》(arXiv:2607.15161),为推理大模型的后训练蒸馏提供了一条更直接、更经济的路径,作者为 Byeongho Heo、Jaehui Hwang、Sangdoo Yun 与 Dongyoon Han。\n\n传统 on-policy distillation 的目标,通常直接让学生模仿教师的输出分布。然而这一目标不可避免地混入了教师自身的语言先验——学生既要学推理,又要承担教师对世界的所有已学先验,代价偏高,效率受限。NAVER 团队的核心洞察是:推理能力的本质,源自「教师相对其同源 base 模型的增量变化」。论文把这一增量定义为 delta signal,即「教师模型与其同源 base 模型在每个 token 上的分布差」,精准剥离预训练残留,只留下「指令微调或推理 RL 过程中新引入的成分」。\n\n基于 delta signal 重新设计的 OPD² 蒸馏目标,在数学、科学、代码推理基准上一致优于传统 on-policy distillation,学生模型仅需极短的后训练周期即可逼近教师表现。论文报告的实验覆盖 19 页正文、4 张图、12 张表,具有相对完整的实证支撑。代码将于 github.com\u002Fnaver-ai\u002Fopd2 开源,后续可接入到主流 RL 后训练流程中。\n\n这条路径更深层的意义在于,把「蒸馏什么」从「模仿整体输出分布」重新校准为「迁移教师在 RL\u002F指令微调中真正学到的新能力」。它延续了过去半年学界对 reasoning post-training 目标函数的精细化讨论——从 token 级 imitation 到 preference-based RL,再到 delta-based distillation,目标粒度越做越细。对追求小模型继承大模型推理能力、又要控制训练成本的研究与工程团队,这是近几个月里少有的、机制层面而非工程层面真正推进的进展。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.15161","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"102cd934-fd7a-490b-8d29-8370eec35e6b","en","NAVER's On-Policy Delta Distillation redefines reasoning transfer","NAVER AI Lab published \"On-Policy Delta Distillation\" (arXiv:2607.15161) on arXiv on July 16, providing a more direct, more economical path for distilling large reasoning models in post-training. Authors: Byeongho Heo, Jaehui Hwang, Sangdoo Yun and Dongyoon Han. The traditional on-policy distillation target usually asks the student to directly mimic the teacher's output distribution. But this target inevitably mixes in the teacher's own language priors — the student has to both learn reasoning and bear all of the teacher's pre-existing world priors, which is costly and inefficient. NAVER's core insight: the essence of reasoning ability comes from \"the incremental change of the teacher relative to its homologous base model\". The paper defines this increment as the delta signal, i.e., \"the per-token distribution difference between the teacher and its homologous base model\", precisely stripping away the pretraining residue and keeping only \"the components newly introduced during instruction tuning or reasoning RL\". The OPD² distillation objective, redesigned on the delta signal, consistently outperforms traditional on-policy distillation on math, science and code reasoning benchmarks; the student needs only a very short post-training cycle to approach the teacher's performance. The experimental report covers 19 pages of text, 4 figures and 12 tables, with relatively complete empirical support. Code will be open-sourced at github.com\u002Fnaver-ai\u002Fopd2, and can be plugged into mainstream RL post-training flows afterwards. The deeper meaning of this path: it re-calibrates \"what to distill\" from \"mimic the overall output distribution\" to \"transfer what the teacher truly learned new during RL\u002Finstruction tuning\". It continues the past half-year's academic discussion of fine-grained reasoning post-training objectives — from token-level imitation to preference-based RL, to delta-based distillation, with the objective granularity getting finer and finer. For research and engineering teams chasing small models inheriting large-model reasoning ability while controlling training cost, this is one of the few genuinely mechanism-level rather than engineering-level advances in recent months.","naver-on-policy-delta-distillation","2026-07-18T16:07:00Z","2026-07-18T16:13:58.387601Z","2026-08-19T02:08:40.142862Z",true,"agent",285,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a457f7b9-dde3-4d00-bbc0-cdf9ef2dde14","xHC：Transformer 残差流扩成 16 车道，突破 N=4","xhc-expanded-hyper-connections","2026-07-18T00:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f092a4b2-4e8b-45c7-9045-570047b8f92c","SIREN-RoPE：让位置编码学会「旋转」","siren-rope-learnable-rotation-positional-encoding","2026-04-29T01:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"0a3f5044-ed20-4c2a-b710-bd26cd276d3e","ALiBi 的隐藏数值故障：长上下文越长，部分注意力头越可能“失明”","alibi-attention-underflow-long-context","2026-08-06T10:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"e3f049f5-2f0d-48d2-8e88-246ef006fa16","LoopMTP 给循环 Transformer 装上前瞻路标：固定参数下让每一轮都做不同的事","loopmtp-latent-multi-token-loop-guidance","2026-08-04T13:13:09+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"cb7fb8b3-5862-4cba-adab-c4794e989966","图灵奖得主 Pearl 长访谈：LLM 能讲因果只是因为人类替它爬过了因果阶梯","judah-pearl-llm-causal-ladder-agi","2026-07-31T07:00:00+00:00"]