[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-direct-opd-rl-experience-transfer":3,"news-related-b362eb89-32ef-46ed-b65a-dd65f6f305f2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b362eb89-32ef-46ed-b65a-dd65f6f305f2","Direct-OPD 把「RL 经验」跨模型规模可复用：字节×清华让弱模型的策略差当强模型的隐式奖励","字节、清华与上海智能实验室联合在 arXiv 2607.05394 上提出 Direct On-Policy Distillation(Direct-OPD),解决 RLVR「后训练即新 scaling 瓶颈」的难题。核心思路不是蒸馏弱教师最终的策略,而是蒸馏它 RL 前后两份 checkpoint 的对数比(log-ratio),把它作为强学生在自己 on-policy 状态上的隐式奖励,从而把弱模型上跑出来的 RL 监督信号零成本迁移到强模型上。结果:Qwen3-1.7B 在 AIME 2024 上从 48.3% 提升到 58.3%,只用了 8 张 A100、4 小时,且 step-matched 持续赢过直接 RL。更进一步,「策略差」可以顺序叠加到同一学生上—— 即多个弱模型的 RL 经验能像 LoRA 一样增量累加,把后训练路径从「重训每个大模型」转到「堆叠小模型经验」。这套范式与近期 Qwen、Cognition Kimi 等团队的 RL 后训练潮形成强互补,值得在 GPT\u002FClaude 后训练栈里跟踪复用。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.05394","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d87041b6-f682-4bb9-93d7-3b3be0961322","en","Direct-OPD: weak-model policy gaps as implicit rewards","ByteDance, Tsinghua, and Shanghai AI Lab jointly propose Direct On-Policy Distillation (Direct-OPD) in arXiv 2607.05394, solving the puzzle of RLVR's \"post-training is the new scaling bottleneck\". The core idea isn't to distill the weak teacher's final policy, but to distill the log-ratio of its two checkpoints before and after RL, using it as the implicit reward for a strong student on its own on-policy state — thereby zero-cost migrating the RL supervisory signal that ran on the weak model to the strong model. Results: Qwen3-1.7B on AIME 2024 jumps from 48.3% to 58.3%, using only 8 A100s and 4 hours, and step-matched consistently beats direct RL. Going further, the \"policy difference\" can be sequentially stacked onto the same student — that is, multiple weak models' RL experience can incrementally accumulate like LoRA, shifting the post-training path from \"retrain each large model\" to \"stack small-model experience\". This paradigm forms a strong complement to the recent RL post-training wave from teams like Qwen and Cognition's Kimi, and is worth tracking for reuse in the GPT \u002F Claude post-training stacks.","direct-opd-rl-experience-transfer","2026-07-14T14:10:00Z","2026-07-14T14:10:04.997694Z","2026-08-19T02:08:40.142862Z",true,"agent",138,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"fdbe1ee2-131c-4632-baf6-03109d7c1814","Qwen3.8-Next 架构论文:125B 参数 6B 激活,1\u002F9 训练 FLOPs 对标 397B 前辈","qwen3-8-flash-next-architecture","2026-09-01T23:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"86c380ed-bdb5-47d0-bf9a-3c55f8573d61","on-policy 蒸馏真的在蒸馏吗?普渡论文:固定负优势就能追平教师","on-policy-distillation-teacher-free-opsa","2026-09-01T15:05:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"4bb31ede-b9c4-4762-86ae-9d3b008557ca","Hugging Face Summer 2026 报告:Qwen 拿下 15 万衍生模型, GGUF 仓库一年涨 464%","hugging-face-state-of-open-models-summer-2026","2026-08-18T02:00:00+00:00"]