[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-single-layer-transformer-rl":3,"news-related-09adc53e-c559-4899-bc68-117b19b717a3":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"09adc53e-c559-4899-bc68-117b19b717a3","单层 Transformer 就能打平全参数 RL 后训练:Qwen 上的「中段层集中增益」现象","arXiv 2607.01232(7 月 1 日提交,2 日更新 v2)抛出一个对 LLM 后训练成本结构颇具冲击力的发现:在 GRPO、GiGPO、Dr.GRPO 三种主流 RL 算法下,只训练 Qwen3 \u002F Qwen2.5 模型中的**单层 Transformer**,就可以恢复绝大部分全参数 RL 增益,某些任务甚至超过全参数训练。\n\n作者把这种现象量化为「层贡献」(layer contribution),即单层训练能恢复的全参数 RL 收益比例。在覆盖 7 个模型、横跨数学推理、代码生成、Agent 决策三类任务的实验里,「贡献度」高的层高度集中在 Transformer 堆栈的中段,而靠输入和靠输出两端的几层,RL 训练带来的增益几乎可以忽略。\n\n更耐人寻味的是,这个「中段集中」模式在数据集、任务、模型族、RL 算法之间都保持高度稳定——也就是说,真正承载 RL 适配能力的那几层,几乎总是落在同一个相对位置上。\n\n这一发现对后训练流程的实际影响是:目前业内普遍采用的全参数 RL 后训练,可能在相当程度上「过度付费」——绝大多数参数更新其实可以省掉,只对中段几个关键层做精细化微调就足以逼近全局 RL 的效果。如果这一结论在大模型上被复现,post-training 的成本曲线有机会出现一次类似预训练 MoE 化那样的台阶式下降,LoRA \u002F IA³ 这类稀疏化方法的设计思路,也会从「节省推理显存」延伸到「节省训练算力」。\n\n当然,论文也坦承:实验主要在 Qwen 系列和 3 种算法上完成,70B+ 级别模型以及 PPO、DPO 等其他算法上的表现,还需要后续工作验证。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.01232","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"2c7b33da-ec26-4ea2-b6dc-59a3bbb6bdec","en","Single-layer Transformer matches full RL post-training on Qwen","arXiv 2607.01232 (submitted July 1, updated v2 on July 2) throws out a finding that's quite impactful for the cost structure of LLM post-training: under three mainstream RL algorithms GRPO, GiGPO, and Dr.GRPO, training only a **single layer Transformer** in Qwen3 \u002F Qwen2.5 models can recover most of the full-parameter RL gain, and on some tasks even surpass full-parameter training. The authors quantify this phenomenon as \"layer contribution\" — the proportion of full-parameter RL gain that single-layer training can recover. In experiments covering 7 models, spanning three categories of tasks: math reasoning, code generation, Agent decision-making, the \"contribution\" high layers are highly concentrated in the middle of the Transformer stack, while the layers near input and output ends get almost negligible gain from RL training. More intriguing is that this \"middle-concentrated\" pattern remains highly stable across datasets, tasks, model families, and RL algorithms — that is, those few layers that truly carry RL adaptation capacity almost always fall in the same relative position. The practical impact of this finding on the post-training pipeline is: the current industry's commonly used full-parameter RL post-training may be \"overpaying\" to a large extent — most parameter updates can actually be saved, and only fine-tuning the few key layers in the middle is enough to approximate the effect of global RL. If this conclusion is replicated on large models, the cost curve of post-training has the chance of a step-down similar to the pretraining MoE-ization, and the design philosophy of sparsification methods like LoRA \u002F IA³ will also extend from \"saving inference memory\" to \"saving training compute\". Of course, the paper also candidly admits: the experiments are mainly done on the Qwen series and 3 algorithms, and the performance on 70B+ models and other algorithms like PPO and DPO still needs follow-up work to verify.","single-layer-transformer-rl","2026-07-03T08:00:00Z","2026-07-03T08:06:49.463327Z","2026-08-19T02:08:40.142862Z",true,"agent",116,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b4f270b3-db43-4586-a0e5-a062320c6d1b","让模型自己声明看哪里:Declarative Attention 零训练砍 52% KV 读取","declarative-attention-kv-cache-declare","2026-09-03T23:07:03+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"fdbe1ee2-131c-4632-baf6-03109d7c1814","Qwen3.8-Next 架构论文:125B 参数 6B 激活,1\u002F9 训练 FLOPs 对标 397B 前辈","qwen3-8-flash-next-architecture","2026-09-01T23:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"86c380ed-bdb5-47d0-bf9a-3c55f8573d61","on-policy 蒸馏真的在蒸馏吗?普渡论文:固定负优势就能追平教师","on-policy-distillation-teacher-free-opsa","2026-09-01T15:05:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"3d922c00-afcb-4f1c-a6d5-8f9d6c10c642","从 Kimi Linear 到 Kimi K3:MoE 推理效率战里被忽略的架构升级","kimi-k3-latentmoe-kda-attnres-nope","2026-07-30T00:30:00+00:00"]