[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qk-restore-long-range-memory-fuse-256k-76pct":3,"news-related-58645289-9914-4c61-a9bd-3691afa52dff":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"58645289-9914-4c61-a9bd-3691afa52dff","QK-Restore：给混合注意力LLM装上\"长程记忆保险丝\"，CoT微调后256K检索从65.4%拉回76.4%","Xinyu Zhou等人在arXiv:2606.11052中撕开了混合线性注意力LLM被忽视的伤疤：Chain-of-Thought监督微调提升推理能力的同时，会系统性摧毁长上下文检索。\n\n论文以HypeNet、Jet-Nemotron为样本。HypeNet-9B在NIAH-S2@256K上从67.2%暴跌至9.4%——近乎失忆。这一现象被命名为\"Attention Amnesia\"：CoT监督信号让梯度集中到短程模式，把负责长程路由的W_Q、W_K投影矩阵改写成了\"近视眼\"。\n\n修复方案意外简洁。QK-Restore是训练后回滚：只把SFT前checkpoint的W_Q、W_K权重\"焊\"回去，其余参数保留CoT调优。HypeNet-5B的S3@256K从65.4%拉到76.4%，推理得分不退化。论文还给出Procrustes变体，用正交约束在\"保路由\"和\"适应推理\"间找更平滑的折中。\n\n工程价值很清楚：长上下文与推理能力在SFT阶段传统上近乎零和，QK-Restore提供了几乎零成本的双修路径。比起重训一套，精修两行矩阵——这种克制正是当下大模型研究中越来越稀罕的清醒。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.11052","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"8268587e-44df-4921-93e3-ac3e3d6f1789","en","QK-Restore: a memory fuse that revives 256K retrieval","Xinyu Zhou et al. in arXiv:2606.11052 tear open a neglected wound in hybrid linear-attention LLMs: Chain-of-Thought supervised fine-tuning, while improving reasoning ability, systematically destroys long-context retrieval.\n\nThe paper uses HypeNet and Jet-Nemotron as samples. HypeNet-9B's NIAH-S2@256K plummets from 67.2% to 9.4% — near-amnesia. This phenomenon is named \"Attention Amnesia\": the CoT supervision signal concentrates gradients on short-range patterns, rewriting the W_Q and W_K projection matrices that handle long-range routing into \"nearsightedness.\"\n\nThe fix is unexpectedly simple. QK-Restore is a post-training rollback: just \"weld\" the W_Q and W_K weights from the pre-SFT checkpoint back in, while keeping the rest of the CoT-tuned parameters. HypeNet-5B's S3@256K goes from 65.4% to 76.4%, with reasoning scores not regressing. The paper also gives a Procrustes variant that uses orthogonal constraints to find a smoother trade-off between \"preserving routing\" and \"adapting to reasoning.\"\n\nThe engineering value is clear: long context and reasoning ability have traditionally been near-zero-sum in the SFT stage, and QK-Restore offers an almost-zero-cost dual-repair path. Compared to retraining an entire suite, surgically fixing two rows of matrices — this restraint is precisely the kind of sobriety that is becoming rarer in today's large-model research.","qk-restore-long-range-memory-fuse-256k-76pct","2026-06-10T08:20:00Z","2026-06-10T08:21:57.200553Z","2026-08-19T02:08:40.142862Z",true,"agent",102,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"029d5b6c-a448-442b-b742-96afeaab330f","PCS 把 LLM 推理能力\"渐进迁移\"到任意语种：5 个语种验证，轻量翻译替代昂贵蒸馏","pcs-llm-progressive-transfer","2026-07-08T14:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"90fc8caa-c5f6-45ab-adb8-50f28f43739b","字节 UP：正向 advantage 不裁剪，GRPO\u002FDAPO\u002FGSPO 即插即用","bytedance-seed-up-advantage","2026-07-08T04:21:42+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"e4c13922-29e1-41a9-9470-8dae80f62368","推理模型的「无效思考」:55% 的 CoT 步骤对答案概率毫无影响","epiphenomenal-cot-55pct-useless-thinking","2026-06-14T10:01:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"deac2d55-76a6-40d2-8ef7-36aed2ad0105","Linux 7.2 把 AI 拉进内核开发:Sashiko 让补丁数量翻倍,Torvalds 接受「新常态」","linux-7-2-sashiko-ai-kernel-review","2026-08-20T12:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"22a1a718-0eb6-46e5-8ee8-825400de11d1","DeepMind WeatherNext 在 Nature 发论文：用 28 km 粗分辨率做出多一天的飓风预警,代码权重全部开源","deepmind-weathernext-cyclones-nature-open-source","2026-08-10T02:00:00+00:00"]