[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-prem-on-demand-refresh-32k":3,"news-related-a1ab01f3-ef5b-4240-aa99-7738f48591aa":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a1ab01f3-ef5b-4240-aa99-7738f48591aa","PReM 用「按需刷新」撕开 LLM 长上下文压缩天花板:阿里团队 32K 上下文做到 16×\u002F32× 压缩仍保住多跳推理","长上下文 LLM 推理的成本与命门都在 KV 缓存——压缩比一拉高,多跳推理就先崩。通义实验室郑博等人提出的 PReM(Preserve and Refresh Memory,arXiv:2607.14327)给出一个清爽回应:不追求一步到位的静态压缩,而是把长上下文当成模型内部的层式 KV 内存,在生成中按需刷新。\n\nPReM 由三个部件组成:Transformer 中间层插入专用「记忆层」对 chunks 实时打分、只留当前步骤真正需要的证据;引入特殊 token `\u003Cm>`,模型一旦输出即触发跨层 KV 内存重选;Top-k chunks 保留原始 KV、其余均值池化为单一代表向量(Preserve-and-Pool),固定预算下兼顾细节与冗余。训练端配套「相位分离刷新训练」,把推理切成内存选择与条件生成两阶段,用对比损失和边界损失逼迫模型识别证据并保证刷新前后生成连贯。\n\n32K 上下文下,PReM 在 16× 与 32× 压缩比同时压制 SnapKV、CAKE、LongLLMLingua、EXIT 等基线;多跳问答增益尤为显著,3B 小模型反超更大方案。这条信号值得记住:动态按需刷新,可能比压缩率竞赛更贴近长上下文推理的真正瓶颈。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.14327","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"89f7053f-da04-4615-8f90-5f85adc755d2","en","PReM: on-demand refresh compresses context 16-32x losslessly","The cost and Achilles' heel of long-context LLM inference both live in the KV cache — the moment compression ratio is pushed up, multi-hop reasoning collapses first. PReM (Preserve and Refresh Memory, arXiv:2607.14327) from Tongyi Lab's Zheng Bo and others gives a clean answer: instead of pursuing one-shot static compression, treat the long context as a layer-wise KV memory inside the model and refresh it on demand during generation. PReM has three components: a dedicated \"memory layer\" inserted in the middle of the Transformer scores chunks in real time, keeping only the evidence truly needed by the current step; a special token `\u003Cm>` is introduced — once the model outputs it, it triggers a cross-layer KV-memory re-selection; the Top-k chunks retain their original KV, the rest are mean-pooled into a single representative vector (Preserve-and-Pool), balancing detail and redundancy under a fixed budget. The training side has a paired \"phase-separated refresh training\", slicing inference into memory-selection and conditional-generation stages, and using contrast loss and boundary loss to force the model to identify evidence and guarantee generation consistency across refreshes. On 32K contexts, PReM simultaneously beats SnapKV, CAKE, LongLLMLingua, EXIT and other baselines at 16× and 32× compression ratios; the gain on multi-hop Q&A is particularly striking, with a 3B small model overtaking larger solutions. The signal worth remembering: dynamic on-demand refresh may be closer to the real bottleneck of long-context inference than the compression-ratio race.","prem-on-demand-refresh-32k","2026-07-18T20:08:00Z","2026-07-18T20:09:47.885403Z","2026-08-19T02:08:40.142862Z",true,"agent",114,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"183fb3be-e062-47e7-9591-7c2372e116c1","LLM 蒸馏的显存瓶颈不只在教师模型：离线 Top-K 与分块 KL 把长上下文训练装回单卡","llm-distillation-offline-top-k-chunked-kl","2026-08-05T20:08:13+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"d2c56430-f9a7-4844-bfab-a8651279a70c","ResKV 不再把 KV 缓存压缩等同于删词：给被淘汰的信息留一份残差账本","reskv-residual-kv-cache-compression","2026-08-03T10:43:23+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"bffbd811-83a4-455a-a441-386dde3661c5","多模态 LLM 边缘推理:压缩、MoE 路由与量化「互锁」才是真战场","multimodal-llm-edge-interlock","2026-07-26T07:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"bdb819d1-09a0-4320-8f78-04dccb15571d","16GB 显卡微调 131K 上下文：Hierarchical Global Attention","hierarchical-global-attention-16gb","2026-07-18T18:00:00+00:00"]