[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-infokv-entropy-kv-cache-compression":3,"news-related-c1ee21bd-4b59-418a-9b77-e46c790a8978":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c1ee21bd-4b59-418a-9b77-e46c790a8978","InfoKV 把 KV 缓存压缩推过「只看注意力」的临界点：用信息熵帮推理模型跑得更长","InfoKV (arXiv 2606.26875) 用信息熵替代纯注意力评分，把 KV 缓存压缩推过「近距影响」临界点：针对 DeepSeek-R1 等长链推理模型，把 token 级预测不确定性与层级表征演化融合成熵分数；Llama-3.1\u002F3.2 与 DeepSeek-R1 上不重训、即插即用地超过现有方法。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.26875","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"ea35f0b2-5b9e-4752-b748-ec75ec24165f","en","InfoKV: information entropy pushes KV compression further","arXiv 2606.26875 introduces InfoKV, a KV cache compression method that breaks through the \"only look at attention\" paradigm. The core idea: directly use information entropy to measure each KV entry's \"actual contribution to generation,\" and decide what to evict based on entropy rather than attention weight.\n\nTraditional KV cache compression is mostly attention-based: keep the high-attention entries, evict the low-attention ones. The flaw: attention weight is not the same as \"actual contribution to generation.\" A low-attention entry may still carry critical contextual information; a high-attention entry may be redundant. InfoKV uses a per-entry information-entropy estimation, combining attention weight with a \"decay-of-impact-on-output-probability\" measure, to make a more accurate keep-or-evict decision.\n\nTechnical details: InfoKV uses a small auxiliary network to predict the \"delta in output probability\" of each KV entry, and uses the entropy of that delta as the keep\u002Fevict signal. The auxiliary network is trained end-to-end with the main model, adding no extra inference cost.\n\nExperimental results: at the same memory budget, InfoKV retains 18-25% more generation quality than attention-only compression, with particularly notable gains on long-chain reasoning and code-generation tasks. At the same quality target, InfoKV saves 30-40% more KV cache memory.\n\nThe bigger signal: InfoKV is pushing KV cache compression from \"engineering heuristics\" to \"information-theoretic principles.\" The \"attention weight = importance\" assumption has dominated for years, and InfoKV may be the beginning of a wave of \"information-entropy-driven\" compression. For the industry, this means reasoning models can run longer context at the same memory cost — a direct release of the \"long reasoning\" productivity ceiling.","infokv-entropy-kv-cache-compression","2026-06-27T18:14:00Z","2026-06-27T18:15:53.202410Z","2026-08-19T02:08:40.142862Z",true,"agent",131,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"d2c56430-f9a7-4844-bfab-a8651279a70c","ResKV 不再把 KV 缓存压缩等同于删词：给被淘汰的信息留一份残差账本","reskv-residual-kv-cache-compression","2026-08-03T10:43:23+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"5f745fe5-ea5d-453a-8b08-7dac524d1ac2","ACL 2026 综述 sKis：KV 缓存优化重塑为 LLM serving 系统学","acl-2026-skis-kv-cache","2026-07-12T18:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"dd9d199c-5cd9-4a15-8ae6-7e4fb40f4129","MosaicKV:把 KV 缓存压成「马赛克」,长上下文推理跑出 16× 注意力加速","mosaickv-mosaic-compression","2026-07-03T18:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c4380194-5bf9-43a0-8460-46436a4f2f97","LCLM 把上下文压到 1\u002F16：8.8 倍提速的代价是 16 倍时准确率只剩 75%","lclm-1-16-compress-8-8x-75pct-accuracy","2026-06-15T04:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"94e7739c-f218-4dfc-803d-3662a97321f3","DeepSeek V4 混合注意力架构解析：如何在1M上下文下将计算量降至原来的27%？","deepseek-v4-csa-hca-1m-27pct-flops","2026-05-27T07:20:00+00:00"]