[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-acl-2026-skis-kv-cache":3,"news-related-5f745fe5-ea5d-453a-8b08-7dac524d1ac2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"5f745fe5-ea5d-453a-8b08-7dac524d1ac2","ACL 2026 综述 sKis：KV 缓存优化重塑为 LLM serving 系统学","近两年关于 KV Cache 压缩、量化、驱逐、prefill\u002Fdecode 拆分的工作呈爆炸式增长,但大多在孤立 benchmark 上自报收益。ACL 2026 Findings 收录的综述「Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization」(arXiv:2607.08057)第一次把这场「算法竞赛」拉回系统工程语境。\n\n论文核心是 sKis 框架(system-aware KV infrastructure for serving LLMs),把现有方法归入三个维度:时间轴覆盖调度、流水线与硬件感知执行;空间轴把布局与迁移分为 GPU 内存层级和跨计算设备两层;结构轴则涵盖量化、低秩近似、结构压缩、驱逐策略与生命周期管理(KVCC \u002F KVRM)。\n\n更有价值的是随附的 behavior × objective 矩阵——表格显式标注每个方法主要改善的是平均延迟、长尾延迟、吞吐、显存还是互联 I\u002FO,并把「质量损失」作为独立维度列入。研究指出,≥70% 的现有论文只报告其中两项收益,对互联争用、能耗、quality impact 几乎不提及。\n\n对工程团队,sKis 提供了「按瓶颈选技术」的判别工具:部署卡在显存时翻 §5.1,卡在跨卡带宽时回到 §4.2。下一篇 KV 论文若不在这张矩阵里写明自己落点,视野就明显窄了。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.08057","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"93dddc42-4a26-45e2-8e20-c904653e05d2","en","ACL 2026 survey sKis: KV cache optimization as a systems discipline","Work on KV cache compression, quantization, eviction, and prefill\u002Fdecode splitting has exploded in the past two years, but most of it self-reports gains on isolated benchmarks. The survey \"Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization\" (arXiv:2607.08057), accepted into ACL 2026 Findings, is the first to pull this \"algorithm race\" back into a systems-engineering context. The paper's core is the sKis framework (system-aware KV infrastructure for serving LLMs), classifying existing methods along three dimensions: the time axis covers scheduling, pipelining, and hardware-aware execution; the space axis splits layout and migration into GPU memory hierarchy and cross-compute-device layers; the structure axis covers quantization, low-rank approximation, structural compression, eviction strategies, and lifecycle management (KVCC \u002F KVRM). The more valuable piece is the accompanying behavior × objective matrix — the table explicitly labels whether each method primarily improves average latency, tail latency, throughput, memory, or interconnect I\u002FO, and lists \"quality loss\" as an independent dimension. The study points out that ≥70% of existing papers only report two of these benefits, with almost no mention of interconnect contention, energy, or quality impact. For engineering teams, sKis provides a \"select technique by bottleneck\" decision tool: when deployment is stuck on memory, flip to §5.1; when stuck on cross-card bandwidth, go back to §4.2. If the next KV paper doesn't locate itself in this matrix, the perspective is clearly narrow.","acl-2026-skis-kv-cache","2026-07-12T18:15:00Z","2026-07-12T18:12:39.838976Z","2026-08-19T02:08:40.142862Z",true,"agent",449,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"d2c56430-f9a7-4844-bfab-a8651279a70c","ResKV 不再把 KV 缓存压缩等同于删词：给被淘汰的信息留一份残差账本","reskv-residual-kv-cache-compression","2026-08-03T10:43:23+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"dd9d199c-5cd9-4a15-8ae6-7e4fb40f4129","MosaicKV:把 KV 缓存压成「马赛克」,长上下文推理跑出 16× 注意力加速","mosaickv-mosaic-compression","2026-07-03T18:01:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"c1ee21bd-4b59-418a-9b77-e46c790a8978","InfoKV 把 KV 缓存压缩推过「只看注意力」的临界点：用信息熵帮推理模型跑得更长","infokv-entropy-kv-cache-compression","2026-06-27T18:14:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c4380194-5bf9-43a0-8460-46436a4f2f97","LCLM 把上下文压到 1\u002F16：8.8 倍提速的代价是 16 倍时准确率只剩 75%","lclm-1-16-compress-8-8x-75pct-accuracy","2026-06-15T04:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"94e7739c-f218-4dfc-803d-3662a97321f3","DeepSeek V4 混合注意力架构解析：如何在1M上下文下将计算量降至原来的27%？","deepseek-v4-csa-hca-1m-27pct-flops","2026-05-27T07:20:00+00:00"]