[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kara-kv-cache-sliding-window":3,"topics-all":36,"news-related-ddbac1f6-08ab-4663-ab2b-da6793947e49":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"ddbac1f6-08ab-4663-ab2b-da6793947e49","Kara:把 KV 缓存压成「滑动窗口」,让推理 LLM 在高并发下不再卡顿","每条 token 都把 K\u002FV 缓存塞进 HBM,推理 LLM 的「长 CoT + 高并发」组合是 KV 缓存压缩研究的真正试炼场。卡内基梅隆的 Han Shen 与 Yuyang Wu 提出的 Kara(arXiv:2607.01237),从「窗口边界」和「保留粒度」两个老问题入手,给出了目前最干净的一组解。\n\nKara 的核心是只压缩最近生成的上下文窗口——这避开了 SnapKV\u002FAdaKV 类「阈值触发 + 全窗口重打分」带来的反复压缩开销;更重要的是,Kara 用双向注意力而不是单向往回看的 query 来打分 KV 对,让保留候选能跨越前后位置,不再被前缀位置主导。然后 Token2Chunk 模块把候选离散 KV 对再扩展成「任意长度的连续 chunk」,既保留离散关键 token 的指向性,又保留 chunk 的语义连续性——这恰好补上了 ChunkKV 「刚性边界」那块短板。\n\n在 PagedAttention 上落地的 KvLLM 框架,设计了周期触发策略而非阈值触发,直接避开了「压缩开销反而压低吞吐」的并发-吞吐反转问题。Qwen3-4B\u002F14B 与 DeepSeek-R1-Distill-Llama-8B 上的实验显示,Kara 在 MATH-500、AIME24、AMC23 上以 30% 保留率几乎保持无压缩精度,NIAH 上的检索表现也明显优于 ChunkKV 与 AdaKV。\n\n观点:Kara 的双向打分 + 灵活 chunk 组合,本质上是把「KV 保留」从一维排序问题升级成二维布局问题。这种升级让 7B\u002F14B 量级推理模型在 8×H100\u002FH200 上跑高并发业务时,首次具备了「压缩不掉精度、吞吐还能涨」的可能。对部署方而言,KvLLM 的周期触发策略比 SnapKV 类阈值触发更适合长 CoT 推理服务——这条路线值得跟进,但工程化落地仍要看 PagedAttention 跨节点时的同步开销是否被周期触发掩盖。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.01237","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"20768567-874e-4d9b-81b1-a428726cbdd0","en","Kara squeezes KV cache into sliding windows for concurrency","Every token stuffs its K\u002FV cache into HBM, making the \"long CoT + high concurrency\" combination for inference LLM the real testing ground for KV cache compression research. Han Shen and Yuyang Wu of Carnegie Mellon, in their Kara (arXiv:2607.01237), take on the two old problems of \"window boundary\" and \"retention granularity\" and give the cleanest set of solutions currently available. Kara's core is to compress only the most recently generated context window — this avoids the repeated compression overhead of SnapKV\u002FAdaKV-class \"threshold-trigger + full-window re-scoring\"; more importantly, Kara uses bidirectional attention rather than the single-direction backward-looking query to score KV pairs, letting retention candidates span forward and backward positions, no longer dominated by prefix position. Then the Token2Chunk module extends candidate discrete KV pairs into \"arbitrary-length continuous chunks\", both retaining the pointer property of discrete key tokens and preserving the semantic continuity of chunks — this neatly fills the \"rigid boundary\" short board of ChunkKV. In the KvLLM framework landed on PagedAttention, the design uses a periodic-trigger strategy instead of threshold-trigger, directly sidestepping the \"compression overhead actually lowers throughput\" concurrency-throughput inversion problem. Experiments on Qwen3-4B\u002F14B and DeepSeek-R1-Distill-Llama-8B show that Kara at 30% retention nearly maintains uncompressed accuracy on MATH-500, AIME24, AMC23, and the retrieval performance on NIAH also significantly outperforms ChunkKV and AdaKV. Opinion: Kara's bidirectional scoring + flexible chunk combination essentially upgrades \"KV retention\" from a one-dimensional sorting problem to a two-dimensional layout problem. This upgrade lets 7B\u002F14B-tier inference models running high-concurrency workloads on 8×H100\u002FH200 for the first time have the possibility of \"compression without losing accuracy, throughput can still rise\". For deployers, KvLLM's periodic-trigger strategy is more suitable for long-CoT inference services than SnapKV-class threshold-trigger — this path is worth following, but engineering landing still depends on whether the synchronization overhead of PagedAttention across nodes is masked by the periodic trigger.","kara-kv-cache-sliding-window","2026-07-05T06:10:00Z","2026-07-04T22:08:43.956647Z","2026-08-19T02:08:40.142862Z",true,"agent",172,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"0bb9aed0-f4f6-4d1a-974a-88e47010815e","长上下文压成答案导向记忆:CMC 让冻结 LLM 省一半显存","cmc-context-memory-embedding-long-context-compression","2026-10-07T00:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"5fa17d26-6090-4805-893a-aa880cf33369","让 Agent 学会遗忘:删掉旧推理,分数反而涨了","agents-forget-reasoning-iclr-compression","2026-09-28T19:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"bc642ebf-9b1c-41cf-98d1-dd55393fd429","HeadWiseKV:无训练KV cache压缩让混合LLM长上下文从114K推到161K","headwisekv-training-free-kv-cache-compression","2026-09-03T03:44:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"7bb4f5ec-14e0-43b6-9913-07cad82a520b","微软内部 AI 账单失控:单员工月烧 2.8 万美元,倒逼默认模型换人","microsoft-internal-ai-bill-explode-default-model-swap","2026-08-28T04:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"6fa1bc74-c98e-476b-bc4c-9ae057105ffb","ParaTempo:免训练并行推理,延迟最高降 32%、token 省三成","paratempo-temporal-confidence-parallel-reasoning","2026-08-24T17:20:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"ddb7bc6c-6b6e-4797-ab76-d1aeab5a3002","压缩得好≠部署得好:树莓派实测边缘 LLM,LoRA恢复模型100题押97个同答案","edge-llm-compression-raspberry-pi","2026-08-23T13:30:00+00:00"]