[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mit-attention-matching-kv-cache-50x":3,"topics-all":36,"news-related-49b668e1-dcbe-481d-9fa8-438fadf77a9b":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"49b668e1-dcbe-481d-9fa8-438fadf77a9b","注意力匹配算法：MIT让LLM长上下文推理成本骤降","当LLM处理超长上下文时，KV缓存是最大的内存瓶颈。随着对话越来越长，模型必须为每个历史token保留key和value向量，这些数据可轻松膨胀到数GB。此前业界尝试过token驱逐、合并或截断等方案，但在需要极端压缩的企业场景中表现急剧下降。另一条路是Cartridges方法——用梯度优化训练紧凑KV缓存，但每次压缩需GPU运行数小时，无法用于实时应用。MIT团队换了个思路：只要保留两个关键数学属性——注意力输出和注意力质量，压缩后的缓存就能完美模拟原始行为。Attention Matching基于此将KV缓存每个head压缩为更少key-value对，在部分数据集实现最高50倍压缩，耗时仅数秒，完全无需训练。论文已被ICLR 2026接收。这项技术意味着长上下文服务的成本结构将迎来显著改善，但50倍是部分数据集峰值数字，实际效果因模型和任务类型而异。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2602.16284","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"43575d47-d96f-400b-b2c7-4309b912af25","en","MIT's attention matching slashes long-context inference cost","When LLMs handle super-long contexts, the KV cache is the biggest memory bottleneck. As dialogues grow, the model must retain key and value vectors for every historical token, and these can easily swell to several GB. The industry has previously tried token eviction, merging, or truncation, but their performance degrades sharply in enterprise scenarios requiring extreme compression. Another path is the Cartridges method — using gradient optimization to train compact KV caches — but each compression requires hours of GPU runtime, unsuitable for real-time applications. The MIT team took a different angle: as long as two key mathematical properties — attention output and attention quality — are preserved, the compressed cache can perfectly simulate the original behavior. Based on this, Attention Matching compresses each head's KV cache to fewer key-value pairs, achieving up to 50× compression on some datasets, taking only seconds, and requiring no training. The paper has been accepted by ICLR 2026. This technology means the cost structure of long-context services is about to see significant improvement, but the 50× number is a peak on some datasets; real-world results vary by model and task type.","mit-attention-matching-kv-cache-50x","2026-05-17T01:00:00Z","2026-05-17T01:09:14.419652Z","2026-08-19T02:08:40.142862Z",true,"agent",140,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"c83af54b-79ed-445c-9482-07d98c26c36b","BeaconKV:长推理会回头看,只压最近窗口的 KV 缓存注定丢东西","beaconkv-beacon-query-kv-cache-compression","2026-09-09T11:25:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"bce9fc16-d31a-49be-b17b-f144619a58e2","LatentPress:上下文压成软令牌直读，7.7 倍压缩反超原文，训练仅动 0.1% 参数","latentpress-soft-token-context-compression","2026-09-05T19:06:09+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"d31bc388-b6c7-41a5-a6e9-6f00657c7616","加GPU还是压KV缓存？arXiv论文：压缩省钱1.2到2倍，但36B是道坎","tensor-parallelism-vs-kv-compression-cost","2026-08-30T17:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"d2c56430-f9a7-4844-bfab-a8651279a70c","ResKV 不再把 KV 缓存压缩等同于删词：给被淘汰的信息留一份残差账本","reskv-residual-kv-cache-compression","2026-08-03T10:43:23+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"5f745fe5-ea5d-453a-8b08-7dac524d1ac2","ACL 2026 综述 sKis：KV 缓存优化重塑为 LLM serving 系统学","acl-2026-skis-kv-cache","2026-07-12T18:15:00+00:00"]