[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cmc-context-memory-embedding-long-context-compression":3,"topics-all":38,"news-related-0bb9aed0-f4f6-4d1a-974a-88e47010815e":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"0bb9aed0-f4f6-4d1a-974a-88e47010815e","长上下文压成答案导向记忆:CMC 让冻结 LLM 省一半显存","圣母大学等团队提出 CMC 软压缩框架,把长输入压成与冻结解码器对齐的 CMEs 嵌入,通过查询导向 Top-K 与局部窗口组成双层 KV 缓存,在 SQuAD 等四个 QA 基准上较基线提升最多 7.3 EM、推理时间与峰值显存最高降 20%、50%。","## 一切从长上下文的隐性成本开始\\n过去两年 LLM 推理的瓶颈,越来越不在参数规模,而在输入长度。每多塞 1k token,自注意力的计算就按平方级往上翻,KV 缓存则按线性级吃显存,延迟、能耗、GPU 占用一起爬坡。架构派(Longformer、LongLoRA、位置插值)改的是注意力本身,硬剪枝派(LLMLingua、SelectiveContext)干脆丢掉看似冗余的 token;软压缩(AutoCompressor、ICAE、Gisting、CCM、500xCompressor、PCC)走中间路线,把整段上下文编码成几组密集向量,让冻结的解码器当成 prompt 看待。但现有方法都缺一块:不是没有查询导向的内存选择,就是训练时没用上答案监督,再就是压缩器死死绑在某个解码器架构上。\\n\\nCMC(Context-to-Answer-Aligned Memory Compression)9 月 22 日挂在 arXiv(2609.25537),由圣母大学 Lucy Family Institute 与日本会津大学联合署名,试图同时把三块短板补齐——压缩器与解码器解耦、训练时引入答案监督、推理时按问题选 Top-K 记忆向量。\\n\\n## 三块组件:ContextEncoder、MemoryBridge、两层 KV 缓存\\nCMC 走的是 encoder-decoder 解耦思路:ContextEncoder 是一个独立的因果语言模型,负责把长上下文切成段,在每段里追加 \u003CMEM>...\u003C\u002FMEM> 占位 token,让 placeholder 通过自注意力吸收整段语义,输出若干隐藏向量——这就是 Context Memory Embeddings(CMEs)。MemoryBridge 是一个带 norm-calibrated 的两层 MLP,负责把 CMEs 从 encoder 的隐空间投到 frozen decoder 的嵌入空间,并把 ℓ2 范数卡在 decoder 自身词嵌入的 2 倍以内,防止注入分布漂移。\\n\\n最关键的是推理期的两层 KV 缓存策略。Tier-1 用问题向量与所有 CME 做余弦相似度,挑 Top-K 与问题最相关的记忆向量塞进 prefix;Tier-2 保留一段局部窗口原文,以全 token 精度进入解码器。这意味着 decoder 不再盯着几千 token 的完整 prompt,而是 Top-K 个 CME + 一小段原文——KV 缓存上界被压到固定预算里。\\n\\n## 训练两阶段:先对齐,再贴答案\\nPhase-1 用 AE(autoencoding)+ AR(autoregressive)重建损失,把 CMEs 拉进 decoder 的嵌入空间;Phase-2 引入 frozen decoder 作为「教师」做知识蒸馏,同时加 KL 对齐与对比学习(contrastive memory-answer alignment),让 CMEs 在嵌入空间里显式贴向答案方向。这个两阶段组合是 CMC 与 PCC 这类方法最显著的区别——PCC 只在文本重建上训,不知道压缩后的向量到底「该不该回答」。\\n\\n训练目标 CE\u002FKL\u002FCL 三项损失必须同时存在,论文里 ablation 显示去掉任何一项,EM 都会从 0.670 掉到 0.589,等于回到只做 Phase-1 的水平。换句话说,KL 是分布级对齐,CL 把 CMEs 钉到答案上,CE 提供 token 级监督——三者一损俱损。\\n\\n## 数字:7.3 EM、20% 时间、50-62.5% 显存\\n九组 encoder-decoder × 四个 QA 基准(SQuAD、AdversarialQA、HotpotQA、CovidQA)的实验里,CMC 几乎在所有配置上都比基线更强。SQuAD 上 EM 最多提升 7.3 点、F1 最多 4.0 点;Llama 与 GPT2-Large 组合下,CMC 的 EM 0.670 相对基线 0.624 净涨 4.6 点。\\n\\n效率侧才是这篇的硬菜。生成 3000 token 时,推理时间从 41 963 秒降到 33 562 秒(降 20%),能耗相应降 20.3%。Llama 解码器的峰值预留显存从 19.7 GB 降到 9.8 GB(降 50%),Mistral 更狠——24.3 GB 降到 9.1 GB,降 62.5%。预算是个 1000 样本的统计均值,不是单条数据。这些节省全部来自「缓存上界固定」这条结构性收益——长生成带来的缓存线性增长,被双层策略切成了常量。\\n\\n## 为什么 ablation 给出明确的工程信号\\n论文给出了三类 ablation。架构层面,去掉图论式上下文去噪(用 cosine sim + 阈值 τ 滤掉与邻居相似度低的 token)直接掉 26.6 EM;随机 CME 选(不用查询导向 Top-K)掉 8.1 EM;去掉 Tier-2 局部窗口是最致命的——EM 直接从 0.670 掉到 0.170。Tier-2 是结构上的不可缺件,不是锦上添花。\\n\\n压缩率 r 的扫描显示 r=4 是大多数组合的最优或并列最优,过密(r=2)反而让 Top-K prefix 噪声变大、过疏(r=8)又丢失关键上下文。值得注意的是,最优 r 略依赖 decoder:Llama 强烈偏好 r=4(对比 r=8 多涨 1.1-1.7 EM),而 Mistral 与 Gemma 在 r=4 和 r=8 之间几乎平。\\n\\n## 局限与边界\\n论文自己列了三条限制:只测了抽取式 QA、没碰摘要和 RAG;压缩率 r ∈ {2,4,8} 是固定的,不会随上下文长度自适应;只在英文上验证。代码已开源(github.com\u002Fmostafiz26\u002FCMC),但训练仍需要把 frozen decoder 作为教师拉进来做蒸馏,工程成本比纯自监督软压缩高一个量级。\\n\\n## 所以呢\\nCMC 把软压缩从「压缩完了就交给 decoder」推进到「压缩方向由答案定、缓存上界由结构定」。它的真正信号是:在不重训 decoder 的前提下,长上下文推理的显存曲线可以被压平,这对 Agent、长文档 RAG、企业级检索这些真把 context 推到 100k+ 的场景是直接相关的工程红利——你不必把上下文压缩进 500xCompressor 那种极致 token 压缩里,只要保留一段原文窗口 + 一些对齐到答案方向的 CME 嵌入,显存就能砍一半。代码是开放的,蒸馏管线可以挂到任何 frozen decoder 上,值得把 frozen Llama\u002FMistral\u002FGemma 跑一遍看看自己的领域数据能不能复现这个 50-62.5% 的显存节省。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25537","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",{"id":19,"name":20,"slug":20,"description":14,"color":14},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"93b26bbd-668e-4e66-b4a8-a55d6859a302","en","CMC: Long Context Compressed into Answer-Aligned Memory, Halving Frozen LLM VRAM","A team from the University of Notre Dame and others propose CMC, a soft-compression framework that compresses long input into CMEs aligned with a frozen decoder's embedding space. A query-guided Top-K plus a local window form a two-tier KV cache. Across four QA benchmarks including SQuAD, CMC lifts EM by up to 7.3 points and cuts inference time and peak reserved memory by up to 20% and 50%.","## The Hidden Cost of Long Context\\nOver the past two years, the bottleneck of LLM inference has shifted away from parameter count and toward input length. Every additional 1k tokens pushes self-attention computation up quadratically and KV cache up linearly; latency, energy, and GPU footprint all climb together. Architectural approaches (Longformer, LongLoRA, positional interpolation) modify attention itself; hard-prompt pruning (LLMLingua, SelectiveContext) drops tokens deemed redundant; soft-compression methods (AutoCompressor, ICAE, Gisting, CCM, 500xCompressor, PCC) take a middle path, encoding the entire context into a small set of dense vectors that a frozen decoder treats as a prompt. But each existing approach is missing a piece: either no query-guided memory selection, no answer-targeted supervision during training, or a compressor tightly coupled to a specific decoder architecture.\\n\\nCMC (Context-to-Answer-Aligned Memory Compression) appeared on arXiv (2609.25537) on September 22, jointly authored by Notre Dame's Lucy Family Institute and the University of Aizu, and attempts to fill all three gaps at once: decouple compressor from decoder, inject answer supervision during training, and pick Top-K memory vectors at inference time.\\n\\n## Three Components: ContextEncoder, MemoryBridge, Two-Tier KV Cache\\nCMC takes an encoder-decoder decoupled route. ContextEncoder is an independent causal language model that chunks long context into segments, appends \u003CMEM>...\u003C\u002FMEM> placeholder tokens to each segment, lets those placeholders absorb segment semantics through self-attention, and outputs a few hidden vectors — these are the Context Memory Embeddings (CMEs). MemoryBridge is a norm-calibrated two-layer MLP that projects CMEs from the encoder's hidden space into the frozen decoder's embedding space, capping the ℓ2 norm at twice the decoder's own token embedding norm to prevent distribution drift.\\n\\nThe most important piece is the inference-time two-tier KV cache strategy. Tier-1 runs cosine similarity between the question vector and all CMEs, picks the Top-K most relevant memory vectors, and injects them into the prefix. Tier-2 retains a local window of original text at full token precision for the decoder. The decoder no longer stares at thousands of tokens of full prompt; instead it gets Top-K CMEs plus a small slice of original text — and the KV cache upper bound collapses to a fixed budget.\\n\\n## Two Training Phases: Align First, Then Anchor to Answers\\nPhase-1 uses AE (autoencoding) plus AR (autoregressive) reconstruction losses to pull CMEs into the decoder's embedding space. Phase-2 brings in the frozen decoder as a teacher for knowledge distillation, while adding KL alignment and contrastive memory-answer alignment, explicitly pinning CMEs toward the answer direction in embedding space. This two-phase setup is the sharpest difference from methods like PCC — PCC only trains on text reconstruction and has no idea whether the compressed vectors actually \"ought to answer.\"\\n\\nThe CE\u002FKL\u002FCL losses must all be present at once: ablations show that removing any one drops EM from 0.670 to 0.589, equivalent to going back to Phase-1 only. KL provides distribution-level alignment, CL pins CMEs to the answer, CE delivers token-level supervision — losing any one takes down the others.\\n\\n## Numbers: 7.3 EM, 20% Time, 50-62.5% VRAM\\nAcross nine encoder-decoder pairings and four QA benchmarks (SQuAD, AdversarialQA, HotpotQA, CovidQA), CMC outperforms the baseline in nearly every configuration. On SQuAD, EM improves by up to 7.3 points and F1 by up to 4.0 points. With the Llama and GPT2-Large pairing, CMC reaches EM 0.670, a net 4.6-point gain over the baseline's 0.624.\\n\\nThe efficiency side is the real headline. At 3,000 generation tokens, inference time drops from 41,963 seconds to 33,562 seconds (down 20%), with energy dropping 20.3% correspondingly. Peak reserved VRAM for the Llama decoder falls from 19.7 GB to 9.8 GB (down 50%), and Mistral is even more dramatic — 24.3 GB to 9.1 GB, down 62.5%. These are 1,000-sample statistical means, not single-point measurements. All the savings come from one structural property: a fixed cache upper bound — the linear cache growth from long generation is cut down to a constant.\\n\\n## What the Ablations Tell Engineers\\nThe paper runs three ablation categories. On the architecture side, removing graph-based context denoising (cosine-similarity + threshold τ to filter low-salience tokens) drops EM by 26.6 points; random CME selection (no query-guided Top-K) drops EM by 8.1 points; and removing the Tier-2 local window is the most fatal — EM falls from 0.670 to 0.170. Tier-2 is structurally indispensable, not a nice-to-have.\\n\\nThe compression-rate r sweep shows r=4 is best or tied-best for most pairings; too dense (r=2) actually adds noise to the Top-K prefix, while too sparse (r=8) loses critical context. Notably, the optimal r is mildly decoder-dependent: Llama strongly prefers r=4 (1.1-1.7 EM over r=8), while Mistral and Gemma are nearly flat between r=4 and r=8.\\n\\n## Limitations and Boundaries\\nThe paper itself lists three: only extractive QA is tested, no summarization or RAG; compression rate r ∈ {2,4,8} is fixed rather than adaptive to context length; and validation is English-only. Code is open (github.com\u002Fmostafiz26\u002FCMC), but training still requires pulling in the frozen decoder as a teacher for distillation — engineering cost is an order of magnitude higher than pure self-supervised soft compression.\\n\\n## So What\\nCMC advances soft compression from \"compress and hand to the decoder\" to \"compression direction is set by the answer, cache upper bound is set by the structure.\" The real signal: without retraining the decoder, the VRAM curve of long-context inference can be flattened. This is directly relevant engineering value for agents, long-document RAG, and enterprise search — scenarios that genuinely push context past 100k. You don't have to go to extreme token compression like 500xCompressor; keeping a local window of original text plus a few answer-aligned CME embeddings cuts VRAM in half. The code is open and the distillation pipeline can plug into any frozen decoder — worth running frozen Llama\u002FMistral\u002FGemma against your own domain data to see whether you can reproduce that 50-62.5% VRAM saving.","cmc-context-memory-embedding-long-context-compression","2026-10-07T00:00:00Z","2026-10-07T09:08:05.049804Z","2026-10-07T09:08:05.049819Z",true,"agent",170,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"ddbac1f6-08ab-4663-ab2b-da6793947e49","Kara:把 KV 缓存压成「滑动窗口」,让推理 LLM 在高并发下不再卡顿","kara-kv-cache-sliding-window","2026-07-05T06:10:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"5fa17d26-6090-4805-893a-aa880cf33369","让 Agent 学会遗忘:删掉旧推理,分数反而涨了","agents-forget-reasoning-iclr-compression","2026-09-28T19:30:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"bc642ebf-9b1c-41cf-98d1-dd55393fd429","HeadWiseKV:无训练KV cache压缩让混合LLM长上下文从114K推到161K","headwisekv-training-free-kv-cache-compression","2026-09-03T03:44:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"7bb4f5ec-14e0-43b6-9913-07cad82a520b","微软内部 AI 账单失控:单员工月烧 2.8 万美元,倒逼默认模型换人","microsoft-internal-ai-bill-explode-default-model-swap","2026-08-28T04:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"6fa1bc74-c98e-476b-bc4c-9ae057105ffb","ParaTempo:免训练并行推理,延迟最高降 32%、token 省三成","paratempo-temporal-confidence-parallel-reasoning","2026-08-24T17:20:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"ddb7bc6c-6b6e-4797-ab76-d1aeab5a3002","压缩得好≠部署得好:树莓派实测边缘 LLM,LoRA恢复模型100题押97个同答案","edge-llm-compression-raspberry-pi","2026-08-23T13:30:00+00:00"]