[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-engramedit-decoupled-knowledge-llm":3,"topics-all":41,"news-related-fc533af5-9e8a-43c2-8408-2f931a6aa398":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"fc533af5-9e8a-43c2-8408-2f931a6aa398","EngramEdit:把事实塞进 LLM 的记忆抽屉","arXiv 2610.10533 提出的 EngramEdit,在 DeepSeek Engram 条件记忆架构上做事实编辑:无需重训 Transformer 主干,实现近 100% 编辑成功率,在 CoT 多跳推理中准确率是最强 baseline 的近 3 倍,且不破坏无关知识与通用能力。","## 为什么大模型「知识更新」是个老难题\n\n把一条新事实写进 LLM,传统做法是继续微调或做知识编辑。问题在于:大模型把事实压进了上千亿参数的 Transformer 权重里,你改一处可能牵动十处——编辑成功率上去了,无关能力就开始塌方。这是 model editing 这个子领域十年来没真正解掉的根结。\n\narXiv 2610.10533 提出的 EngramEdit 把解题思路换了一条:不碰 Transformer 主干,只改「条件记忆」里那一份外挂的 n-gram embedding 表。\n\n## 什么是条件记忆\n\n条件记忆架构(代表工作:DeepSeek Engram)在 Transformer 之外,挂一张可学习的 n-gram embedding 表。推理时,模型先用输入的 n-gram 去查表,取出向量再和 Transformer 一起算。这相当于给 LLM 加了一个「外挂硬盘」,只增加有限计算,就把模型容量撑了上去。Engram 的早期工作把这条路用来做 scaling,这篇新论文想做的是另一个方向:把那张表变成「可编辑的事实接口」。\n\n## 两步走的编辑方法\n\n直接改条件记忆的 embedding 也不是简单事。同一条事实在不同表达下会激活不同的 n-gram;而一个 n-gram embedding 又往往被多条事实共用——粗暴更新必然误伤。这套新方法的解法分两步:\n\n第一步,它先用当前模型,反推出「要让模型在新事实的多条表达上都预测正确,目标 memory 向量应该是怎样的」。这一步得到的是一组虚拟的目标 embedding。第二步,它联合更新共享的 n-gram embedding,让它们尽量逼近这些目标,但对「高频共用」的 embedding 加更大的惩罚——高频被多事实依赖,改它会牵动太多无关知识,所以得让它更稳。\n\n这个「按使用频率加权」的思路,本质是把编辑问题当成一个带正则项的最小二乘,而不是一次单点改写。\n\n## 论文里报的数字\n\narXiv 2610.10533 给出的实验结果是:事实编辑的编辑成功率接近 100%。被编辑过的知识,在未见过的表达上能用,在多跳推理里也能用——CoT prompting 下,准确率接近最强 baseline 的 3 倍。多次编辑叠加时,无关知识和通用能力基本不掉。\n\n需要注意的是,这些数字来自 DeepSeek Engram 这类带条件记忆模块的架构,不是对通用 Transformer 的 silver bullet——没装 Engram 的标准 LLM 用不上这套机制。\n\n## 为什么这件事对 LLM 工程重要\n\n把「存储事实」从「通用计算」里剥出来,一直是模型架构想要的方向。这篇工作的价值不在某一个 benchmark 数字,而在它证明了:条件记忆这种外挂结构,既能用来扩容量,也能用来做外科手术式的事实修订。配合 RAG 和工具调用,这种「主模型不动、外挂记忆可改」的范式,可能会成为长寿命 LLM 的标配——明年 Qwen5、Claude 7 这种计划跑很多年的模型,如果要在不重新预训练的前提下吸收 2027 年的新事实,EngramEdit 这种思路就是基础设施级别的拼图。\n\n(本文事实均来自 arXiv 2610.10533 论文摘要与正文,引用见 https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.10533)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.10533","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b52db7e9-7c58-42c3-9536-5132cb2f8f72","deepseek",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":25,"name":26,"slug":26,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"4c93a861-9e5b-419f-b109-84f281ab0406","en","EngramEdit: surgical knowledge edits via conditional memory","arXiv 2610.10533: edit DeepSeek Engram conditional memory, no retraining. ~100% edit success, ~3x baseline on CoT multi-hop.","## Why knowledge updates are an old problem for LLMs\n\nThe conventional way to add a new fact to an LLM is continued fine-tuning or model editing. The issue is that LLMs compress facts into hundreds of billions of Transformer weights, so editing one place tends to disturb ten others. Edit success goes up while unrelated capability quietly collapses. This is the root tension that the model-editing subfield has been unable to truly resolve for a decade.\n\nEngramEdit (arXiv 2610.10533) takes a different route: leave the Transformer backbone untouched and edit only the external n-gram embedding table inside the conditional memory module.\n\n## What conditional memory actually is\n\nConditional memory architectures (the canonical example being DeepSeek Engram) attach a learnable n-gram embedding table outside the Transformer. At inference, the model looks up embeddings by input n-grams and feeds them into the Transformer alongside the usual tokens. It is effectively an external hard drive bolted onto the LLM: model capacity grows while extra compute stays bounded. Earlier Engram work used this path for scaling; the new paper points it at a different problem — turning that table into an editable factual interface.\n\n## The two-step edit\n\nNaively rewriting the conditional-memory embeddings is itself hard. The same fact expressed differently activates different n-grams, and a single n-gram embedding is typically shared by many facts, so a blunt update will collateral-damage unrelated knowledge. The proposed method proceeds in two steps.\n\nFirst, given the current model, it back-computes the target memory vectors that would make the model predict the updated fact across multiple expressions. That yields a set of virtual target embeddings. Second, it jointly updates the shared n-gram embeddings to match these targets, while penalising updates to frequently reused embeddings more strongly — high-frequency embeddings are depended on by many facts, so they need to be more conservative.\n\nThis 'weighted by usage frequency' framing is essentially least-squares with a regularisation term, not a single-point rewrite.\n\n## What the paper reports\n\nThe arXiv 2610.10533 experiments report an edit success rate near 100%. Edited knowledge is usable on unseen expressions and in multi-hop reasoning — under chain-of-thought prompting, accuracy is nearly three times that of the strongest baseline. Stacking many edits does not visibly erode unrelated knowledge or general capability.\n\nIt is worth flagging that these numbers come from architectures equipped with a conditional memory module like DeepSeek Engram; they do not transfer to vanilla Transformers that lack the Engram machinery.\n\n## Why this matters for LLM engineering\n\nDecoupling factual storage from general-purpose computation has long been a goal of model architecture. The value of this paper is not any single benchmark number but the proof that the conditional-memory structure can serve both as a capacity lever and as a surgical fact-editing surface. Combined with RAG and tool use, a paradigm of 'frozen backbone, editable external memory' is likely to become a default for long-lived LLMs — models like Qwen 5 or Claude 7, planned to run for many years, will need to absorb new facts in 2027 without a full retraining run, and this style of work is exactly the infrastructure piece that fits.\n\n(Facts in this piece are drawn from the abstract and body of arXiv 2610.10533; reference at https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.10533)","engramedit-decoupled-knowledge-llm","2026-10-09T04:00:00Z","2026-10-09T07:19:21.342236Z","2026-10-09T07:19:21.342250Z",true,"agent",44,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"cb7fb8b3-5862-4cba-adab-c4794e989966","图灵奖得主 Pearl 长访谈：LLM 能讲因果只是因为人类替它爬过了因果阶梯","judah-pearl-llm-causal-ladder-agi","2026-07-31T07:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"f8436dd3-d6fc-4ea7-9f2e-1086026c11d0","Transformer 的几何之眼：arXiv 2607.17146 把注意力炼成薛定谔桥，把 SGD 写成伊藤扩散","transformer-geometry-schrodinger-bridge","2026-07-23T12:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"e54f030e-14ed-4262-9dd9-8685fdbb03ab","DiscoLoop 把循环 Transformer 的「表征瓶颈」焊死:双通道架构让多跳推理一步到位","discoloop-dual-channel-recurrent","2026-07-20T08:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"f8ea285b-f717-4f82-aafd-6a096ca6cf46","DeepLoop：Princeton\u002FUCLA 修对 Looped Transformer 残差缩放","princeton-ucla-deeploop","2026-07-17T22:13:48+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"e40de2f9-9ee4-47f3-b1f1-7bf93a4870a9","一个动词翻转工具调用决策,LLM 内部向量现形","llm-tool-call-decision-vector","2026-10-09T13:10:00+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"c368ad9f-9308-4a0b-8f5c-3ae4601b48b9","D2K-Bench: 专家设计把 LLM 写 GPU 核提速 33.9%","d2k-bench-llm-gpu-kernel-design-guidance","2026-10-07T03:00:00+00:00"]