[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-lema-exact-match-attention":3,"topics-all":38,"news-related-0bdbdcd9-fbb5-4123-9854-57b8f28b385b":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"0bdbdcd9-fbb5-4123-9854-57b8f28b385b","把注意力退回查字典:LEMA让KV缓存离开显存","图宾根大学的 LEMA 把注意力推到极端:查询与键全部二值化,每个查询只取最近一次精确匹配的键,KV 缓存变成主内存里的哈希表,生成速度恒定。理论上它与 word-RAM 双向等价;834M 语言模型长程召回超过 gated DeltaNet,但整体损失仍落后 softmax 注意力。","KV 缓存是长上下文推理绕不开的两难:上下文越长,softmax 注意力要读的键值越多,生成越慢;把它们全塞进显存,显存又先撑不住。线性注意力和状态空间模型把上下文压进固定大小的状态,速度和显存都恒定了,但固定状态的容量上限让召回能力随信息量增长而崩掉。德国图宾根大学数学系的 Moritz Brösamle 在 arXiv 提出的 LEMA(Latest Exact Match Attention)选了第三条路:状态可以无限增长,但不放显存、也不逐条扫描——直接当字典查。(https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25802)\n\n## 注意力退化成一次哈希查找\n\nLEMA 的规则极简:查询和键先二值化,每个查询只取最近一次精确匹配的键,匹配不到就返回零。论文给出的生成步伪代码只有几行——查一次字典、写一次字典,同键覆盖旧值。于是每个注意力头的 KV 缓存就是一个字典:状态大小随内容增长,理论上限是 2 的头维度次方个条目,实际上只存出现过的键。\n\n## 理论:与 word-RAM 双向等价\n\n论文证明,带思维链的 LEMA transformer 可以模拟 word-RAM(现代计算机的抽象模型);反过来,word-RAM 也能以每 token 与上下文长度无关的开销模拟 LEMA transformer。作者称据其所知,这种双向对应在此前的注意力变体里没有建立过——硬注意力家族的表达力结果通常只证单向。\n\n## 训练:从软注意力退火到硬规则\n\n精确匹配不可导,训练是最大难点。论文用直通估计器传二值化梯度,再用 stick-breaking 注意力做软代理,逐步退火到 LEMA。作者坦承这只是第一次尝试:退火对学习率敏感,且训练仍需与序列长度平方成正比的计算量。\n\n## 实测:召回超 GDN,整体仍落后 softmax\n\n在合成联想召回任务上(词表 4096),只在 8 对关联上训练的 LEMA 几乎完美外推到 4096 对,固定状态的 gated DeltaNet(GDN)则在关联数变大后失效。在 FineWeb-Edu 上训练的 29M 到 834M 参数语言模型里,LEMA 的验证损失大约追平参数量为其 55%-57% 的 softmax 模型。两个长程召回代理测试更能说明差异:重复稀有 bigram 的最远距离桶里,LEMA 损失比自身基线低 2.1 nats,GDN 只低 0.7;RULER 单针检索(S-NIAH-1)上,LEMA 一旦检索成功,重复填充句不再改变状态,检索可以无限持续而不增长状态。\n\n## 工程:哈希表进内存,生成速度恒定\n\n推理实现把所有头的 KV 缓存做成主内存里一张开放寻址哈希表(线性探测),显存只留权重和激活。在 RTX 3090 上与 vLLM 基准对比,只要哈希表不接近容量上限,LEMA 生成速度恒定、与 GDN 相当。834M 训练模型每头字典在 16k token 后平均 1.7k 条,256k 后 18k 条——论文测试用了 50 GB 的主内存哈希表。代价也有:约 10% 的头从未找到匹配。代码已开源(github.com\u002Fmoritzbroe\u002Flatest_exact_match_attention)。\n\n把\"更像计算机的注意力\"和\"能训练的注意力\"接起来,LEMA 给出了一条少见的路线:召回靠无限状态,速度靠 O(1) 查找,显存压力靠内存转嫁。834M 规模的差距说明训练方法远未成熟,但 KV 缓存焦虑的解法未必是更聪明的压缩——也可能是干脆换一种数据结构。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25802","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"b7ab5718-6d3f-4de3-87e3-31a4f318d3ac","en","LEMA Reduces Attention to Dictionary Lookup and Moves KV Cache Off VRAM","LEMA, from the University of Tubingen, takes attention to an extreme: queries and keys are binarized, and each query retrieves only the latest exactly matching key. The KV cache becomes a hash table in main memory, keeping generation speed constant. In theory it is bidirectionally equivalent to word-RAMs; an 834M language model beats gated DeltaNet on long-range recall, though overall loss still trails softmax attention.","The KV cache is the unavoidable dilemma of long-context inference: as context grows, softmax attention must read more keys and values, slowing generation; storing them all in VRAM breaks memory first. Linear attention and state space models compress context into a fixed-size state, making speed and memory constant, but the fixed capacity ceiling makes recall collapse as information grows. Moritz Brosamle from the Department of Mathematics at the University of Tubingen proposes LEMA (Latest Exact Match Attention) on arXiv, choosing a third path: the state can grow without bound, but it lives neither in VRAM nor gets scanned token by token — it is simply looked up as a dictionary. (https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.25802)\n\n## Attention reduced to one hash lookup\n\nThe LEMA rule is minimal: queries and keys are binarized first, and each query attends only to the latest exactly matching key, returning zero when nothing matches. The generation-step pseudocode in the paper is a few lines — one dictionary lookup, one dictionary insert, with an existing key overwriting its old value. The KV cache of each attention head thus becomes a dictionary: state size grows with content, with a theoretical bound of 2 to the power of head-dimension entries, while in practice only keys that occur are stored.\n\n## Theory: bidirectional equivalence with word-RAMs\n\nThe paper proves that LEMA transformers with chain of thought can simulate word-RAMs, an abstraction of modern computers; conversely, word-RAMs can simulate LEMA transformers at a cost per token independent of context length. The author states that, to their knowledge, this two-way correspondence has not been established for other attention variants — expressivity results for the hard-attention family usually prove only one direction.\n\n## Training: annealing from soft attention to the hard rule\n\nExact matching is non-differentiable, making training the hardest part. The paper passes gradients through binarization with a straight-through estimator, then uses stick-breaking attention as a soft surrogate, slowly annealing it towards LEMA. The author admits this is merely a first attempt: the annealing is sensitive to learning rates, and training still requires compute quadratic in sequence length.\n\n## Results: beats GDN on recall, still trails softmax overall\n\nOn a synthetic associative recall task (vocabulary 4096), LEMA trained only at 8 associations extrapolates almost perfectly to 4096, while the fixed-state gated DeltaNet (GDN) fails once associations grow too numerous. On language models from 29M to 834M parameters trained on FineWeb-Edu, LEMA matches the validation loss of softmax transformers at 55-57% of its parameters. Two long-range recall proxies show the difference more clearly: in the farthest bucket of repeated rare bigrams, LEMA loss sits 2.1 nats below its own baseline versus 0.7 for GDN; on RULER single-needle retrieval (S-NIAH-1), once LEMA retrieves successfully, repeated filler sentences no longer change its state, so retrieval persists indefinitely without state growth.\n\n## Engineering: hash table in memory, constant generation speed\n\nThe inference implementation packs all heads KV caches into a single open-addressing hash table (linear probing) in main memory, leaving only weights and activations in VRAM. Benchmarked against vLLM on an RTX 3090, LEMA generates at constant speed comparable to GDN as long as the hash table stays away from capacity. For the trained 834M model, per-head dictionaries hold on average 1.7k entries after 16k tokens and 18k after 256k — the paper tests with a 50 GB hash table in main memory. The cost: around 10% of heads never find a match. Code is open-sourced (github.com\u002Fmoritzbroe\u002Flatest_exact_match_attention).\n\nConnecting \"attention that resembles a computer\" with \"attention that can be trained\", LEMA offers a rare route: recall via unbounded state, speed via O(1) lookup, and VRAM relief via main memory. The gap at the 834M scale shows the training method is far from mature, but the cure for KV-cache anxiety may not be smarter compression — it may be a different data structure altogether.","lema-exact-match-attention","2026-09-25T15:15:00Z","2026-09-25T15:11:29.971080Z","2026-09-25T15:11:29.971091Z",true,"agent",1253,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"282f3cb4-429c-4432-b632-5e7288192eb9","MassAlloc注意力:按质量分配算力,反传快3倍","massalloc-attention-mala","2026-09-29T21:08:59+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"b4f270b3-db43-4586-a0e5-a062320c6d1b","让模型自己声明看哪里:Declarative Attention 零训练砍 52% KV 读取","declarative-attention-kv-cache-declare","2026-09-03T23:07:03+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"3d922c00-afcb-4f1c-a6d5-8f9d6c10c642","从 Kimi Linear 到 Kimi K3:MoE 推理效率战里被忽略的架构升级","kimi-k3-latentmoe-kda-attnres-nope","2026-07-30T00:30:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"926f89fc-5ed5-4170-bfac-931d3a31b6a4","腾讯混元开源 AngelSpec 投机解码框架：DFly 在 Hy3-A21B 上取得 1.98–2.40× 加速","tencent-angelspec-spec-decoding-hy3-dfly","2026-07-30T00:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"070aef27-5fdb-4f0a-8b90-99afc1ea34fb","Jet-Long 用「动态双焦 RoPE」让 Qwen3 免训练扩到 128K,RULER 直接多涨 4.79 pp","jet-long-dynamic-dual-rope","2026-07-12T02:30:00+00:00"]