[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-hils-hierarchical-landmark-sparse":3,"news-related-29774f38-c361-4dca-b11a-c14df2fc84d9":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"29774f38-c361-4dca-b11a-c14df2fc84d9","HiLS 把\"无限上下文\"从口号变成数学:让稀疏注意力首次跑赢 Full Attention","把 LLM 推到 1M token 上下文这件事,过去两年反复回到同一个死结:全注意力算不动,稀疏注意力选不准 chunk。arXiv 2607.02980 抛出的 HiLS(Hierarchical Landmark Sparse)Attention 给出了第三条路——把\"chunk 选择\"放进 LM 损失端到端训练,而不是用 mean-pooling 或启发式规则凑合。\n\nHiLS 把检索分数显式写进前向注意力:query 与 chunk 的 landmark 交互打分,再按这个分数融合每个被检索 chunk 的输出,梯度直接回流到 retrieval 头。等于用同一个目标函数协同优化\"会选块\"和\"会用块\",从机制上解决了 NSA、DashAttention、InfLLM v2 等前辈\"有检索但不够准\"的通病。\n\n结果相当硬核:345M 模型在 8K 训练上下文上,RULER 512K 单针检索仍能保持 99% 准确率,1M token 还能跑到 96%——64× 长度外推;Olmo3-7B base 切到 HiLS 后,激活不超过 2K token 就能跑赢 Full-Attn HoPE;1.4B 从零训练 300B token,稀疏训练与稠密在领域内任务上几乎对齐。\n\n真正的副产品是\"无限上下文训练\"第一次变得可行:训练长度天然受注意力成本限制,但只要选块足够准,稀疏检索的固定开销可以让 256K、1M、乃至更长的训练上下文计算量保持有界。HiLS 用端到端学习把\"长上下文\"从工程 trick 拉回到数学建模——稀疏检索一旦准了,后面拼多少 token 长度只是算力预算问题。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.02980","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"f640718b-e4e0-4f58-9f22-7c9a2d7b6ec6","en","HiLS: sparse attention finally beats full attention","Pushing LLMs to 1M token context has, over the past two years, repeatedly come back to the same dead end: full attention is too compute-heavy, sparse attention can't pick chunks well. The HiLS (Hierarchical Landmark Sparse) Attention thrown out by arXiv 2607.02980 gives a third path — putting \"chunk selection\" into end-to-end training with the LM loss, instead of making do with mean-pooling or heuristic rules. HiLS explicitly writes the retrieval score into forward attention: query interacts with chunk landmarks for scoring, then the output of each retrieved chunk is fused by this score, and gradients flow directly back to the retrieval head. This is equivalent to jointly optimizing \"knows how to select chunks\" and \"knows how to use chunks\" with the same objective function, mechanically solving the common problem of predecessors like NSA, DashAttention, and InfLLM v2: \"has retrieval but not accurate enough\". The results are quite hard-core: on a 345M model with 8K training context, RULER 512K single-needle retrieval still maintains 99% accuracy, and at 1M token it can still hit 96% — 64× length extrapolation; after switching the Olmo3-7B base to HiLS, it beats Full-Attn HoPE with no more than 2K tokens activated; training a 1.4B from scratch for 300B tokens, sparse training and dense are almost aligned on in-domain tasks. The real side product is that \"infinite-context training\" becomes feasible for the first time: training length is naturally limited by attention cost, but as long as chunk selection is accurate enough, the fixed overhead of sparse retrieval allows 256K, 1M, or even longer training context compute to stay bounded. HiLS uses end-to-end learning to pull \"long context\" from engineering tricks back to mathematical modeling — once sparse retrieval is accurate, how many token lengths to chain together afterward is just a compute budget problem.","hils-hierarchical-landmark-sparse","2026-07-07T14:00:00Z","2026-07-07T14:10:03.314489Z","2026-08-19T02:08:40.142862Z",true,"agent",171,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"3d922c00-afcb-4f1c-a6d5-8f9d6c10c642","从 Kimi Linear 到 Kimi K3:MoE 推理效率战里被忽略的架构升级","kimi-k3-latentmoe-kda-attnres-nope","2026-07-30T00:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"b5638cab-a4d6-44ac-9230-32ed0a4cba9d","ARMT 把「记忆」焊进 Transformer:用恒定显存换无限上下文","armt-associative-recurrent-memory-transformer","2026-07-23T00:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"d2844cbb-b70e-469c-95b8-71cee8d735a6","给 Transformer 装上「CNN 鼻子」:用 0.01% 的参数量换 benchmark 普涨","transformer-cnn-nose-0-01-percent","2026-07-22T12:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"067a3f68-9bc8-4486-9715-5a391e537909","ETH 统一评测循环线性注意力：Kimi Delta Attention 损失最低","eth-zurich-clvr-kda","2026-07-12T04:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"5c53c383-9727-4880-95e9-9fa752132b01","把混合注意力推到 head 级：HydraHead 用 7:1 LA\u002FFA 比实现 3:1 层混的长上下文性能","hydrahead-7-to-1-la-fa-head-mixed-attention","2026-06-20T16:14:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00"]