[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-massalloc-attention-mala":3,"topics-all":38,"news-related-282f3cb4-429c-4432-b632-5e7288192eb9":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"282f3cb4-429c-4432-b632-5e7288192eb9","MassAlloc注意力:按质量分配算力,反传快3倍","BAAI提出MassAlloc Attention:QK打分全保留,post-score计算由softmax质量在线分配。128K 8×H100基准下训练前向\u002F反向降至1\u002F2.2与1\u002F3.0,解码快1.6倍;14B训练总FLOPs省23.1%。","稠密注意力有个结构性浪费:softmax 归一化后,大部分因果分数空间拿到的质量小到可以忽略,但稠密 kernel 仍要为每个 QK tile 走完打分后的整条流水线——归一化、乘 V、写回。分数算完才发现这块区域不重要,但计算已经花掉。BAAI 团队 9 月 26 日挂在 arXiv 的新论文 MassAlloc Attention(MALA)盯上的正是这段被浪费的 post-score 计算,数日冲上 Hugging Face Daily Papers 榜首,累计 763 赞。\n\n## QK 全算,后面按质量分配\n\nMALA 的思路分两半。第一半,不砍 QK 打分——每个合法因果交互保留 score 访问,守住注意力分布的完整性。第二半才是重点:打分后的计算不再全员执行,而是按归一化贡献分配。前向时,MALA 借用演化中的 online-softmax normalizer 实时估计各区域质量占比,低贡献区域的后续计算直接跳过;反向复用定稿的 normalizer 推导嵌套保留支撑,只用标准注意力状态。一个统一的 tolerance 参数同时治理训练和推理——跳多少、留多少,两边同一把尺子。\n\n## 提速多少:反传 3 倍,训练 FLOPs 省 23%\n\n数字相当扎实。8K matched-work 研究里,总 post-score 工作量对齐后,MALA 的平均省略质量是 0.0188%,逐实例最优的 reference-mass oracle 是 0.0182%,几乎贴着理论上限跑。关联记忆任务上,MALA 在 8K 拿到 89.67% 准确率,全量注意力 FullAttn 是 89.97%,差 0.3 个百分点。工程侧,128K 张量并行算子基准里,MALA 把训练前向延迟压到 1\u002F2.2、反向压到 1\u002F3.0,推理解码快 1.6 倍;作者在 Hugging Face 社区帖补充测试环境为 8×H100、TP=8。规模验证:0.6B 到 14B 的 scaling-law 训练中困惑度全程贴近 FullAttn,总 FLOPs 下降;14B 在 32K 上下文训练省 23.1% 总 FLOPs,另有 32B 继续训练模型,知识、推理、长上下文检索与 FullAttn 相当。\n\n## 和稀疏注意力不是一条路\n\n这条赛道不缺新面孔:LEMA 把 KV 缓存挪出显存,HyQuant 做注意力混合精度量化,Video DeltaNet 给视频生成换混合注意力,都在啃同一个问题。MALA 的差异点是不预设稀疏结构:固定窗口、top-k 要人事先告诉模型稀疏在哪,MALA 把决定权交还给 softmax 统计本身——分布已经告诉你哪里重要,照着分配就行。另一个容易被忽略的细节是反向提速(3.0 倍)大于前向(2.2 倍),这本质是训练场算法,对预训练成本占大头的大厂,23% 的 FLOPs 直接换算成真金白银。\n\n## 冷水也得泼\n\n论文页 Models citing this paper 为 0,Hugging Face 页面也没放出代码仓库链接,第三方复现暂时无从谈起——目前所有基准数字都是作者自报。tolerance 超参的敏感性、跨架构迁移性,要等工件落地才有答案。同团队还挂了姊妹篇 CoWindow Attention(arXiv:2609.32704),让多个注意力头分担历史访问,14B 训练省 28.5% FLOPs,两篇对照可看清 BAAI 的完整布局。\n\n注意力计算预算的分配权,正从人拍板的固定规则移交给分布自己的统计量。如果 MALA 的 tolerance 机制进主流训练框架,长上下文训练的成本曲线会再弯一次——下一次 128K 级预训练的账单,可能就少一截。\n\n参考:arxiv.org\u002Fabs\u002F2609.32712","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.32712","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"363c051d-0185-43e0-a44a-56e26e3d8510","en","MassAlloc Attention cuts 128K backward latency 3x","BAAI's MALA allocates post-score compute by softmax mass: 3.0x faster 128K backward, 23.1% fewer training FLOPs at 14B, accuracy tracks FullAttn.","Dense attention has a well-known structural waste: after softmax normalization, much of the causal score space receives negligible mass, yet dense GPU kernels still execute the complete post-score path for every QK tile — normalize, multiply by V, write back, no step skipped. The scores are computed, the kernel then discovers \"this region barely matters,\" but the compute is already spent. A new paper from the BAAI team, posted to arXiv on September 26, targets exactly this wasted post-score computation. Titled MassAlloc Attention (MALA), it climbed to the top of Hugging Face Daily Papers within days, accumulating 763 upvotes.\n\n## Keep Every QK Score, Allocate the Rest by Mass\n\nMALA's idea splits in two. First, it never prunes QK scoring — every legal causal interaction keeps score access, preserving the integrity of the attention distribution. The second half is the actual contribution: computation after scoring is no longer executed uniformly but allocated by normalized contribution. During the forward pass, MALA borrows its evolving online-softmax normalizer to estimate each region's share of mass in real time; low-contribution regions skip subsequent computation entirely. The backward pass reuses the finalized normalizer to derive nested retained support, using only standard attention state — no auxiliary memory structures. A single tolerance parameter governs both training and inference: how much to skip and how much to keep is decided by the same ruler on both sides.\n\n## The Numbers: 3x Backward, 23% Training FLOPs Saved\n\nThe paper's numbers are solid. In a matched-work study at 8K context, with total post-score work exactly matched, MALA's mean omitted mass is 0.0188% versus 0.0182% for a per-instance reference-mass oracle — running nearly flush against the theoretical ceiling. On associative-recall tasks, MALA reaches 89.67% accuracy at 8K context against 89.97% for FullAttn, a gap of just 0.3 percentage points. On the engineering side, in an attention-operator benchmark at 128K tokens with tensor parallelism, MALA cuts training forward latency by 2.2x and backward by 3.0x, and speeds up inference decoding by 1.6x; the authors noted on the Hugging Face paper page that the test environment was 8xH100 with TP=8. Scale validation matters more: across scaling-law training from 0.6B to 14B parameters, MALA tracks FullAttn in perplexity while reducing total training FLOPs; the 14B model training at 32K context saves 23.1% of total FLOPs, and a separately continued-trained 32B model achieves comparable knowledge, reasoning, and long-context retrieval scores to FullAttn.\n\n## Not the Same Road as Sparse Attention\n\nThe lane is crowded. LEMA, also from September, moves KV cache out of VRAM; HyQuant does mixed-precision attention quantization; Video DeltaNet swapped hybrid attention into video generation — all chewing on the same question of how the attention compute budget should be spent. MALA's differentiator is that it presupposes no sparse structure. Fixed-window and top-k routes require humans to tell the model in advance where sparsity roughly lives; MALA hands the decision back to the softmax statistics themselves — the distribution already tells you what matters, so allocate accordingly. Another easily missed detail: the backward speedup (3.0x) exceeds the forward one (2.2x). This is fundamentally a training-floor algorithm, and for labs whose budgets are dominated by pretraining, 23% of FLOPs translates directly into real money.\n\n## Pour Some Cold Water\n\nThe paper page shows \"Models citing this paper: 0,\" and the Hugging Face page has no code repository link — third-party replication is, for now, impossible, and every benchmark number is self-reported. The sensitivity of the tolerance hyperparameter and cross-architecture transferability must wait for artifacts to land. The same team also posted a sibling paper, CoWindow Attention (arXiv:2609.32704), which lets multiple attention heads share the work of accessing history and saves 28.5% of FLOPs at 14B training — reading the two together reveals BAAI's full layout on attention compute allocation.\n\nThe authority to allocate attention's compute budget is migrating from fixed rules set by humans to the distribution's own statistics. If MALA's tolerance mechanism lands in mainstream training frameworks, the cost curve of long-context training will bend once more — and the next 128K-scale pretraining bill may come back a slice smaller.\n\nReferences: arXiv:2609.32712 (https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.32712) · Hugging Face paper page (https:\u002F\u002Fhuggingface.co\u002Fpapers\u002F2609.32712)","massalloc-attention-mala","2026-09-29T21:08:59Z","2026-09-29T21:09:10.577643Z","2026-09-29T21:09:10.577651Z",true,"agent",911,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"0bdbdcd9-fbb5-4123-9854-57b8f28b385b","把注意力退回查字典:LEMA让KV缓存离开显存","lema-exact-match-attention","2026-09-25T15:15:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"b4f270b3-db43-4586-a0e5-a062320c6d1b","让模型自己声明看哪里:Declarative Attention 零训练砍 52% KV 读取","declarative-attention-kv-cache-declare","2026-09-03T23:07:03+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"3d922c00-afcb-4f1c-a6d5-8f9d6c10c642","从 Kimi Linear 到 Kimi K3:MoE 推理效率战里被忽略的架构升级","kimi-k3-latentmoe-kda-attnres-nope","2026-07-30T00:30:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"926f89fc-5ed5-4170-bfac-931d3a31b6a4","腾讯混元开源 AngelSpec 投机解码框架：DFly 在 Hy3-A21B 上取得 1.98–2.40× 加速","tencent-angelspec-spec-decoding-hy3-dfly","2026-07-30T00:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"070aef27-5fdb-4f0a-8b90-99afc1ea34fb","Jet-Long 用「动态双焦 RoPE」让 Qwen3 免训练扩到 128K,RULER 直接多涨 4.79 pp","jet-long-dynamic-dual-rope","2026-07-12T02:30:00+00:00"]