[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tapered-lm-tapered-mlp-free-upgrade":3,"topics-all":36,"news-related-a886afac-eb9a-4666-82a2-03fabc82a29f":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a886afac-eb9a-4666-82a2-03fabc82a29f","Tapered Language Models：Mila\u002FCornell\u002FUdeM 用「锥形 MLP」把 LLM 的容量分配「免费升舱」","过去十年，从原始 Transformer 到 Gated Attention、Hope-attention、Titans，所有现代语言模型几乎都共享一个默认骨架：把 L 层相同结构的 block 均匀堆叠，每层分到相同份额的参数。这个从原始 transformer 继承下来的设计几乎从未被质疑过。\n\nMila、Cornell、Université de Montréal 和 CIFAR AI Chair 的 Reza Bayat 等人在 arXiv:2606.23670 中把它打了个问号。他们做了个简单对照实验：在 440M 参数 transformer 上保持总参数不变，只把 MLP 中间宽度按三段（前\u002F中\u002F后）重新分配。结果非常不对称——「前宽后窄」的 perplexity 比均匀基线低 0.32，反向「前窄后宽」却高出一整个点以上。方向错了就是浪费预算，方向对了就是白送的收益。\n\n基于这个观察，他们提出 Tapered Language Models (TLMs)：在固定参数与 FLOPs 预算下，用平滑的 cosine schedule 把 MLP 宽度沿深度单调递减，把容量前置到浅层。设计空间有三种 schedule（linear \u002F cosine \u002F sigmoid），cosine 因为两端都有平台、过渡最平滑，结果最稳。\n\n核心结论很硬：TLM 在 Transformer、Gated Attention、Hope-attention、Titans 四种 token-mixing 截然不同的架构上，440M \u002F 760M \u002F 1.3B 三个规模都稳定降低 perplexity 并提升下游基准——没有任何额外参数或计算开销。机理也讲清楚了：作者测了每层 MLP 输出与残差的对齐度，发现越深的层越倾向于「重述」残差而不是「写入」新特征，把宽度用在这类冗余层上就是在浪费。\n\nlayer-skipping、ShortGPT、早期退出、模型剪枝这一系列研究已经反复暗示「深层 MLP 不那么重要」，TLMs 把这条线索从「能不能砍」推进到「按什么比例分配」。它不需要新算力、不需要新数据，是任何团队只要换个初始化 schedule 就能拿到的免费杠杆。下一步值得期待的是，这套「沿深度非均匀」思路扩散到 attention head 数、KV 维度、recurrent state 大小乃至 MoE 的专家数——这些维度也都存在「前后层贡献不均」的嫌疑。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.23670","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"73b36ced-1941-40ea-bb22-043656ce4667","en","Tapered LMs: tapered MLPs give free capacity upgrades","For the past decade, from the original Transformer to Gated Attention, Hope-attention, and Titans, almost all modern language models share a default skeleton: L layers of the same-structure block stacked uniformly, with each layer getting an equal share of parameters. This design inherited from the original transformer has hardly ever been questioned.\n\nReza Bayat et al. at Mila, Cornell, Université de Montréal, and the CIFAR AI Chair challenge it in arXiv:2606.23670. They ran a simple controlled experiment: on a 440M-parameter transformer, holding total parameters constant, they simply redistribute the MLP intermediate width across three segments (front \u002F middle \u002F back). The result is strikingly asymmetric — \"wide front, narrow back\" lowers perplexity by 0.32 versus the uniform baseline, while the reverse \"narrow front, wide back\" runs more than a full point higher. Get the direction wrong and you waste budget; get it right and the gain is free.\n\nBased on this observation, they propose Tapered Language Models (TLMs): under fixed parameter and FLOPs budgets, use a smooth cosine schedule to monotonically decrease the MLP width along depth, front-loading capacity to the shallow layers. The design space has three schedules (linear \u002F cosine \u002F sigmoid); cosine is the most stable because it has plateaus at both ends and the smoothest transition.\n\nThe core result is hard: TLMs stably lower perplexity and improve downstream benchmarks on four token-mixing-different architectures (Transformer, Gated Attention, Hope-attention, Titans) and three scales (440M \u002F 760M \u002F 1.3B) — with no extra parameters or compute. The mechanism is also explained: the authors measure the alignment between each layer's MLP output and the residual stream, and find that the deeper the layer, the more it tends to \"restate\" the residual rather than \"write\" new features — so spending width on such redundant layers is wasteful.\n\nLayer-skipping, ShortGPT, early exit, and model pruning have all repeatedly hinted that \"deep MLPs aren't that important.\" TLMs push that thread from \"can we cut\" to \"in what proportion should we allocate.\" It needs no new compute and no new data — a free lever any team can pick up by changing the initialization schedule. The next step worth watching is whether this \"non-uniform along depth\" thinking spreads to attention-head counts, KV dimensions, recurrent state sizes, and even the number of MoE experts — all of which also have a \"front and back layers contribute unevenly\" smell.","tapered-lm-tapered-mlp-free-upgrade","2026-06-28T14:15:00Z","2026-06-28T14:13:41.187654Z","2026-08-19T02:08:40.142862Z",true,"agent",201,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"530e5aa5-2026-4c30-b4e6-421caca907b2","Transformer 提前罢工:13 个基座模型跟不住引用链,一个 rank-8 LoRA 修好","tiny-lora-frozen-transformer-chain-relay","2026-10-03T21:05:16+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"d6ec8624-ad4e-41ce-9566-22d1f0926d49","循环解码器+并行编码器:RLT长度外推翻盘","recurrent-looped-transformer-length-generalization","2026-10-08T23:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"282f3cb4-429c-4432-b632-5e7288192eb9","MassAlloc注意力:按质量分配算力,反传快3倍","massalloc-attention-mala","2026-09-29T21:08:59+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"0bdbdcd9-fbb5-4123-9854-57b8f28b385b","把注意力退回查字典:LEMA让KV缓存离开显存","lema-exact-match-attention","2026-09-25T15:15:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"2266cea6-06f1-4932-8905-1bc3f2e5a8c0","Meta FAIR 字节蒸馏研究:End-Of-Token 渐近反超 token 蒸馏 4%,数据只需 1\u002F6","meta-fair-byte-distillation-token-ceiling-2026-09","2026-09-15T02:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"644ac630-d5e5-4ee6-9d63-eccc85d811c4","NCP-ArchPreview：一半 token 追平 OLMo-3 训练损失，概念级预测改写预训练经济学","ncp-archpreview-next-concept-prediction","2026-09-11T15:10:00+00:00"]