[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tiny-lora-frozen-transformer-chain-relay":3,"topics-all":41,"news-related-530e5aa5-2026-4c30-b4e6-421caca907b2":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"530e5aa5-2026-4c30-b4e6-421caca907b2","Transformer 提前罢工:13 个基座模型跟不住引用链,一个 rank-8 LoRA 修好","佐治亚理工实测 13 个基座模型:直接作答时只能可靠跟随 1.4-3.6 行引用链,预训练更多循环也没用。而在单个早期层挂一个 rank-8 LoRA、冻结其余全部权重后,Qwen3-8B 在 24 行链上精确准确率从 15.5% 升至 99%。LoRA 启动了中间层的接力传递——计算原本就在,只是没被激活。","先给一个会让人有点不舒服的结论:你每天在用的预训练大模型,可能连 `K = apple; B = K; D = B; print(D)` 这类几行长的引用链都跟不住。佐治亚理工的 Zehao Jin、Ruixuan Deng 和 Junran Wang 三人组在 9 月 29 日提交到 arXiv 的论文(2609.36585)里给出了测量:十三个 0.6B 到 32B 的基座模型,要求不写思维链、直接作答时,能可靠跟随的引用链平均只有 1.4 到 3.6 行;更扎心的是,给模型预训练额外的循环结构几乎不解决问题。模型的深度摆在那里,但它默认不会用。\n\n## 一个 LoRA,全模型冻结\n\n论文的修法克制到近乎吝啬:在单个早期层挂一个 rank-8 LoRA,其余全部权重冻结,只训这一个适配器。效果却是数量级式的——Qwen3-8B 在 24 行引用链上的精确准确率从 15.5% 拉到 99%,训练更久的版本能跟住 50 行;1.4B 的 Ouro 在四次循环后到 60 行,八次循环后至少 160 行。团队在 GitHub 仓库(Lunamos\u002Fstop-thinking-too-early,代码 Apache 2.0 开源)给出了完整复现脚本,headline 那个 LoRA(Qwen3-8B,第 14 层)在单张 A100 上训练加评估只要约 16 分钟——这个成本,任何做推理优化的团队都掏得起。\n\n## 接力机制:计算本来就在\n\n为什么一个单层适配器能撬动几十层的冻结网络?论文的分析给出了一个漂亮的答案:LoRA 启动的是一场「接力」。程序行把各自的链身份经由一小段中间层向下传递,冻结的注意力头沿着链条逐级向上读取;一旦移除对父行的注意力,接力立刻中断。这坐实了论文标题的判断——Transformer 不是不会算,是「停止思考得太早」:默认前向传播只用了深度的一小截,计算能力其实潜伏在那里,等人唤醒。团队还展示了一个在冻结模型上的测量方法,能在四个 held-out 模型中的三个里定位「最后一层有用的干预位置」;任务特定的 LoRA 在 MuSiQue 问答任务上也有提升,说明这套机制不只在合成任务上成立。\n\n## 对行业的三点意味\n\n第一,「模型不会」和「模型没被激活」是两回事。评测算子给出低分时,先别急着换更大的模型——可能是深度利用率问题,一个 16 分钟的 LoRA 就能翻盘。第二,推理成本叙事要修:如果少量参数编辑就能释放冻结网络里的既有计算,那「买更深模型」和「激活现有深度」之间,存在一条便宜的中间路线。第三,可解释性有了工程抓手:接力机制画出了一张清晰的层间分工图,后续做推理加速或结构搜索的人可以直接在上面做文章。\n\n所以呢:下次看到模型答错长链条问题,别急着骂容量不够——它可能只是没被叫醒。论文、代码与在线演示都在下面的链接里,16 分钟一张 A100,值得亲手跑一遍。\n\n参考:arxiv.org\u002Fabs\u002F2609.36585 · github.com\u002FLunamos\u002Fstop-thinking-too-early","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36585","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":25,"name":26,"slug":26,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"c2ebe1ba-5e37-4558-bdbe-88e085ff63b6","en","Tiny LoRA wakes frozen transformers: 15.5% to 99% on chains","Georgia Tech: base LLMs follow only 1.4-3.6 reference-chain lines. One rank-8 LoRA, all weights frozen, lifts Qwen3-8B to 99% on 24-line chains.","Start with an uncomfortable finding: the pretrained LLMs you use every day likely cannot follow even a short chain of references like `K = apple; B = K; D = B; print(D)` when asked to answer directly. A Georgia Tech team — Zehao Jin, Ruixuan Deng and Junran Wang — measured this in a paper submitted to arXiv on September 29 (2609.36585): thirteen base models from 0.6B to 32B parameters reliably follow only 1.4 to 3.6 lines of such chains when no chain-of-thought is allowed. More uncomfortably, adding extra pretrained loops barely helps. The depth is sitting right there; the model just does not use it by default.\n\n## One LoRA, everything else frozen\n\nThe fix is almost austere: attach a single rank-8 LoRA at one early layer, freeze every other weight, and train only that adapter. The effect is an order-of-magnitude jump — Qwen3-8B goes from 15.5% to 99% exact accuracy on 24-line chains, and a longer-trained LoRA reaches 50 lines. The 1.4B Ouro model reaches 60 lines after four loops and at least 160 after eight. The team open-sourced the full reproduction pipeline (Lunamos\u002Fstop-thinking-too-early, Apache 2.0); the headline LoRA (Qwen3-8B, layer 14) trains and evaluates in about 16 minutes on a single A100 — a price any inference-optimization team can afford.\n\n## The relay mechanism: the computation was already there\n\nWhy can a one-layer adapter mobilize dozens of frozen layers? The paper's analysis offers an elegant answer: the LoRA starts a relay. Program lines pass their chain identity down through a short range of middle layers, and frozen attention heads read progressively further up the chain; remove attention to the parent line and the relay stops immediately. This confirms the paper's title — transformers do not lack the ability to compute, they simply stop thinking too early: the default forward pass uses only a small slice of the available depth. The team also demonstrates a frozen-model measurement that locates the last useful intervention layer within tolerance in three of four held-out models, and task-specific LoRAs also improve results on MuSiQue — the mechanism is not confined to synthetic tasks.\n\n## Three takeaways for the industry\n\nFirst, \"the model cannot do it\" and \"the model was never activated\" are different failure modes. When an eval returns a low score, do not rush to buy a bigger model — it may be a depth-utilization problem that a 16-minute LoRA can flip. Second, the inference-cost narrative needs revision: if a tiny parameter edit releases computation already latent in a frozen network, there is a cheap middle road between \"buy more depth\" and \"activate what you have\". Third, interpretability gains an engineering handle: the relay mechanism maps a clean division of labor across layers, which anyone working on inference acceleration or architecture search can build on.\n\nSo the next time a model flubs a long-chain question, do not blame capacity too quickly — it may simply not have been woken up. Paper, code and an interactive demo are linked below; 16 minutes on one A100 is worth spending yourself.\n\nReferences: arxiv.org\u002Fabs\u002F2609.36585 · github.com\u002FLunamos\u002Fstop-thinking-too-early","tiny-lora-frozen-transformer-chain-relay","2026-10-03T21:05:16Z","2026-10-03T21:06:56.553906Z","2026-10-03T21:06:56.553917Z",true,"agent",1232,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"d6ec8624-ad4e-41ce-9566-22d1f0926d49","循环解码器+并行编码器:RLT长度外推翻盘","recurrent-looped-transformer-length-generalization","2026-10-08T23:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"926f89fc-5ed5-4170-bfac-931d3a31b6a4","腾讯混元开源 AngelSpec 投机解码框架：DFly 在 Hy3-A21B 上取得 1.98–2.40× 加速","tencent-angelspec-spec-decoding-hy3-dfly","2026-07-30T00:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"7312935f-7432-4dbd-9cff-c665fa7b4765","FlashMorph:ByteDance Seed 把混合注意力的\"层选择\"做成预算约束优化","flashmorph-hybrid-attention-layer","2026-07-05T02:01:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"a886afac-eb9a-4666-82a2-03fabc82a29f","Tapered Language Models：Mila\u002FCornell\u002FUdeM 用「锥形 MLP」把 LLM 的容量分配「免费升舱」","tapered-lm-tapered-mlp-free-upgrade","2026-06-28T14:15:00+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"c8ba186e-e7a6-40e1-9485-41aeb4de388e","Sebastian Raschka 发布 LLM 架构图谱：40+ 开源模型一站式横向对比","sebastian-raschka-llm-arch-gallery-40","2026-05-19T07:01:00+00:00"]