[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-discoloop-dual-channel-recurrent":3,"news-related-e54f030e-14ed-4262-9dd9-8685fdbb03ab":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e54f030e-14ed-4262-9dd9-8685fdbb03ab","DiscoLoop 把循环 Transformer 的「表征瓶颈」焊死:双通道架构让多跳推理一步到位","LLM 想在单次前向里完成多跳推理,几乎一定会撞上「深度局部存储」这堵墙——早期层学到的桥接实体,等到第二跳检索时已经找不回来。Looped Transformer 用循环复用参数缓解了内存问题,但泛化一直不干净。UC Berkeley Stuart Russell 团队 7 月 1 日挂出的 DiscoLoop(arXiv:2607.00341)把症结直接归到表征上:第一个循环其实已经把桥接实体解码得几乎完美,但对应的隐状态却和这个 token 的 embedding 对不齐。一个零训练成本的 realignment 干预,就能把泛化 gap 拉满。基于这一观察,作者提出双通道循环架构——同时跑一条离散 embedding 通道和一条连续隐状态通道。在符号化和合成语言多跳任务上,DiscoLoop 用远少于基线的训练步数拿到近 100% 准确率;迁移到真实预训练后,训练 loss 更低,基准也更强。最有意思的不是 SOTA,而是「一个免费 realignment 就能补上大部分 gap」这个发现——它说明 Looped Transformer 一直在输的不是参数,而是表征通道太单薄。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.00341","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"ee50ff0a-4341-47c3-9210-c32d0a8cac34","en","DiscoLoop welds shut the looped-Transformer bottleneck","If LLMs are to complete multi-hop reasoning in a single forward pass, they'll almost certainly hit the \"shallow local storage\" wall — bridging entities learned in early layers are lost by the time the second hop retrieves them. Looped Transformers, which re-use parameters to alleviate memory issues, haven't generalized cleanly. UC Berkeley's Stuart Russell team posted DiscoLoop (arXiv:2607.00341) on July 1, attributing the issue directly to representation: the first loop actually decodes the bridging entity nearly perfectly, but the corresponding hidden state is misaligned with this token's embedding. A zero-training-cost realignment intervention can close most of the generalization gap. Based on this observation, the authors propose a dual-channel loop architecture — running a discrete-embedding channel and a continuous hidden-state channel simultaneously. On symbolic and synthetic language multi-hop tasks, DiscoLoop reaches near-100% accuracy with far fewer training steps than the baseline; when transferred to a real pretrained model, training loss is lower and benchmarks are stronger. The most interesting point isn't the SOTA, but the discovery that \"a free realignment can fill most of the gap\" — it means Looped Transformers have been losing not from parameters but from a too-thin representation channel.","discoloop-dual-channel-recurrent","2026-07-20T08:00:00Z","2026-07-19T20:06:28.740379Z","2026-08-19T02:08:40.142862Z",true,"agent",129,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"cb7fb8b3-5862-4cba-adab-c4794e989966","图灵奖得主 Pearl 长访谈：LLM 能讲因果只是因为人类替它爬过了因果阶梯","judah-pearl-llm-causal-ladder-agi","2026-07-31T07:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"f8436dd3-d6fc-4ea7-9f2e-1086026c11d0","Transformer 的几何之眼：arXiv 2607.17146 把注意力炼成薛定谔桥，把 SGD 写成伊藤扩散","transformer-geometry-schrodinger-bridge","2026-07-23T12:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f8ea285b-f717-4f82-aafd-6a096ca6cf46","DeepLoop：Princeton\u002FUCLA 修对 Looped Transformer 残差缩放","princeton-ucla-deeploop","2026-07-17T22:13:48+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"22a1a718-0eb6-46e5-8ee8-825400de11d1","DeepMind WeatherNext 在 Nature 发论文：用 28 km 粗分辨率做出多一天的飓风预警,代码权重全部开源","deepmind-weathernext-cyclones-nature-open-source","2026-08-10T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00"]