[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-recurrent-looped-transformer-length-generalization":3,"topics-all":38,"news-related-d6ec8624-ad4e-41ce-9566-22d1f0926d49":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"d6ec8624-ad4e-41ce-9566-22d1f0926d49","循环解码器+并行编码器:RLT长度外推翻盘","普林斯顿等机构在 arXiv 发布 Recurrent Looped Transformer：网络层拆成并行编码器与循环解码器，计算路径随序列增长而单 token 成本固定。仅在 40 位奇偶序列上训练，256 位测试全种子 100% 准确，同规模 Transformer 停在 50%；消融显示增益完全依赖反馈通路。","Transformer 有个被长期回避的结构性短板：每个 token 分到的计算深度是固定的，八层就是八层，序列再长也不会多算一步。但状态跟踪任务恰恰要求每来一个输入就更新一次状态。普林斯顿大学的 Yifan Zhang 与宾夕法尼亚大学的两位合作者 10 月 6 日在 arXiv 发布 Recurrent Looped Transformer（RLT），把「并行」和「循环」这对老对手缝合进了同一个架构（[arXiv:2610.07591](https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.07591)）。\n\n## 八层拆两半：并行编码器加循环解码器\n\nRLT 的做法直白：把八层拆成并行因果编码器和循环解码器。编码器照常并行处理所有 token，产出表征和全局 KV 记忆；解码器在每个 token 处把编码器输出与上一个 token 的最终解码状态做门控合并，再配上每层的滑动窗口注意力缓存。效果是计算路径随序列长度增长——处理完 t 个 token，循环路径已经穿过 t 乘以解码器深度个解码器块——但每个 token 的计算成本保持固定。GitHub 仓库上线三天已获 918 星，代码以 Apache 2.0 开源。\n\n## 40 位训练，256 位全种子 100%\n\n长度外推的数字最能说明问题。只在至多 40 位奇偶校验上训练，5+3 和 7+1 两种切分在 256 位测试上所有种子都拿到 100% 准确率，同规模八层 Transformer 停在 50.07%。置换群 S5 追踪任务外推到训练长度八倍（256 次操作）时，4+4 切分达 97.30%，Transformer 只有 0.85%。模算术任务 RLT 最高 93%，Transformer 为 33%。16 层系列同样如此：8+8、9+7、11+5、16+0 四种切分在 256 位全部保持 100%，Transformer 16 是 49.41%。\n\n## 命门是反馈通路，不是层数\n\n消融把因果链钉死了：RLT-0 直接砍掉反馈通路，奇偶和 S5 全部跌回随机水平（64 位上 50.23%），每个切分都一样——增益完全依赖循环状态本身。作者还给出工程折中方案 RLT-2：每四个 token 才更新一次反馈状态，已知 token 可以在块内并行，CPU 训练步提速 2.27 倍，64 位奇偶保住约 99%；但 S5 追踪从 100% 掉到约 20%，说明置换跟踪必须逐 token 反馈。反馈间隔由此变成可调旋钮：预训练用大块抢并行度，后训练再把间隔收到 1。\n\n## 训练推理统一，社区已开始复现\n\n仓库给出 prefill、生成、预训练、SFT、RL replay 五种模式的统一执行表，全梯度穿过循环输出、解码器 KV 和编码器记忆。社区开发者已用约 7.9 万参数的独立小实现复现方向性结论：训练长度内全部拟合，但 128 次操作外推衰减到 60.8%，提醒我们这些结论目前限于算法任务、监督训练，强化学习表现未评估。\n\n这篇工作的定位值得划重点：它不是又一篇刷榜论文，而是把「深度换状态」变成了显式的设计轴。做长上下文、agent 记忆维护的人，值得把 108 组对照实验的原文完整读一遍。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.07591","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"91b57279-5876-4ded-b632-2580af023c74","en","Recurrent Looped Transformer: Depth That Grows With Sequence","Princeton's RLT splits layers between a parallel encoder and a recurrent decoder, hitting 100% parity at 256 bits where a Transformer stays at 50%.","Transformers have a structural blind spot: the compute depth each token receives is fixed—eight layers is eight layers, no matter how long the sequence grows. State tracking, however, demands an update at every input. The Recurrent Looped Transformer (RLT) from Yifan Zhang at Princeton and two University of Pennsylvania collaborators, posted to arXiv on October 6, stitches parallel and recurrent computation into one architecture ([arXiv:2610.07591](https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.07591)).\n\n## Eight layers, split in half\n\nRLT splits its layers between a parallel causal encoder and a recurrent decoder. The encoder processes tokens in parallel and produces representations plus global KV memory; at each token the decoder merges the encoder output with the previous token's final decoder state through a gated merge, backed by per-layer sliding-window attention caches. The computation path grows with sequence length—after t tokens the recurrent path has traversed t times the decoder depth in blocks—while per-token cost stays fixed. The GitHub repo picked up 918 stars within three days, Apache 2.0 licensed.\n\n## Trained on 40 bits, 100% at 256\n\nLength generalization tells the story. Trained only on parity of at most 40 bits, the 5+3 and 7+1 splits hit 100% accuracy at 256 bits in every seed, while a same-size eight-layer Transformer sits at 50.07%. On swap-based S5 permutation tracking at eight times the training length (256 operations), the 4+4 split reaches 97.30% versus 0.85%. Modular arithmetic tops out at 93% against 33%. The sixteen-layer series repeats the pattern: 8+8, 9+7, 11+5 and 16+0 all hold 100% at 256 bits; Transformer 16 manages 49.41%.\n\n## The feedback path is the whole game\n\nAblations nail the causal chain: RLT-0, which removes the feedback path, drops parity and S5 back to chance (50.23% at 64 bits) at every split. RLT-2 offers an engineering compromise: update the feedback state once per four-token chunk so known tokens run in parallel—CPU training steps speed up 2.27x and 64-bit parity stays at about 99%, but S5 tracking falls from 100% to about 20%. Permutation tracking genuinely needs per-token feedback. The feedback interval thus becomes a tunable knob: large chunks for pretraining parallelism, then squeeze toward 1 in mid- and post-training.\n\n## One execution across training and inference\n\nThe repo ships a unified schedule for prefill, generation, pretraining, SFT, and RL replay, with full gradients flowing through recurrent outputs, decoder KV, and encoder memory. A community developer has already reproduced directional results with an independent implementation of roughly 79K parameters: everything fits at training length, but accuracy decays to 60.8% at 128 operations—a reminder that these conclusions cover algorithmic tasks under supervised training, and RL is not evaluated.\n\nThe takeaway: this is not another benchmark-chasing paper. It turns depth-versus-state into an explicit design axis, and anyone working on long context or agent memory should read the 108-run comparison in full.","recurrent-looped-transformer-length-generalization","2026-10-08T23:30:00Z","2026-10-08T23:09:08.688836Z","2026-10-08T23:09:08.688850Z",true,"agent",153,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"530e5aa5-2026-4c30-b4e6-421caca907b2","Transformer 提前罢工:13 个基座模型跟不住引用链,一个 rank-8 LoRA 修好","tiny-lora-frozen-transformer-chain-relay","2026-10-03T21:05:16+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"926f89fc-5ed5-4170-bfac-931d3a31b6a4","腾讯混元开源 AngelSpec 投机解码框架：DFly 在 Hy3-A21B 上取得 1.98–2.40× 加速","tencent-angelspec-spec-decoding-hy3-dfly","2026-07-30T00:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"c8ba186e-e7a6-40e1-9485-41aeb4de388e","Sebastian Raschka 发布 LLM 架构图谱：40+ 开源模型一站式横向对比","sebastian-raschka-llm-arch-gallery-40","2026-05-19T07:01:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"95e663a4-2cdf-454c-9787-b154fbd41909","TokenRouter:token 级路由提速 64 倍","tokenrouter-token-level-llm-routing","2026-10-09T21:11:28+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"e40de2f9-9ee4-47f3-b1f1-7bf93a4870a9","一个动词翻转工具调用决策,LLM 内部向量现形","llm-tool-call-decision-vector","2026-10-09T13:10:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"f1ccbea4-5749-4976-bfbd-9fc835318226","Foil 重构循环 MoE:专家压进单层,循环次数翻八倍","foil-loop-moe-flatten-untie","2026-10-07T19:09:13+00:00"]