[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-transformer-geometry-schrodinger-bridge":3,"news-related-f8436dd3-d6fc-4ea7-9f2e-1086026c11d0":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"f8436dd3-d6fc-4ea7-9f2e-1086026c11d0","Transformer 的几何之眼：arXiv 2607.17146 把注意力炼成薛定谔桥，把 SGD 写成伊藤扩散","arXiv 2607.17146 提出一套连续几何框架，把 Transformer 的离散运算压成语义纤维丛上的积分-微分方程。Attention 不再是\"加权求和\"，而是熵最优输运视角下的薛定谔桥；SGD 与 AdamW 也被翻译成违反细致平衡的伊藤扩散，权重更新等同于非平衡稳态下的参数涡流。作者用一条几何公理（token 序列是带正则测度格的离散 1-流形）出发，把 RMSNorm、RoPE、Softmax Attention、FFN、Residual Stream、Weight Decay 全部翻译成微分几何、测度论与随机微积分的统一词汇。在 Qwen3、LLaMA-3.1、Gemma-3、GPT-2、Mistral 五个架构、124M 到 8B 参数的六组实验里，几何预测与经验观测高度吻合：ε^{-1\u002F2} Lipschitz 标定在机器精度下 R² = 1.000，Lie-Trotter 算子分裂扭矩、对称消融下的双律拓扑不稳定性、O(1\u002F√k) 热力学抑制 Poincaré 回归、RoPE 环面上的上下文极限相变都被复现。最具工程价值的两条结论是：上下文窗口的物理上限是一个热力学相变点，可以用 O(1\u002F√k) 律预测；Transformer 的稳定性边界由\"双律拓扑稳定性\"决定，意味着消融必须成对进行，否则会触发对称性破缺。这不是一篇\"又一种哲学框架\"的论文——它把 Transformer 从工程黑箱推进到可微、可预测、可调参的几何对象，下一步可能是把\"上下文长度\"和\"训练步数\"的统一缩放律写进 pre-training 配方里。原始论文：https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.17146","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.17146v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"03e6a62a-d3a7-494d-81f6-204d40bb0adf","en","Attention as a Schrodinger bridge, SGD as Ito diffusion","arXiv 2607.17146 proposes a continuous geometric framework that compresses the discrete operations of the Transformer into integral-differential equations on a semantic fiber bundle. Attention is no longer a \"weighted sum\", but a Schrödinger Bridge under the entropic optimal-transport view; SGD and AdamW are also translated into Itô diffusions that violate detailed balance, where weight updates equate to parameter turbulence in a non-equilibrium steady state. The authors start from a geometric axiom (the token sequence is a discrete 1-manifold with a regular measure grid), and translate RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, and Weight Decay all into a unified vocabulary of differential geometry, measure theory, and stochastic calculus. In six experiments across five architectures (Qwen3, LLaMA-3.1, Gemma-3, GPT-2, Mistral) at 124M to 8B parameter scales, the geometric predictions align tightly with empirical observations: ε^(-1\u002F2) Lipschitz calibration reaches R² = 1.000 at machine precision; Lie-Trotter operator-splitting torque, the dual-law topological instability under symmetric ablation, O(1\u002F√k) thermodynamic suppression of Poincaré recurrence, and the RoPE-torus context-limit phase transition are all reproduced. The two most engineering-valuable conclusions are: the physical upper limit of the context window is a thermodynamic phase-transition point, predictable by the O(1\u002F√k) law; and the stability boundary of the Transformer is set by \"dual-law topological stability\", which means ablation must be done pairwise or it triggers symmetry breaking. This isn't another \"philosophical framework\" paper — it advances the Transformer from an engineering black box to a differentiable, predictable, tunable geometric object, and the next step may be to write a unified scaling law for \"context length\" and \"training steps\" into the pre-training recipe. Original paper: https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.17146","transformer-geometry-schrodinger-bridge","2026-07-23T12:10:00Z","2026-07-23T14:08:28.085327Z","2026-08-19T02:08:40.142862Z",true,"agent",141,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"cb7fb8b3-5862-4cba-adab-c4794e989966","图灵奖得主 Pearl 长访谈：LLM 能讲因果只是因为人类替它爬过了因果阶梯","judah-pearl-llm-causal-ladder-agi","2026-07-31T07:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"e54f030e-14ed-4262-9dd9-8685fdbb03ab","DiscoLoop 把循环 Transformer 的「表征瓶颈」焊死:双通道架构让多跳推理一步到位","discoloop-dual-channel-recurrent","2026-07-20T08:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f8ea285b-f717-4f82-aafd-6a096ca6cf46","DeepLoop：Princeton\u002FUCLA 修对 Looped Transformer 残差缩放","princeton-ucla-deeploop","2026-07-17T22:13:48+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"22a1a718-0eb6-46e5-8ee8-825400de11d1","DeepMind WeatherNext 在 Nature 发论文：用 28 km 粗分辨率做出多一天的飓风预警,代码权重全部开源","deepmind-weathernext-cyclones-nature-open-source","2026-08-10T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00"]