[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-subliminal-clocks-dlm-latent":3,"news-related-79c21822-4e80-43f7-afea-baa14af4ba0a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"79c21822-4e80-43f7-afea-baa14af4ba0a","Subliminal Clocks: 扩散语言模型里那块\"潜时钟\"被找到了","Sapienza \u002F EPFL 等机构的 Rulli 等人 7 月 2 日挂在 arXiv 的论文,把当下正在崛起的\"扩散语言模型(DLM)\"架构翻开了一道新的解释性切口。LLaDA、Dream 7B、Gemini Diffusion、Mercury 这类模型,虽然走的是 BERT 式把 [mask] 一格一格揭开的生成路径,显式层面看不到任何 timestep 输入,但作者证明:DLM 的残差流里确实编码着一条与\"去噪进度\"对应的低维子空间——通过线性 probe,可以在多层稳定读出;沿这条子空间\"推动\"模型,会让输出的置信度和熵发生可预测变化。\n\n关键的图景是几何化的:LLaDA 把这条潜时间信号组织成一条从\"全部 [mask]\"到\"完全无 [mask]\"的低维流形曲线,而不是散乱分布。这说明 DLM 在内部已经自发学会计时——只是没被命名,也没被接口暴露。对做调度、做加速、做安全的人来说,这条\"潜时钟\"等于一个全新的可操控旋钮:可以在不改权重的条件下,通过激活空间干预去改变 DLM 的\"去噪节奏\"和\"自我确信度\";而这一点在过去,通常只能黑箱调温度或采样步数。\n\nDLM 阵营从 LLaDA 到商用 Gemini Diffusion 都在快速发展,但可解释性几乎空白。这篇论文给出的不是又一份 benchmark,而是一把能插进 DLM 内部的\"探针\",对学界和工业界都值得一读。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.01774","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"5a766984-4265-4db0-8955-f0b53af2dfa8","en","Subliminal Clocks: the hidden timer inside diffusion LMs","The paper, posted to arXiv on July 2 by Rulli et al. of Sapienza \u002F EPFL and other institutions, opens up a new explanatory window into the currently-rising \"Diffusion Language Model (DLM)\" architecture. LLaDA, Dream 7B, Gemini Diffusion, Mercury — these models, though they follow a BERT-style generation path of unmasking [mask] tokens one by one, with no explicit timestep input at the visible level, the authors prove: DLM's residual stream does encode a low-dimensional subspace corresponding to \"denoising progress\" — through linear probes, it can be stably read across multiple layers; \"pushing\" the model along this subspace causes predictable changes in the output's confidence and entropy. The key picture is geometric: LLaDA organizes this latent time signal into a low-dimensional manifold curve from \"all [mask]\" to \"completely no [mask]\", rather than scattered distribution. This shows that DLM has spontaneously learned to time itself internally — just hasn't been named, and hasn't been exposed via interface. For those doing scheduling, acceleration, or safety, this \"latent clock\" is a brand new manipulable knob: without changing weights, one can intervene in activation space to alter DLM's \"denoising rhythm\" and \"self-certainty\"; previously, this typically required black-box temperature tuning or sampling step adjustments. The DLM camp is developing rapidly from LLaDA to commercial Gemini Diffusion, but interpretability is almost blank. This paper gives not another benchmark, but a \"probe\" that can be inserted into DLM's interior, worth reading for both academia and industry.","subliminal-clocks-dlm-latent","2026-07-06T06:00:00Z","2026-07-06T06:12:15.255968Z","2026-08-19T02:08:40.142862Z",true,"agent",92,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"028e11f1-c29d-47dd-9c66-d4d90bcc4a26","MMOE 之外:AIGC 团队重新算账,单卡 8×H100 也能跑赢参数堆叠","mmoe-diffusion-transformer-efficient-experts-reproducibility-budget","2026-08-02T08:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"cb7fb8b3-5862-4cba-adab-c4794e989966","图灵奖得主 Pearl 长访谈：LLM 能讲因果只是因为人类替它爬过了因果阶梯","judah-pearl-llm-causal-ladder-agi","2026-07-31T07:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"3d922c00-afcb-4f1c-a6d5-8f9d6c10c642","从 Kimi Linear 到 Kimi K3:MoE 推理效率战里被忽略的架构升级","kimi-k3-latentmoe-kda-attnres-nope","2026-07-30T00:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f8436dd3-d6fc-4ea7-9f2e-1086026c11d0","Transformer 的几何之眼：arXiv 2607.17146 把注意力炼成薛定谔桥，把 SGD 写成伊藤扩散","transformer-geometry-schrodinger-bridge","2026-07-23T12:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b5638cab-a4d6-44ac-9230-32ed0a4cba9d","ARMT 把「记忆」焊进 Transformer:用恒定显存换无限上下文","armt-associative-recurrent-memory-transformer","2026-07-23T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"d2844cbb-b70e-469c-95b8-71cee8d735a6","给 Transformer 装上「CNN 鼻子」:用 0.01% 的参数量换 benchmark 普涨","transformer-cnn-nose-0-01-percent","2026-07-22T12:10:00+00:00"]