[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-sunta-surprise-chunking-video":3,"news-related-a2e8ac5b-ca51-4ddb-88d4-54373d1f0774":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a2e8ac5b-ca51-4ddb-88d4-54373d1f0774","SUNTA 用\"惊奇度\"切分视频预测:东京大学让模型在 250 步后仍不崩溃","长时序视频预测是世界模型绕不开的硬骨头。东京大学 Matsuo 团队在 arXiv 公开的 SUNTA,从最容易被忽略的角度切入:分层状态空间模型(HSSM)的分段边界到底该由谁决定。\n\n过去的 HSSM 用固定长度切片,或用帧间相似度找切换点,但这些启发式规则常常和数据本身的时序结构错位。SUNTA 提出用\"惊奇度\"(surprise-based chunking)驱动分段——预测误差突然升高的时刻,就是该调用更长程上下文的时刻。这一改看似朴素,实则解决了两个老问题:端到端训练时容易出现的\"层次塌缩\",以及开环预测时惊奇信号消失的尴尬。\n\n具体做法双管齐下:用解耦训练策略保留惊奇信号;在想象展开的预测里,用模型内部的\"不一致性\"作为顶层惊奇指标决定何时换段。效果立竿见影——2D\u002F3D 视频预测任务上,SUNTA 是唯一能在 250 步之后仍保持准确预测的方法,所有 baseline 在前 10 步就开始退化。\n\n这条思路对今天拼长视频一致性的世界模型(Sora、Veo、可灵等)是直接的\"技术借鉴清单\":分层抽象不该再交给人工设计的窗口,要让模型自己学着\"惊讶\"。当 AI 真正学会在被意外打断时换档,长视频才有可能从 5 秒连贯走向 5 分钟连贯。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.02087","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"ac5b09b0-95e9-4735-8c4f-816440dd0a25","en","SUNTA splits video prediction by surprise, stable at 250 steps","Long-horizon video prediction is the hard nut that world models can't avoid. Matsuo's team at the University of Tokyo, in SUNTA publicly posted to arXiv, cuts in from the most easily overlooked angle: who exactly decides the segmentation boundary of the Hierarchical State Space Model (HSSM). Past HSSM used fixed-length slicing, or used inter-frame similarity to find transition points, but these heuristic rules often misalign with the data's own temporal structure. SUNTA proposes using \"surprise-based chunking\" to drive segmentation — the moment the prediction error suddenly rises, is when a longer-range context should be invoked. This change may seem plain, but it actually solves two old problems: \"hierarchy collapse\" that easily occurs during end-to-end training, and the embarrassment of the surprise signal disappearing during open-loop prediction. The specific approach is a two-pronged: use a decoupled training strategy to preserve the surprise signal; in the prediction expanded by imagination, use the model's internal \"inconsistency\" as the top-level surprise indicator to decide when to switch segments. The effect is immediate — on 2D\u002F3D video prediction tasks, SUNTA is the only method that can still maintain accurate prediction after 250 steps, with all baselines starting to degrade in the first 10 steps. This line of thinking is a direct \"technology learning list\" for today's world models competing on long-video consistency (Sora, Veo, Kling, etc.): hierarchical abstraction should no longer be handed to human-designed windows, the model should learn to \"be surprised\" on its own. When AI truly learns to switch gears when interrupted by surprise, can long video move from 5 seconds of coherence to 5 minutes of coherence.","sunta-surprise-chunking-video","2026-07-04T16:00:00Z","2026-07-04T16:08:38.681993Z","2026-08-19T02:08:40.142862Z",true,"agent",95,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"82af5716-322e-46cf-9de5-b85e8cdd5712","微软首推安全专用模型 MAI-Cyber-1-Flash:小模型+多智能体,把漏洞挖掘成本砍半","microsoft-mai-cyber-1-flash-mdash","2026-07-28T01:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"2657cbe0-7743-43f2-9332-ee18b84b1229","Directing the World: 中国电信 TeleAI 把自回归视频世界模型推到\"组合控制\"","teleai-directing-the-world","2026-07-01T10:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"e4776508-8e3b-4eba-a804-ff4ee7e8a76d","「Holo-World」用一张图控制相机、物体和天气：视频世界模型首次把\"环境状态\"做成独立控制轴","holo-world-camera-object-weather-control","2026-06-21T16:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"123a3181-4a6b-499c-9575-a54aa6804620","蔚来把世界模型推到 NT2\u002FNT3 双平台：自研 AI Infra 让 4000 万公里「影子测试」装进每一辆车","nio-nt2-nt3-world-model-40m-km-shadow","2026-06-18T20:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"99319884-dac1-4ce8-81dd-8b5cb97bd91a","MoVerse 实时视频世界模型：用「全景高斯脚手架」把单图漫游跑进 8 FPS，扩散-3D-渲染三段式终于打通","moverse-8-fps-panoramic-gaussian-scaffold","2026-06-13T06:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"ff7fa3e1-6737-4bd9-85c8-8010d13a44f3","字节跳动 Bernini 开源：用 MLLM 当\"语义规划师\"，拆开视频生成的\"思考\"与\"渲染\"","bytedance-bernini-mllm-semantic-planner-diT","2026-05-21T00:00:00+00:00"]