[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-hdr-video-multi-step-planning":3,"news-related-79ed2e02-2fe4-43ca-a9b2-847740969424":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"79ed2e02-2fe4-43ca-a9b2-847740969424","HDR 把视频模型的多步推理硬拉出新手感:层级隐变量让经典规划任务成功率从 34% 跳到 60%","arXiv:2607.15278(7 月 16 日挂的)给出视频扩散模型做「多步视觉推理」的一个相当硬核的解法:HDR(Hierarchical Denoising for Visual Reasoning)。\n\n核心观察很直白——流式自回归 diffusion 跑得快但不会做长程规划,双向 diffusion 会做规划但每帧全重建代价太高,两边都卡死在逻辑一致性上。HDR 的 trick 是把视频隐变量搭成树状层级:粗粒度层先保留若干假设用于全局规划,细粒度层再把这些假设逐级具象化成具体视觉状态,中间用 SHAP(Sparse Hierarchical Attention Pattern)把时序注意力成本压下来。\n\n数字也很硬:6 个 OOD 任务(maze、Hanoi、一笔画、滑动拼图、Sokoban、倒水)的平均成功率从基线的 34.22 拉到 60.29(相对 +76.2%),平均进度从 76.00 拉到 89.56;延迟稳定在 0.70 秒\u002Flatent,比双向 diffusion 快 54.2 倍;只用 2% 训练数据仍能保留 82.9% 的全数据性能,而双向 diffusion 只剩 52.0%。作者还把模型搬到真机上做机器人实验,把它推向物理交互和世界建模。\n\n值得讨论的是这件事的战略意义。视频生成模型过去两年一直在卷「画面能不能再真一点」,但要进入「视觉基础模型」这一层级,真正的护城河是长程、可纠错、可规划的推理——HDR 这种「先在隐空间里把思路想清楚,再一帧一帧画出来」的范式,大概率会成为下一波视频推理工作的标准动作;而 2% 数据保留 82.9% 性能这件事,更是把「视觉推理需要海量演示」的老假设悄悄掀翻了一页。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.15278","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"2e3c13a1-1da2-4c1a-aa24-5e6f6bb539d7","en","HDR lifts video planning success from 34% to 60%","arXiv:2607.15278 (posted July 16) gives a fairly hardcore solution for video diffusion models doing \"multi-step visual reasoning\": HDR (Hierarchical Denoising for Visual Reasoning). The core observation is direct — streaming autoregressive diffusion is fast but can't do long-horizon planning, bidirectional diffusion can plan but each full frame reconstruction is too costly, both get stuck on logical consistency. HDR's trick is to layer the video latents into a tree hierarchy: coarse-grained layers first hold several hypotheses for global planning, fine-grained layers then concretize these hypotheses step by step into specific visual states, with SHAP (Sparse Hierarchical Attention Pattern) compressing temporal attention cost in between. The numbers are tough: average success rate on 6 OOD tasks (maze, Hanoi, one-stroke drawing, sliding puzzle, Sokoban, water-pouring) goes from the baseline's 34.22 to 60.29 (relative +76.2%), average progress from 76.00 to 89.56; latency steady at 0.70 s\u002Flatent, 54.2× faster than bidirectional diffusion; with only 2% of the training data, 82.9% of the full-data performance is preserved, while bidirectional diffusion retains only 52.0%. The authors also moved the model to a real robot for embodied experiments, pushing it toward physical interaction and world modeling. The strategic significance of this event is worth discussing. Video generation models have spent the past two years competing on \"can the picture be more real\", but to enter the \"visual foundation model\" tier, the real moat is long-horizon, corrigible, plannable reasoning — HDR's \"think through the plan in the latent space first, then draw it out frame by frame\" paradigm will most likely become the standard move of the next wave of video-reasoning work; and \"2% data preserves 82.9% performance\" quietly overturns the old assumption that \"visual reasoning needs massive amounts of demonstration\".","hdr-video-multi-step-planning","2026-07-18T12:00:00Z","2026-07-18T12:07:22.339281Z","2026-08-19T02:08:40.142862Z",true,"agent",92,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"028e11f1-c29d-47dd-9c66-d4d90bcc4a26","MMOE 之外:AIGC 团队重新算账,单卡 8×H100 也能跑赢参数堆叠","mmoe-diffusion-transformer-efficient-experts-reproducibility-budget","2026-08-02T08:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"9f443e5f-4ca1-4dfa-b8c7-a5d7ba7aaf6e","SANA-Video 2.0：用混合线性注意力把视频生成推到单卡可用","sana-video-2-mixed-linear-attention","2026-07-24T04:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f884bd95-631b-4904-8e8d-a8db751f9314","MV-Forcing：用「4D 几何桥」打通「长 × 多视角」,单一扩散模型端到端跑出 4D 视频","mv-forcing-4d-bridge","2026-07-08T04:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"36e9e98a-7d93-4e87-ad9e-a72af72a2c1c","ICML 2026 荣誉提名 Motive:首个『运动归因』框架,让视频生成学会挑运动片段","icml-2026-motion-attribution-motive","2026-07-07T08:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"0e44f256-e66e-495c-82e3-aae4dd5e2374","LiveEdit 把扩散视频编辑推到 12.66 FPS：清华让 AR 实时编辑走出 PPT","liveedit-ar-video-editing","2026-07-01T06:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"62b2e6d9-4ac7-457c-a30c-548c713730ad","PhyCo：让视频生成模型「理解」物理世界","phyco-cvpr-2026-physics-video-controlnet-vlm-reward","2026-05-01T16:00:00+00:00"]