[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mv-forcing-4d-bridge":3,"news-related-f884bd95-631b-4904-8e8d-a8db751f9314":33},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":20,"news_slug":26,"published_at":27,"created_at":28,"modified_at":29,"is_published":30,"publish_type":31,"image_url":13,"view_count":32},"f884bd95-631b-4904-8e8d-a8db751f9314","MV-Forcing：用「4D 几何桥」打通「长 × 多视角」,单一扩散模型端到端跑出 4D 视频","视频扩散两条赛道一直咬不上:时间自回归把单视角拉到分钟级(Sora、Veo 3.1);双向注意力做多视角一致(VideoMV、4Diffusion),却只能撑几秒静态。Cornell Tech 的 Fiebelman 等人在 **arXiv:2607.05376** 提出 MV-Forcing:以自回归 3D 重建作「4D 几何桥」传递视角间先验,让单一扩散模型同时吃下「长」与「多视角一致」。  机制分三层。 几何桥负责对齐——源视角 3D 重建后渲染下一视角的深度、法线、位姿先验,交给扩散做高频细节,3D 守一致、扩散守保真。联合去噪让两视角槽位都从噪声起互给先验,绕开 teacher 固定窗口,使生成真正无界。DMD + Spatio-Temporal Self-Forcing 把 few-step student 蒸馏出来,用视频级损失修补 exposure bias,延续 Xun Huang 团队 Self Forcing 思路,这次同时盖住时间与视角两个自回归轴。  为什么值得看。 世界模型与自动驾驶仿真要的正是「任意长度、任意视点、物理一致」的 4D 场景,此前只能堆算力或拼多通道管线。MV-Forcing 提出「几何归 3D,纹理归扩散」的轻量范式,若被更大模型验证,工业级 4D 视频生成的成本曲线有望再下一档。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.05376","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[21],{"id":22,"lang":23,"title":24,"summary":25,"content":13},"ad3b5fd8-e0cd-4914-a686-53702588be9e","en","MV-Forcing: one diffusion model end-to-end for 4D video","Two tracks in video diffusion have never quite connected: temporal autoregression pulls single-view to minute-level (Sora, Veo 3.1); bidirectional attention does multi-view consistency (VideoMV, 4Diffusion), but only stretches a few seconds of static content. Fiebelman et al. of Cornell Tech in **arXiv:2607.05376** propose MV-Forcing: using autoregressive 3D reconstruction as a \"4D geometric bridge\" to transfer inter-view priors, letting a single diffusion model simultaneously take in \"length\" and \"multi-view consistency\". The mechanism is in three layers. The geometric bridge handles alignment — after the source view's 3D reconstruction, the depth, normal, and pose priors of the next view are rendered and handed to the diffusion for high-frequency details; 3D guards consistency, diffusion guards fidelity. Joint denoising lets both view slots start from noise and provide each other with priors, side-stepping the teacher's fixed window and making generation truly unbounded. DMD + Spatio-Temporal Self-Forcing distills a few-step student, with video-level loss patching exposure bias, continuing the Self Forcing line of Xun Huang's team, this time covering both temporal and view autoregression axes at the same time. Why it's worth watching. What world models and autonomous-driving simulation want is exactly \"any length, any viewpoint, physical consistency\" 4D scenes, which previously could only be achieved by stacking compute or splicing multi-channel pipelines. MV-Forcing proposes a lightweight paradigm of \"geometry to 3D, texture to diffusion\" — if validated by larger models, the cost curve of industrial-grade 4D video generation is expected to drop another notch.","mv-forcing-4d-bridge","2026-07-08T04:00:00Z","2026-07-08T04:08:17.182511Z","2026-08-19T02:08:40.142862Z",true,"agent",93,{"items":34},[35,40,45,50,55,60],{"id":36,"title":37,"news_slug":38,"published_at":39},"79ed2e02-2fe4-43ca-a9b2-847740969424","HDR 把视频模型的多步推理硬拉出新手感:层级隐变量让经典规划任务成功率从 34% 跳到 60%","hdr-video-multi-step-planning","2026-07-18T12:00:00+00:00",{"id":41,"title":42,"news_slug":43,"published_at":44},"36e9e98a-7d93-4e87-ad9e-a72af72a2c1c","ICML 2026 荣誉提名 Motive:首个『运动归因』框架,让视频生成学会挑运动片段","icml-2026-motion-attribution-motive","2026-07-07T08:30:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"0e44f256-e66e-495c-82e3-aae4dd5e2374","LiveEdit 把扩散视频编辑推到 12.66 FPS：清华让 AR 实时编辑走出 PPT","liveedit-ar-video-editing","2026-07-01T06:15:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"62b2e6d9-4ac7-457c-a30c-548c713730ad","PhyCo：让视频生成模型「理解」物理世界","phyco-cvpr-2026-physics-video-controlnet-vlm-reward","2026-05-01T16:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"f4c705fd-47c9-481a-807f-8001820070f8","InfinityEdit:三注意力轻量适配器,把视频编辑推进无界流时代","infinityedit-infinite-video-editing-adapter","2026-08-25T13:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"18d2aa73-7244-4b10-b611-46475e17327e","ForgeWM开源:一步去噪72FPS的可玩世界模型,8张卡复现全流程","forgewm-few-step-playable-world-model","2026-08-24T21:10:00+00:00"]