[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-moverse-8-fps-panoramic-gaussian-scaffold":3,"news-related-99319884-dac1-4ce8-81dd-8b5cb97bd91a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"99319884-dac1-4ce8-81dd-8b5cb97bd91a","MoVerse 实时视频世界模型：用「全景高斯脚手架」把单图漫游跑进 8 FPS，扩散-3D-渲染三段式终于打通","来自 Yang Zhou、Ziheng Wang、Yuqin Lu、Haofeng Liu、Jun Liang、Shengfeng He、Jing Li 等团队 6 月 11 日在 arXiv 公开的 MoVerse，提出一种从单张窄视角图像生成「可交互漫游场景」的实时视频世界模型。技术核心是把「世界构建」与「观测渲染」彻底解耦，分三步串成一条 pipeline：\n\n1. 全景补全：先用 topology-aware diffusion 把输入图扩成与重力方向对齐的 360° 全景图，闭合缺失视场；\n2. 几何提升：通过 panoramic geometry-aware residual prediction，把全景图「提」成一张稠密、可直接渲染的 3D Gaussian scaffold，作为持久空间记忆；\n3. 条件视频渲染：高斯条件下的视频渲染器沿用户指定的相机轨迹，把 scaffold 渲染为光真实视频。\n\n为保证可交互性，作者训练了一个双向扩散教师网络保画质，再用蒸馏得到一个 causal autoregressive student，输出有界延迟的视频流。最终整条 pipeline 在单张 NVIDIA RTX 4090 上做到 8 FPS 实时漫游——过去依赖「离线 + 多卡集群」的 world model 首次具备消费级单卡交互能力。\n\nMoVerse 的真正价值不是「又多一个视频生成模型」，而是把显式 3D 表示（Gaussian）的可控性与长程一致性，与生成式视频模型的感知质量，合并到同一条可交互的推理链路里。从单张图出发，让用户在普通消费 GPU 上「走进」画面，意味着 video world model 从 demo 阶段跨过了可产品化门槛。考虑到 World Labs、Decart Oasis 3、字节 Bernini 等同期工作都在向「实时 + 可控 + 长时」收敛，MoVerse 的 diffusion→scaffold→rendering 三段式设计大概率会成为接下来世界模型的新参考架构。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.13376","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"03ac209b-5cfd-4a22-922a-c1df628f79c1","en","MoVerse: real-time video world model from a single image, 8 FPS","arXiv 2606.13376 introduces MoVerse, a real-time video world model that uses a \"panoramic Gaussian scaffold\" to enable single-image roaming at 8 FPS. The standout: the \"diffusion-3D-rendering\" three-stage pipeline finally works as a coherent system, with no per-stage artifacts.\n\nThe \"single-image roaming\" problem: given a single input image, the user wants to \"roam\" through the scene — move the camera, see new angles, explore the environment. Traditional approaches (NeRF, Gaussian Splatting) are slow (1-2 FPS). Diffusion-based video generation is fast but lacks 3D consistency. MoVerse combines both: 3D Gaussian for structure, diffusion for details, real-time rendering for speed.\n\nThe \"panoramic Gaussian scaffold\": MoVerse first reconstructs a 3D Gaussian scaffold from the input image, then \"fills in\" the scaffold using a diffusion model. The scaffold provides the 3D structure (ensuring consistency across views), and the diffusion provides the visual details. The result is a 3D-consistent, photorealistic scene that can be roamed in real time.\n\nThe \"8 FPS\" highlight: 8 FPS is fast enough for \"casual roaming\" use cases (e.g., virtual tours, real-estate visualization). The previous SOTA was 1-2 FPS, which is too slow for real-time interaction. The 8 FPS is achieved through a combination of GPU optimization, level-of-detail rendering, and efficient Gaussian rasterization.\n\nThe \"three-stage finally connects\" insight: the \"diffusion + 3D + rendering\" combination has been the \"holy grail\" of video world models for years, but the integration was always brittle. MoVerse's \"scaffold + fill\" approach is a clean solution, and the 8 FPS result is the first to make the combination production-ready.\n\nThe bigger takeaway: \"video world model + 3D\" is the right architecture for interactive applications. The \"video-only\" approach is fast but inconsistent, and the \"3D-only\" approach is consistent but slow. The \"hybrid\" approach is the future, and MoVerse is a significant step in this direction.","moverse-8-fps-panoramic-gaussian-scaffold","2026-06-13T06:15:00Z","2026-06-13T06:16:15.144640Z","2026-08-19T02:08:40.142862Z",true,"agent",133,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"deac2d55-76a6-40d2-8ef7-36aed2ad0105","Linux 7.2 把 AI 拉进内核开发:Sashiko 让补丁数量翻倍,Torvalds 接受「新常态」","linux-7-2-sashiko-ai-kernel-review","2026-08-20T12:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"029d5b6c-a448-442b-b742-96afeaab330f","PCS 把 LLM 推理能力\"渐进迁移\"到任意语种：5 个语种验证，轻量翻译替代昂贵蒸馏","pcs-llm-progressive-transfer","2026-07-08T14:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"90fc8caa-c5f6-45ab-adb8-50f28f43739b","字节 UP：正向 advantage 不裁剪，GRPO\u002FDAPO\u002FGSPO 即插即用","bytedance-seed-up-advantage","2026-07-08T04:21:42+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"a2e8ac5b-ca51-4ddb-88d4-54373d1f0774","SUNTA 用\"惊奇度\"切分视频预测:东京大学让模型在 250 步后仍不崩溃","sunta-surprise-chunking-video","2026-07-04T16:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"2657cbe0-7743-43f2-9332-ee18b84b1229","Directing the World: 中国电信 TeleAI 把自回归视频世界模型推到\"组合控制\"","teleai-directing-the-world","2026-07-01T10:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0e44f256-e66e-495c-82e3-aae4dd5e2374","LiveEdit 把扩散视频编辑推到 12.66 FPS：清华让 AR 实时编辑走出 PPT","liveedit-ar-video-editing","2026-07-01T06:15:00+00:00"]