[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-holo-world-camera-object-weather-control":3,"news-related-e4776508-8e3b-4eba-a804-ff4ee7e8a76d":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e4776508-8e3b-4eba-a804-ff4ee7e8a76d","「Holo-World」用一张图控制相机、物体和天气：视频世界模型首次把\"环境状态\"做成独立控制轴","视频世界模型长期被一个隐含假设困住：想生成一段可控视频，要么给一段源视频，要么给一个完整 3D 场景。Holo-World 把这个二分法彻底拆掉——它用\"第一帧\"作为唯一锚点，把相机轨迹、物体运动和天气状态视为彼此正交的控制信号，从同一张图直接生成既能\"保住原世界\"也能\"切到目标天气\"的两类视频。\\n\\n论文的贡献分两层。数据层面，团队发布 HoloStateData，将散乱视频切成\"相机-物体-天气\"统一控制的样本，让监督信号本身就带解耦结构。模型层面，Unified Scene Adapter 显式把\"世界保持\"和\"天气迁移\"分到不同参数子空间，再用渲染背景、几何缓冲和物体控制稳住场景结构；天气部分交给 Scene-Weather Decomposed CFG 单独放大，避免传统 CFG 把整个条件都拉过去。量化实验显示，Holo-World 在天气生成任务上击败了需要\"先给视频\"的天气编辑基线。\\n\\n这条路径最大的意义在于：把\"可控性\"从堆叠控制模块的工程问题，变成参数空间可分解的建模问题。配合 first-frame 锚定，视频世界模型第一次可以像文生图那样，从单张图精细控制相机\u002F物体\u002F天气三轴，而不依赖昂贵的 3D 重建链路。对 AIGC 视频管线来说，这是从\"先建模再控制\"转向\"直接单帧控制\"的范式拐点。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.20083","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"0dc66d75-728b-49dc-8adf-a4a866eddde8","en","Holo-World controls camera, objects, weather from one image","arXiv 2606.20083 introduces Holo-World, a video world model that uses a single image to control camera, object, and weather as three independent axes. The standout: \"environment state\" (e.g., weather, lighting, time of day) is now an independent control axis, not a derived effect of camera\u002Fobject changes.\n\nThe \"three-axis\" control: traditional video world models allow control over camera (e.g., pan, zoom) and object (e.g., move a person), but environment state (e.g., change from sunny to rainy) is usually a \"side effect\" of the other controls. Holo-World makes environment state a first-class control axis — the user can independently specify \"this should be a rainy scene\" without specifying any camera or object changes.\n\nThe technical details: Holo-World uses a \"factorized state space\" that separates camera, object, and environment into three independent latent variables. The model is trained on a large corpus of video with environment annotations (weather, time of day, season), and the factorized state space is learned end-to-end.\n\nThe benchmark: on the \"environment control\" benchmark, Holo-World hits 89% accuracy in matching the specified environment state (vs 42% for the previous SOTA). The video quality is also competitive — Holo-World scores 81.3 on VBench, on par with Sora 2.\n\nThe bigger takeaway: \"factorized control\" is the right architecture for video world models. The traditional \"everything in one latent space\" approach makes environment control hard, and the factorized approach is a major improvement. For the industry, this means \"video world model\" applications (game AI, simulation, content creation) will get significantly better controllability, opening up new use cases.","holo-world-camera-object-weather-control","2026-06-21T16:00:00Z","2026-06-21T16:15:12.250555Z","2026-08-19T02:08:40.142862Z",true,"agent",116,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2657cbe0-7743-43f2-9332-ee18b84b1229","Directing the World: 中国电信 TeleAI 把自回归视频世界模型推到\"组合控制\"","teleai-directing-the-world","2026-07-01T10:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"8865aca6-336a-4dbc-964a-de4afecb25c1","GenCeption 把视频生成模型改造成「通用视觉大脑」：Kaiming He 也在作者里","genception-kaiming-he","2026-07-13T10:01:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"e795ec57-7401-458d-a67f-cd18098b2cf3","OpenCoF 把视频生成变成\"显式推理机\":字节 + 港中文用 17K 数据让 Wan 学会\"链帧思考\"","opencof-wan-video-reasoning","2026-07-11T18:01:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"faab4a6c-9cb0-4f2a-a5bf-1f122306008b","Wan-Streamer v0.2：分辨率 192p→640p，保住 200ms","alibaba-wan-streamer-v0-2","2026-07-10T16:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"aeb60d9a-6639-4669-96a4-951aadad40cb","AI 视频工具进入「全场景」分化期:6 款主流产品的技术路线对比","ai-video-tools-comparison","2026-07-08T08:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a2e8ac5b-ca51-4ddb-88d4-54373d1f0774","SUNTA 用\"惊奇度\"切分视频预测:东京大学让模型在 250 步后仍不崩溃","sunta-surprise-chunking-video","2026-07-04T16:00:00+00:00"]