[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-dreamx-world-1-0-amap-controllable-camera":3,"news-related-1d5771ce-dbfa-4a66-8f20-efff9b7ba3b2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1d5771ce-dbfa-4a66-8f20-efff9b7ba3b2","DreamX-World 1.0：把通用世界模型拉回「可控相机 + 长程记忆」的真问题","高德地图 AMAP-ML 团队在 arXiv 发布 DreamX-World 1.0，提出一种支持可控相机导航和长程场景记忆的通用交互式世界模型。它用 E-PRoPE 投影式位置编码实现低成本相机控制，用 Memory-Conditioned Scene Persistence（MCSP）从相机几何回拉历史视角抑制长视频的颜色与风格漂移，并通过 DMD 蒸馏 + 因果强制训练 + 长 rollout 训练 + 强化学习对齐，把双向视频生成器压成少步自回归世界模型。8 张 RTX 5090 上跑到 16 FPS，5 秒评估整体分 84.76，超过 HY-WorldPlay 1.5（80.79）和 LingBot-World（80.45）。\\n\\n技术上有三个亮点：\\n\\n**E-PRoPE**——一种轻量化的投影式位置编码，把相机几何以 attention 注入到空间压缩后的 token 上，免去全分辨率相机控制的开销，同时保留 PRoPE 的射影几何性质。\\n\\n**Memory-Conditioned Scene Persistence（MCSP）**——用相机几何检索历史帧，把已生成过的视角拉回来当 conditioning；残差回收机制让 conditioning 路径对不完美的记忆 latent 更鲁棒，是抑制长视频累积漂移（颜色偏移、风格走样）的关键招。\\n\\n**DMD 蒸馏 + 因果强制训练 + 长 rollout 训练 + RL 对齐**——把双向视频生成器改成少步自回归世界模型：自生成的长程上下文让模型反复接触自己的历史，再用 RL 找回蒸馏丢掉的相机精度与画质。\\n\\n实测在 8 张 RTX 5090 上能跑到 16 FPS，5 秒评估的整体分 84.76，超过 HY-WorldPlay 1.5（80.79）和 LingBot-World（80.45）。配合混合精度 DiT、75% 剪枝的 VAE 解码、异步流水线并行，整套推理栈做了系统级优化。\\n\\n最值得说的还是思路：之前很多「世界模型」演示稿都把力气花在「逼真度」上，DreamX 团队却把工程重心放在「可控相机 + 长程记忆」这两件更接近实用门槛的事情上。这两条若真站稳，下游物理 AI 训练的合成环境、消费级交互创作工具才有底座可用。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.16993","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d8173b73-6a3a-432e-bb0a-0b0816b92562","en","DreamX-World 1.0: controllable camera, long-horizon memory","arXiv 2606.16993 introduces DreamX-World 1.0, a general-purpose world model that focuses on the \"controllable camera + long-range memory\" combination. The standout: unlike most world models that focus on \"video quality\" or \"physics accuracy,\" DreamX-World 1.0 focuses on \"controllability\" — the ability to specify the camera path and the memory of past events, and have the model follow them consistently.\n\nThe \"controllable camera\" highlight: DreamX-World 1.0 allows the user to specify a precise camera path (e.g., \"pan left 30 degrees, then zoom in 2x, then orbit around the object\"). The model generates video that strictly follows the camera path, with no drift. This is a significant improvement over previous world models, where the camera path was approximate at best.\n\nThe \"long-range memory\" highlight: DreamX-World 1.0 maintains a \"memory\" of past events in the generated video, allowing the user to query \"what happened 30 seconds ago\" and get a consistent answer. The memory is stored as a 3D scene representation, and the model uses it to maintain consistency across long videos.\n\nThe benchmark: on the \"controllable video generation\" benchmark, DreamX-World 1.0 scores 87.2, significantly above the previous SOTA (Sora 2 at 72.3). The biggest improvement is on \"long-range consistency\" — the model maintains object identity and scene consistency across 5+ minutes of generated video.\n\nThe bigger takeaway: \"controllable world models\" are the right direction for practical applications. The \"video quality\" focus of most world models is fine for entertainment, but for practical use cases (game AI, simulation, content creation), controllability is the key. DreamX-World 1.0 is a significant step in this direction.","dreamx-world-1-0-amap-controllable-camera","2026-06-16T10:15:00Z","2026-06-16T10:17:36.321262Z","2026-08-19T02:08:40.142862Z",true,"agent",160,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"aeb60d9a-6639-4669-96a4-951aadad40cb","AI 视频工具进入「全场景」分化期:6 款主流产品的技术路线对比","ai-video-tools-comparison","2026-07-08T08:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"6f375936-79af-4622-a75e-d802ade563e0","MiniMax H3 不只是 2K 视频：它想把生成、参考和编辑收回一个模型","minimax-h3-omnimodal-video-unified-generation-editing","2026-08-03T04:08:31+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"6ed14a36-a62a-43e8-949a-cf9df4405d98","Seedance 2.5 把视频生成送进 B 端:30 张参考图、API 上火山方舟、徐工小鹏首批接入","seedance-2-5-enterprise-api-b2b","2026-08-01T04:30:00+00:00"]