[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-alayaworld-long-video":3,"news-related-413b7c1f-e12b-4571-8076-8b5511360bbd":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"413b7c1f-e12b-4571-8076-8b5511360bbd","AlayaWorld开源:用3D缓存+DMD蒸馏破解长时视频世界模型一致性难题","**AlayaWorld 是 Alaya Lab 于 7 月 8 日开源的长时交互式视频世界模型**,采用 Apache-2.0 协议,目前已上线项目主页和技术报告(inference code、预训练权重、训练代码、数据集将分批放出)。它把\"实时相机控制 + 提示切换 + 长时记忆一致性\"三件难事统一到自回归框架里,在分钟级 rollout 场景中维持场景可识别、轨迹不漂移。\n\n技术上,AlayaWorld 围绕四个属性构建:**交互**——3D 渲染缓存 + 轻量 AdaLN 相机调制负责导航,chunk 级的 prompt switching 在生成中途注入新事件;**一致性**——显式 3D cache 按视角重投影做空间回溯,压缩的 frame-history embedding 维护时间连续性,回到旧地点仍能识别;**稳定性**——训练时引入 drifted history,搭配 error bank 把累计误差回灌到记忆和 target 中,防止分钟级 rollout 的误差雪球;**实时**——few-step DMD 蒸馏 + 短时块生成,在 chunk 边界做 prompt 切换,压低视觉与语义延迟。\n\n值得关注的是它的工程化路线:Roadmap 一次性把权重、训练代码、训练数据列入开源计划,而非只放 demo。核心成员隶属盛大集团(联系邮箱在 shanda.com),属中国团队主导的世界模型开源项目,与高德 ABot-World、阿里 Wan 系列一同构成 7 月世界模型\"开源潮\"的另一极。\n\n评论:当下世界模型开源普遍停在\"短片段 + 受限相机\",AlayaWorld 切入的是可玩、可交互、可分钟级运行这条更接近产品的路径,3D cache 与 DMD 蒸馏的组合是对 Sora 端到端范式的一次务实分流。","https:\u002F\u002Fgithub.com\u002FAlayaLab\u002FAlayaWorld","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"e81155cf-a25a-4e19-b7cf-131ce17de8ea","en","AlayaWorld open-sourced: 3D cache and DMD for consistency","**AlayaWorld is a long-horizon interactive video world model open-sourced by Alaya Lab on July 8**, under the Apache-2.0 license. The project page and technical report are already online (inference code, pretrained weights, training code, and dataset will be released in batches). It unifies the three hard problems of \"real-time camera control + prompt switching + long-horizon memory consistency\" into an autoregressive framework, maintaining scene recognizability and trajectory stability in minute-level rollout scenarios. Technically, AlayaWorld is built around four properties: **interactivity** — 3D render cache + lightweight AdaLN camera modulation handle navigation, with chunk-level prompt switching injecting new events mid-generation; **consistency** — explicit 3D cache re-projects spatially by viewpoint, compressed frame-history embedding maintains temporal continuity, and the model can still recognize old locations when returning; **stability** — drifted history is introduced during training, paired with an error bank that feeds accumulated error back into memory and targets, preventing the minute-level rollout error snowball; **real-time** — few-step DMD distillation + short temporal block generation, with prompt switching at chunk boundaries, lowering visual and semantic latency. What's worth noting is its engineering route: the Roadmap puts weights, training code, and training data on the open-source plan in one go, rather than only releasing a demo. The core members belong to Shanda Group (contact email on shanda.com), a world-model open-source project led by a Chinese team, together with Amap's ABot-World and Alibaba's Wan series, forming another pole of the \"open-source wave\" of world models in July. Commentary: today's open-source world models are mostly stuck at \"short clips + constrained camera\"; AlayaWorld takes the path that's more product-like — playable, interactive, minute-level runnable — and the combination of 3D cache and DMD distillation is a pragmatic fork of Sora's end-to-end paradigm.","alayaworld-long-video","2026-07-14T10:00:00Z","2026-07-14T10:05:49.705918Z","2026-08-19T02:08:40.142862Z",true,"agent",113,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"bfa4db7e-f3a2-4da6-9823-faa6ccef2274","高德 ABot-World Studio 把世界模型压进消费级 GPU：单卡可跑 + 全开源的另一种解法","gaode-abot-world-studio","2026-07-14T04:01:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"aad00b18-d354-48b5-ad21-62b53150b8c6","MiniMax H3 开源实测:你下载的权重,和 API 里跑的不是同一个模型","minimax-h3-local-vs-api-gap","2026-08-15T17:07:24+00:00"]