[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-jd-echowm-omnimodal-world-model":3,"news-related-b4214f43-353e-42e3-b48e-92dd4fc64290":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","京东 Echo 团队开源全模态世界模型 EchoWM:跟随连续 6-DoF 导航轨迹,同步生成 720p 视频、环境音、音乐与语音,支持第一\u002F第三人称交互,长时程 rollout 保持音画同步。权重与推理代码已开放,限学术与非商业用途。","世界模型赛道这两年的共同瓶颈:画面越来越清晰,交互却停留在离散指令——按方向键、点目标点,模型响应一帧,谈不上\"进入\";而音频在多数工作里长期缺位,开源世界模型大多只出画面不出声。8 月 24 日,京东 Echo 团队(Joy Future Academy, JD)发布 EchoWM,试图把这两块一起补上:一个跟随连续 6-DoF 导航轨迹、同步生成 720p 视频、环境音、音乐与语音的全模态世界模型([arXiv:2608.23189](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.23189))。\n\n## 能\"走进去\",而不是只准看\n\nEchoWM 的定位是 omnimodal(全模态)world model:在 720p 分辨率下,画面、环境音、音乐、语音四路输出同步生成,同时响应连续的 6 自由度导航轨迹,第一人称与第三人称视角都支持。论文摘要的表述是 \"responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech\"。\n\n交互组织方式是它比较有意思的设计:论文提出用 camera intent(相机意图)统一两类场景——第一人称时它指定观察者怎么动;第三人称时,相机与角色的动力学直接从数据里学,不需要为每个视角单独做控制器。离散指令和连续位姿都被映射到一条共享的、米制尺度的相对 6-DoF 轨迹上,再用数据集级校准(dataset-level calibration)保证混用不同来源数据训练后运动幅度依然一致。\n\n## 8 步长视频,4 步因果世界模型\n\n训练侧,EchoWM 构建了一套互补数据引擎来联合学音画生成与轨迹控制,先渐进式训练、再做自回归后训练以拉长生成时程。开源仓库给出的能力包括:8 步采样的长视频生成、4 步采样的因果世界模型、跨镜头记忆库维持角色外观与声音的连续性,以及在米制 6-DoF 控制下 60 秒稳定 rollout。\n\n开源的部分足够实在:GitHub 仓库 jd-opensource\u002FJoyAI-Echo(目前 1.9k star、168 fork)放出了推理代码与完整权重——约 46GB 的 safetensors 主模型加约 24GB 的 Gemma-3-12B 文本编码器;默认 25fps、241 帧、1280×736 配置下峰值显存约 46–50GB,一张 48GB 卡或 80GB 的 H100\u002FA100 能跑通。论文提交一天,已在 Hugging Face Daily Papers 拿下 64 个 upvote,冲上 8 月 25 日热榜前排。\n\n## 开源有诚意,但边界也清楚\n\n需要泼冷水的点同样明显。第一,底座是 Lightricks 的 LTX2.3 视频生成器加 Google Gemma 文本编码器——EchoWM 是站在 LTX 肩膀上的改造与后训练,不是从零预训练。第二,许可证限学术研究与非商业使用(继承 LTX-2 Community License),商用需联系 Lightricks,想直接拿去做产品的团队要先读条款。第三,当前版本只支持文生视频与多镜头长视频(带音视频记忆),图生视频(I2V)尚未支持,官方称后续版本会补上。\n\n仓库更新里,JoyAI-Echo 1.5 已宣布、代码\"即将发布\",TODO 列表里还有超分模块 Echo-SR 和 Director Agent。社区侧已有第三方 ComfyUI 节点,支持逐镜头改提示词、跨镜头记忆串联。\n\n## 所以呢\n\n世界模型的下一轮竞争,大概率不在\"更清晰的一帧\",而在两件事:连续控制(你能走进去)与音频补全(走进去之后有声音)。EchoWM 把这两件事打包进了一次开源发布——许可证虽然限制了商用,但它给了社区一个能复现、能改造的全模态基线。对做交互式内容、机器人仿真数据生成或游戏原型的团队,这 46GB 权重值得拉下来看一眼,参考实现与论文都在 [GitHub](https:\u002F\u002Fgithub.com\u002Fjd-opensource\u002FJoyAI-Echo)。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.23189","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"3da21e89-9d96-44eb-aedf-c09fac270ab6","en","JD Open-Sources EchoWM: An Omnimodal World Model With 720p Synced Audio-Visual Generation","JD's Echo team open-sources EchoWM, an omnimodal world model that follows continuous 6-DoF navigation while jointly generating 720p video, environmental sound, music, and speech, supporting first- and third-person interaction with synchronized audio over long rollouts. Weights and inference code are available for academic, non-commercial use.","The shared bottleneck of the world-model race over the past two years: frames keep getting sharper, but interaction stays stuck on discrete commands — press a key, click a waypoint, the model responds one frame at a time. That is \"watching\", not \"entering\". And audio has long been missing from most open work: most open-source world models render video with no sound. On August 24, the Echo team at JD (Joy Future Academy, JD) released EchoWM, an attempt to fix both at once: an omnimodal world model that follows continuous 6-DoF navigation trajectories while jointly generating 720p video, environmental sound, music, and speech ([arXiv:2608.23189](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.23189)).\n\n## A world you walk into, not just look at\n\nEchoWM positions itself as an omnimodal world model: at 720p, video, environmental sound, music, and speech are generated in sync, while the model responds to continuous 6-DoF navigation in both first-person and third-person views. The paper's own phrasing: \"responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech\".\n\nThe interaction design is the interesting part. The paper organizes control around \"camera intent\": in first-person scenes it specifies how the observer moves; in third-person scenes, camera–character dynamics are learned directly from data, with no view-specific controllers. Discrete commands and continuous poses are both mapped onto a shared metric-scale relative 6-DoF trajectory, and dataset-level calibration keeps motion magnitude consistent across heterogeneous training data.\n\n## 8-step long video, 4-step causal world model\n\nOn the training side, the team built a complementary data engine to jointly learn audio-visual generation and trajectory control, with progressive training followed by autoregressive post-training for long-horizon generation. The open repository lists: 8-step long-video generation, a 4-step causal world model, composable cross-shot memory that preserves character appearance and voice, and stable 60-second rollouts under metric 6-DoF control.\n\nThe open-source part is substantial: the GitHub repo jd-opensource\u002FJoyAI-Echo (currently 1.9k stars, 168 forks) ships inference code and full weights — a ~46 GB safetensors main model plus a ~24 GB Gemma-3-12B text encoder. At the default 25 fps, 241 frames, 1280×736 setting, peak VRAM sits around 46–50 GB, so a single 48 GB card or an 80 GB H100\u002FA100 is enough. One day after submission, the paper picked up 64 upvotes on Hugging Face Daily Papers, reaching the front of the August 25 trending list.\n\n## Generous release, clear limits\n\nThe cold-water points are equally clear. First, the base is Lightricks' LTX2.3 video generator plus a Google Gemma text encoder — EchoWM is a modification and post-training effort built on LTX shoulders, not a from-scratch pretrain. Second, the license restricts use to academic research and non-commercial purposes (inheriting the LTX-2 Community License); commercial use requires contacting Lightricks, so product teams should read the terms first. Third, the current release supports text-to-video and multi-shot long video with paired audio-video memory only — image-to-video (I2V) is not yet supported, and the team says a future version will add it.\n\nPer the repo updates, JoyAI-Echo 1.5 has been announced with code \"coming soon\", and the TODO list still includes the Echo-SR super-resolution module and a Director Agent. On the community side, a third-party ComfyUI node already exists, supporting per-shot editable prompts and cross-shot memory chaining.\n\n## So what\n\nThe next round of world-model competition probably won't be about \"a sharper frame\". It will be about two things: continuous control (can you actually walk in?) and audio completeness (once you're in, is there sound?). EchoWM bundles both into one open release — the license limits commercial use, but it hands the community a reproducible, modifiable omnimodal baseline. For teams working on interactive content, robot-simulation data generation, or game prototyping, these 46 GB of weights are worth a look; the reference implementation and paper are both on [GitHub](https:\u002F\u002Fgithub.com\u002Fjd-opensource\u002FJoyAI-Echo).","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00Z","2026-08-25T23:10:52.482702Z","2026-08-25T23:10:52.482716Z",true,"agent",42,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"65c90f98-0017-45ff-869e-a8cd251d7498","SCAIL-2：智谱+清华用端到端架构改写角色动画\"骨架法则\"","scail-2-zhipu-tsinghua-end-to-end-skeleton","2026-06-10T06:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"aad00b18-d354-48b5-ad21-62b53150b8c6","MiniMax H3 开源实测:你下载的权重,和 API 里跑的不是同一个模型","minimax-h3-local-vs-api-gap","2026-08-15T17:07:24+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00"]