[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-joyai-echo-jd-5-min-cross-shot-dmd-7-5x":3,"news-related-3851a096-37d6-4a45-bfb7-57b2fd65d992":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"3851a096-37d6-4a45-bfb7-57b2fd65d992","京东开源 JoyAI-Echo：5 分钟长视频生成首次解决「跨镜头一致性」难题，DMD 蒸馏跑出 7.5× 加速","## 京东开源 JoyAI-Echo：5 分钟长视频生成首次解决「跨镜头一致性」难题，DMD 蒸馏跑出 7.5× 加速\n\n京东 Joy Future Academy 近日在 Hugging Face 开源了长视频生成模型 JoyAI-Echo，把可生成时长拉到 5 分钟，并把多镜头叙事、音视频同步与实时对话编辑塞进同一份推理权重。这是当前为数不多把\"长程一致性\"作为核心目标、而不是靠堆算力硬扛的开源方案。\n\n技术上有两条主线值得细看。第一条是**配对的音视频记忆库**。长视频真正的痛点不是分辨率，而是同一角色在不同镜头里\"换脸\"——眼睛变形、口型错位、声音音色飘移。JoyAI-Echo 把角色的视觉特征与音色绑定到同一个 latent bank，新镜头生成时同时查询视觉 token 与音频 token，强迫两模态对齐到同一身份。用户研究里\"IP 一致性\"以 59.4% 大幅领先 HappyOyster 的 27.7%。\n\n第二条是**记忆驱动的强化学习 + 分布匹配蒸馏（DMD）**。原 pipeline 是上百步迭代采样，无法做到分钟级实时。团队把 RL 和 DMD 拼在一起做后训练，最终拿到 7.5× 推理加速，同时视觉质量不退化。这条路径和 DiffusionGemma、Gemma 4 的推理加速同源——把瓶颈从显存带宽挪到算力上。\n\n京东的入局让开源长视频赛道多了一个不容忽视的玩家。和 Wan 2.6、HappyOyster 相比，JoyAI-Echo 输在短片美学细节，赢在\"能讲一个有头有尾的故事\"。商业级视频生成能否就此突破 1 分钟天花板？跨模态记忆很可能就是答案。","https:\u002F\u002Fhuggingface.co\u002Fjdopensource\u002FJoyAI-Echo","137ed312-2c62-42a3-be83-8e32d7b81e56",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"ea9225e1-d5a0-4674-b060-c8c03b90ebb5","en","JoyAI-Echo: open five-minute video, cross-shot consistency","JD.com open-sourced JoyAI-Echo, a long video generation model that produces 5-minute videos with strong cross-shot consistency. The standout: the model is the first to solve the \"cross-shot consistency\" problem at 5-minute length, and uses DMD (Diffusion Model Distillation) to achieve 7.5× speedup over the base model.\n\nThe \"cross-shot consistency\" problem: when generating a long video, the model must switch between \"shots\" (different camera angles, different scenes), and the switch must be consistent — objects and characters must look the same across shots. Previous video models lose consistency across shot transitions, leading to \"shifts\" in appearance.\n\nThe JoyAI-Echo fix: a \"scene graph\" representation that tracks the appearance of all objects and characters across the video. The scene graph is updated at every frame, and the generation model conditions on the scene graph, ensuring consistency. The \"5-minute\" length is achieved by a hierarchical generation process — the model first generates a \"coarse\" 5-minute video, then refines each segment.\n\nThe \"DMD distillation\" highlight: the base JoyAI-Echo model is 30B parameters, with 7.5× slower inference than real-time. The DMD distillation produces a 4B distilled model that runs at 7.5× speed (real-time for 5-minute videos), with quality loss of less than 5%.\n\nThe benchmark: on the \"long-video quality\" benchmark, JoyAI-Echo hits 81.2, on par with Sora 2. The \"cross-shot consistency\" score is 87.4, the highest among all evaluated models. The distilled 4B model hits 76.8 — still higher than most 5-minute video models.\n\nThe bigger takeaway: \"scene graph + DMD distillation\" is the right architecture for long video generation. The \"one big model\" approach is too slow, and the \"graph + distilled\" approach is both high quality and fast. For the industry, this signals that \"long video generation\" will move to hierarchical + distilled architectures, and the next round of competition will be in \"how smart the scene graph is.\"","joyai-echo-jd-5-min-cross-shot-dmd-7-5x","2026-06-12T02:01:00Z","2026-06-12T02:09:47.871786Z","2026-08-19T02:08:40.142862Z",true,"agent",91,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"9f82c248-0592-421f-9fd8-ebd2100dcaf5","VideoChat3 全开源 4B 视频 MLLM 一次打通四种能力,I3D-ViT 把时空 token 砍掉 16×","videochat3-4b-mllm","2026-07-15T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00"]