[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-jd-joyai-video-edit-realtime-diffusion":3,"news-related-6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","京东开源实时流式视频编辑模型 JoyAI-Video-Edit,基于 16B 参数多模态扩散 Transformer,端到端 720p 视频达到 30.19 FPS,让视频创作从「先有素材再修改」变为可实时互动编辑,代码、权重与在线 Demo 一并放出。","京东( JD.com ) 近日在 GitHub 开源了 **JoyAI-Video-Edit**,一个面向开放域流式视频的实时编辑系统。该模型由 MLLM 条件编码器、因果视频 VAE 与 16B 参数多模态扩散 Transformer(MMDiT)组成,通过自回归扩散 + 蒸馏的方式,把「边播放、边用自然语言改视频」这件事,从离线批量变成了实时的流式生成。\n\n## 关键能力\n\n- **实时开放域编辑**:对摄像头流或上传视频,模型随帧到达就因果式地编辑,不需要预先知道总长度,也不需要回头处理未来帧。\n- **指令多样性**:支持主体编辑、局部编辑、背景替换、风格迁移、动作调整、参考图引导等常见视频编辑任务。\n- **720p 高吞吐部署**:官方部署基准显示,整条端到端管线在 720×1280 分辨率下达到 **30.19 FPS**,把视频编辑从离线批处理推进到交互式流式生成。\n- **训练-推理一致性**:通过 aligned autoregressive distribution matching distillation、长时序优化与 bounded KV-state inference 缓解 train-inference mismatch 与累积时序漂移。\n\n## 技术结构\n\n系统的三个核心组件各司其职:\n\n- **MLLM 条件编码器**:把用户的自然语言指令(以及参考图)编码为条件信号,交给下游扩散 Transformer。\n- **因果视频 VAE**:在 token 空间下做时空压缩与还原,是「实时流式」得以成立的关键 —— 不能每帧重算一整段视频。\n- **16B MMDiT 骨干**:多模态扩散 Transformer 负责在压缩后的 token 序列上做扩散式编辑。\n\n部署侧提供了 persistent TorchInductor \u002F Triton \u002F CUDA cache 复用、bounded KV-state 等调度策略,目标是让 720p 编辑保持稳定 per-chunk 算力。\n\n## 落地与开源\n\n京东同时放出了完整部署代码、技术报告( arXiv:2608.03974 )、在线 Demo( joyai-labs.jd.com\u002Fv2v ) 与 Hugging Face 权重,协议为 **Apache 2.0**。仓库的 `DEPLOYMENT.md` 里写了完整的环境、checkpoint 与自定义部署流程,默认监听 `http:\u002F\u002Flocalhost:8080`。\n\nGitHub 仓库的 TODO 也直接列出了下一阶段动作:\n\n- 面向消费级 GPU(如 GeForce RTX 5090)的部署优化;\n- 更强版本正在训练,重点是参考图引导视频编辑( RV2V );\n- 完整训练与数据 pipeline 开源。\n\n## 我的看法\n\n视频生成的「最后一公里」正从「离线剪」转向「实时改」。JoyAI-Video-Edit 把这件事用自回归扩散 + 流式 VAE 的组合做成可工业部署的形态 —— 30.19 FPS 在 720p 端到端是工程上能讲得通的数字,关键是它开源了权重、训练报告和部署代码,而不是只放了论文。\n\n值得关注的两个信号:\n\n- **大厂押注物理世界模型矩阵**:京东的措辞是「物理世界模型矩阵进一步完善」,JoyAI-Video-Edit 同时也为具身智能的大规模数据合成积累底层技术 —— 实时编辑 + 因果流式 = 合成数据流水线上的关键一环。\n- **多模态扩散的「实时化」拐点**:Sora、Veo 那一代主打的是「质量」,但当 16B 模型已经能在 720p 跑 30+ FPS,「实时性」就接棒成为下一个差异化战场。对个人创作者和直播电商来说,这意味着「改镜头」会从后期工种变成现场操作。\n\n下一步看的是消费级 GPU 适配和 RV2V 强参考图能力 —— 谁能先把这两个做扎实,谁就能把「实时视频编辑」从 demo 推进到日常工具。\n\n参考:\n- GitHub 仓库: https:\u002F\u002Fgithub.com\u002Fjd-opensource\u002FJoyAI-Video-Edit\n- 技术报告: https:\u002F\u002Farxiv.org\u002Fpdf\u002F2608.03974\n- Hugging Face 权重: https:\u002F\u002Fhuggingface.co\u002Fjdopensource\u002FJoyAI-Video-Edit\n- 在线 Demo: https:\u002F\u002Fjoyai-labs.jd.com\u002Fv2v\u002F","https:\u002F\u002Fgithub.com\u002Fjd-opensource\u002FJoyAI-Video-Edit","137ed312-2c62-42a3-be83-8e32d7b81e56",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":25,"name":26,"slug":26,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"cf1b0df1-d2b1-48d9-977c-166e23a04eac","en","JoyAI-Video-Edit: JD.com's 16B model for streaming edits","JD.com has open-sourced JoyAI-Video-Edit, a real-time streaming video editing model built on a 16B-parameter multimodal diffusion Transformer. End-to-end 720p video reaches 30.19 FPS, turning video creation from a \"shoot first, edit later\" workflow into an interactive live-editing experience. Code, weights, and an online demo ship together.","JD.com has open-sourced **JoyAI-Video-Edit** on GitHub — a real-time, instruction-guided video editing system aimed at open-ended video streams. The model combines an MLLM-based condition encoder, a causal video VAE, and a 16B-parameter multimodal diffusion Transformer (MMDiT). Through autoregressive diffusion plus distillation, it turns \"edit videos with natural language as they play\" from an offline batch task into a real-time streaming generation workflow.\n\n## Key capabilities\n\n- **Real-time open-ended editing.** For live camera streams or uploaded videos, the model edits frames causally as they arrive — no need to know the total length up front, and no need to revisit future frames.\n- **Diverse instruction control.** Supports subject edits, local edits, background replacement, style transfer, motion changes, and reference-guided editing.\n- **High-throughput 720p deployment.** The official deployment benchmark shows the end-to-end pipeline reaching **30.19 FPS at 720×1280**, pushing video editing from offline batch processing into interactive streaming generation.\n- **Train-inference consistency.** Aligned autoregressive distribution matching distillation, long-horizon optimization, and bounded KV-state inference mitigate train-inference mismatch and accumulated temporal drift.\n\n## Architecture\n\nThe three core components divide the work cleanly:\n\n- **MLLM condition encoder** encodes the user's natural language instructions (and reference images) into condition signals for the downstream diffusion Transformer.\n- **Causal video VAE** performs spatiotemporal compression and reconstruction in the token space — this is what makes \"real-time streaming\" viable, because the model can't afford to recompute the whole video every frame.\n- **16B MMDiT backbone** performs diffusion-style editing on the compressed token sequence.\n\nOn the deployment side, the project ships with persistent TorchInductor \u002F Triton \u002F CUDA cache reuse, bounded KV-state scheduling, and a stable per-chunk compute budget — the goal is to keep 720p editing consistent under load.\n\n## Release and availability\n\nJD.com released the full deployment code, technical report (arXiv:2608.03974), online demo (joyai-labs.jd.com\u002Fv2v), and Hugging Face weights under the **Apache 2.0** license. The repository's `DEPLOYMENT.md` covers environment setup, checkpoint preparation, and custom deployment. The default server listens on `http:\u002F\u002Flocalhost:8080`.\n\nThe GitHub TODO also points at the next steps:\n\n- Consumer-GPU deployment optimization (e.g. GeForce RTX 5090).\n- A stronger version is in training, with a focus on reference-image-guided video editing (RV2V).\n- Full training and data pipeline open-sourcing.\n\n## Why it matters\n\nThe \"last mile\" of video generation is shifting from \"offline editing\" to \"live editing.\" JoyAI-Video-Edit turns that shift into a deployable form through the combination of autoregressive diffusion and a streaming VAE — 30.19 FPS end-to-end at 720p is a number that holds up in engineering terms, and crucially the team released weights, a training report, and deployment code, not just a paper.\n\nTwo signals worth watching:\n\n- **A big-tech bet on a physical-world model matrix.** JD.com's framing — \"the physical-world model matrix continues to grow\" — and the explicit link to embodied-AI data synthesis suggest that real-time editing + causal streaming is a key building block of the next-generation synthetic data pipeline.\n- **The \"real-time\" inflection in multimodal diffusion.** The Sora \u002F Veo generation was about \"quality.\" Now that a 16B model can run at 30+ FPS at 720p, real-timeness is the next differentiator. For individual creators and live e-commerce, that means \"change the shot\" stops being a post-production step and becomes something you do live.\n\nWhat to watch next: consumer-GPU support and stronger RV2V reference-image capability. Whoever nails those two first gets to move real-time video editing from demo to daily tool.\n\n## References\n\n- GitHub repository: https:\u002F\u002Fgithub.com\u002Fjd-opensource\u002FJoyAI-Video-Edit\n- Technical report: https:\u002F\u002Farxiv.org\u002Fpdf\u002F2608.03974\n- Hugging Face weights: https:\u002F\u002Fhuggingface.co\u002Fjdopensource\u002FJoyAI-Video-Edit\n- Online demo: https:\u002F\u002Fjoyai-labs.jd.com\u002Fv2v\u002F","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00Z","2026-08-10T10:06:09.210092Z","2026-08-10T10:06:09.210102Z",true,"agent",196,{"items":42},[43,48,53,58,63,68],{"id":44,"title":45,"news_slug":46,"published_at":47},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"5612d186-46ee-4509-9a93-94045ba004ae","LTX-2.5 开放权重视频模型:4K 反而在 Fast 端点,EXR 色彩管线也焊进去了","ltx-2-5-open-weights-video","2026-08-18T15:20:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"d3e01f3d-745b-4c98-9289-38081a3f5f06","FLUX 3：图像\u002F视频\u002F音频统一进 flow matching","bfl-flux-3-flow-matching","2026-07-27T10:00:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"9f82c248-0592-421f-9fd8-ebd2100dcaf5","VideoChat3 全开源 4B 视频 MLLM 一次打通四种能力,I3D-ViT 把时空 token 砍掉 16×","videochat3-4b-mllm","2026-07-15T02:00:00+00:00"]