[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-jd-joyai-image":3,"news-related-f1397080-206a-469f-846c-932a4b3ab8f9":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"f1397080-206a-469f-846c-932a4b3ab8f9","京东开源 JoyAI-Image：统一多模态基础模型，把「理解-生成-编辑」拧成一个闭环","京东（jd-opensource）正式开源 JoyAI-Image——一款真正意义上的统一多模态基础模型。它用一个 8B 多模态大语言模型（MLLM）做理解中枢，搭配 16B 多模态扩散 Transformer（MMDiT）做生成引擎，把「图像理解、文生图、指令式图像编辑」三件事拧进同一个模型族。\n\n最值得关注的是它的「理解⇄生成」闭环设计：以往的多模态系统理解归理解、生成归生成；JoyAI-Image 把两边打通——更强的空间理解反哺生成质量与可控编辑，而生成出的新视角又能为空间推理提供证据。这种双向激励的范式在国内开源多模态里相当少见。\n\n技术细节上，它在长文本排版、多视角生成、几何感知的空间编辑（Object Move \u002F Object Rotation \u002F Camera Control）上做了专项优化，在 spatial reasoning 与 Qwen-Image-Edit、Nano Banana Pro 等同台对标。模型权重、Diffusers 集成、ComfyUI 工作流、HuggingFace Demo 一并释放，Apache 2.0 完全开源。\n\n7 月 17 日最新更新：JoyAI-Image-Edit 与 Edit-Plus 原生支持 ComfyUI，可直接拖入工作流运行，无需额外依赖。\n\n这条路径的真正意义在于：当行业还在争论「理解派」和「生成派」谁主导，京东用统一架构给出了答案——边界本来就不必分。短期内，这会给 Qwen-Image-Edit、Seedream 等开源图像模型压力；长期看，「理解⇄生成」的双向激励范式有可能成为下一代多模态模型的标配。","https:\u002F\u002Fgithub.com\u002Fjd-opensource\u002FJoyAI-Image","137ed312-2c62-42a3-be83-8e32d7b81e56",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"434cdef7-9db5-4188-8d22-2637aa642756","en","JoyAI-Image: JD.com's unified understanding-generation-edit loop","JD.com (jd-opensource) officially open-sources JoyAI-Image — a true unified multimodal foundation model. It uses an 8B multimodal LLM (MLLM) as the understanding center, paired with a 16B multimodal diffusion Transformer (MMDiT) as the generation engine, fusing \"image understanding, text-to-image, instruction-based image editing\" into a single model family. The most noteworthy piece is its \"understanding ⇄ generation\" closed-loop design: in traditional multimodal systems, understanding is understanding and generation is generation; JoyAI-Image bridges the two — stronger spatial understanding feeds back into generation quality and controllable editing, while newly generated viewpoints provide evidence for spatial reasoning. This bidirectional-incentive paradigm is rare in domestic open-source multimodal work. On the technical side, it's specifically optimized for long-text typesetting, multi-view generation, geometry-aware spatial editing (Object Move \u002F Object Rotation \u002F Camera Control), and stands against Qwen-Image-Edit, Nano Banana Pro on spatial reasoning. Model weights, Diffusers integration, ComfyUI workflow, and HuggingFace Demo are all released, fully open-source under Apache 2.0. Latest update on July 17: JoyAI-Image-Edit and Edit-Plus natively support ComfyUI and can be dragged directly into the workflow to run, with no extra dependencies. The real meaning of this path: when the industry is still arguing over \"understanding camp\" vs \"generation camp\", JD.com gives the answer with a unified architecture — the boundary didn't have to be drawn in the first place. In the short term, this puts pressure on open-source image models like Qwen-Image-Edit and Seedream; in the long term, the bidirectional-incentive \"understanding ⇄ generation\" paradigm may become the standard for the next generation of multimodal models.","jd-joyai-image","2026-07-20T06:00:00Z","2026-07-20T06:06:49.626150Z","2026-08-19T02:08:40.142862Z",true,"agent",130,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"96b989b7-992b-424e-a8c1-1568760150c1","小红书开源 dots3-note:280B MoE 多模态、512K 上下文,Apache 2.0 直接放行","dots3-note-preview-280b-open-weights","2026-08-18T23:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"fb97a60d-69a1-4988-8de6-d1540ba63359","2.4B 参数读懂整页 A4:Cohere Labs 把最小的多模态模型挂上了 Apache 2.0","cohere-north-micro-vision-open-vlm","2026-08-18T13:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b1e41506-8ce0-4bbd-a11a-89d823998130","B 站 IndexTTS-2.5 开放权重:0.8B 参数零样本克隆五语种音色,8 维情感向量把情绪做成旋钮","indextts-2-5-bilibili-zero-shot-tts","2026-08-17T15:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"cb64371f-62b6-473d-8150-b576001d3f56","Qwen3.8-27B 开源权重上线:单卡跑得动的 Qwen3.8,还塞了个视觉编码器","qwen3-8-27b-open-weights-release","2026-08-14T19:30:00+00:00"]