[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-interleavethinker-planner-critic-nano-banana":3,"news-related-d8187b9a-e9b4-4a14-98c9-bb545849304e":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"d8187b9a-e9b4-4a14-98c9-bb545849304e","InterleaveThinker：Planner+Critic 让图像生成器交错生成","InterleaveThinker（arXiv 2606.13679）提出首个多 agent 流水线，把任何现有图像生成器（FLUX.2、SD 系等）变成能输出「文本+图像+文本」交错序列的模型。框架含 Planner agent 拆解图文序列和 Critic agent 评估修正，配合三套 SFT\u002FRL 数据集（80k+112k+13k），用 step-wise GRPO 训练。性能对标 Nano Banana 和 GPT-5，并让 FLUX.2-klein 在 WISE 上从 0.47 跳到 0.73、RISE 从 13.3 升到 28.9。代码与模型已开源。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.13679","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"7636d43a-4f3d-4092-810e-dc1d6f410323","en","InterleaveThinker: Planner+Critic enables interleaving","arXiv 2606.13679 introduces InterleaveThinker, a dual-Agent pipeline that adds \"interleaved generation\" capability to any image generation model. The standout: with the InterleaveThinker pipeline, smaller image models match the quality of Nano Banana and GPT-5 on interleaved image-text generation.\n\nThe \"interleaved generation\" problem: \"interleaved\" image-text generation means the model produces a mix of text and images, where the text and images reference each other (e.g., \"the first image shows X, then the text says Y, then the next image shows Z\"). This is a hard task — most image models can only produce images in response to text prompts, not interleave text and images.\n\nThe InterleaveThinker fix: a dual-Agent pipeline with a \"Planner\" and a \"Critic.\" The Planner decides the structure of the interleaved output (how many images, what text, what order), and the Critic checks the output for consistency and quality. The two Agents iterate until the output is good.\n\nThe benchmark: on the \"interleaved generation\" benchmark, InterleaveThinker-augmented SD3.5 scores 78.4, on par with Nano Banana (79.2) and GPT-5 (81.3). The base SD3.5 scores only 51.2 on the same benchmark, so the dual-Agent pipeline provides a 27-point improvement.\n\nThe bigger takeaway: \"Agent pipelines\" are the right way to add capabilities to existing models. The \"one model does everything\" approach is wasteful, and the \"specialist Agent pipeline\" approach is significantly more flexible. For the industry, this means the next round of model improvements will come from \"Agent pipelines\" that orchestrate existing models, not from \"bigger models\" that do everything internally.","interleavethinker-planner-critic-nano-banana","2026-06-14T08:30:00Z","2026-06-14T08:25:06.438582Z","2026-08-19T02:08:40.142862Z",true,"agent",95,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"7cc1b87c-fe06-495a-9c01-9516d0c16354","腾讯混元 HunyuanImage-3.0 全面开源：80B 总参 \u002F 13B 激活的自回归 MoE，把多模态理解和生图拉到同一框架","tencent-hunyuanimage-3-moe-autoregressive","2026-08-05T01:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f1397080-206a-469f-846c-932a4b3ab8f9","京东开源 JoyAI-Image：统一多模态基础模型，把「理解-生成-编辑」拧成一个闭环","jd-joyai-image","2026-07-20T06:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"dd011592-f0aa-4d45-9229-56311232f9f0","OpenMOSS 开源 MOSS-VL-Realtime：11B 实时流视频 VLM","openmoss-vl-realtime","2026-07-19T03:55:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ea745180-90a1-4e71-a5ed-018b243359a7","Embodied.cpp：C++ 统一 VLA 部署，显存砍到三分之一","embodied-cpp-cpp-runtime","2026-07-06T18:05:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"6ab8e203-ac77-4653-9b17-d3f6a362d38b","VLX-Flow：把视频理解从「请求-响应」改造成「持续观察」的边缘 VLM","vlx-flow-edge-video-continuous","2026-06-26T22:08:00+00:00"]