[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-imagewam-sjtu-world-action-model-1-6-flops":3,"news-related-844a2f3b-682d-4d03-b8ab-f02ec1c16dcf":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"844a2f3b-682d-4d03-b8ab-f02ec1c16dcf","ImageWAM 抛弃视频生成：用图像编辑做世界动作模型，FLOPs 降到 1\u002F6","ImageWAM 把「世界动作模型」（World Action Model，WAM）的视觉部分从视频生成换成图像编辑，思路直击当前主流方案的痛点。\n\n现有视频 WAM 三大痼疾：1）密集多帧未来 token 推起来昂贵；2）视频预测在动作无关的时序\u002F外观细节上浪费容量；3）长时想象易引入误差，误导动作预测。SJTU 团队由此发问：WAM 真的需要视频生成吗？\n\nImageWAM 的核心做法是复用预训练图像编辑模型（FLUX.2-4B\u002F9B）做单帧目标变换预测，并仅在编辑去噪的 KV cache 上挂一个 flow-matching 动作专家。推理时甚至不解码目标帧，让编辑 cache 直接充当「世界-动作上下文」。\n\n为什么有效？其一，图像编辑天然是「当前帧→目标帧」的变换先验，与动作预测需求天然对齐；其二，编辑模型经过指令-局部视觉变化的专门预训练，能聚焦任务相关区域；其三，注意力分析显示编辑 cache 自动集中在任务相关变化区域，而非无关纹理。\n\n实验结果亮眼：在 RoboTwin、LIBERO 仿真与真机任务上，ImageWAM 超过标准 VLA 基线和同等规模视频 WAM，无需额外策略预训练；FLOPs 仅为视频方案的 1\u002F6，延迟降至 1\u002F4。\n\n这反映了一个更深层趋势：把「通用预测器」收敛为「任务相关生成器」。当目标只是让当前机器人完成一个动作，多帧视频预测里的帧间冗余与无关纹理是浪费；精准的图像编辑先验 + KV cache 复用，让算力消耗降一个数量级的同时性能更好。\n\n值得跟进的问题：编辑 cache 在长视野、多步规划上的稳定性；能否扩展到灵巧手、导航等更广任务；以及当指令从单步变成多步复合时，编辑先验的精度是否会被放大成误差。这条路线把世界模型从「无条件全帧预测」解放出来，给视频-图像混合架构开辟了更经济的工程范式。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.19531","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"94573676-efbf-43a3-acad-f0744f0f8f3b","en","ImageWAM: world action model via image editing, 1\u002F6 the FLOPs","arXiv 2606.19531 introduces ImageWAM, a \"world action model\" that uses image editing (rather than video generation) as the underlying primitive. The result: 1\u002F6 the FLOPs of equivalent video-generation models, with comparable downstream task performance.\n\nThe \"abandon video generation\" angle: most world models for robotics and game AI are video-generation models — they predict the next frame. ImageWAM's insight: the \"next frame\" is mostly redundant with the \"current frame\" — only the changed regions matter. So instead of generating a full new frame, ImageWAM predicts the \"edit map\" (a sparse set of pixel changes) between the current and next frame.\n\nThe technical details: ImageWAM is trained on a large corpus of \"frame pairs\" (current frame + next frame), with the training objective being to predict the \"edit map\" — a low-resolution per-pixel change map. At inference, the model generates the edit map and applies it to the current frame to produce the next frame. The edit map is 1\u002F16 the size of a full frame, so the FLOPs drop by 6×.\n\nThe benchmark: on a set of robotics and game-AI tasks, ImageWAM matches the performance of video-generation world models at 1\u002F6 the compute. The \"edit-based\" representation also has a nice side effect: the model can be queried for \"what would change if I do X\" — i.e., counterfactual reasoning, which is hard for video-generation models.\n\nThe bigger takeaway: \"world models don't have to be video models.\" The assumption that \"world model = video generation\" has been dominant, but ImageWAM shows that a sparse edit-based representation can be just as effective at 1\u002F6 the cost. For the industry, this means world models for robotics and game AI can be deployed on much cheaper hardware, opening up new use cases.","imagewam-sjtu-world-action-model-1-6-flops","2026-06-22T06:20:00Z","2026-06-22T06:19:23.248385Z","2026-08-19T02:08:40.142862Z",true,"agent",243,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"18d2aa73-7244-4b10-b611-46475e17327e","ForgeWM开源:一步去噪72FPS的可玩世界模型,8张卡复现全流程","forgewm-few-step-playable-world-model","2026-08-24T21:10:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"fb97a60d-69a1-4988-8de6-d1540ba63359","2.4B 参数读懂整页 A4:Cohere Labs 把最小的多模态模型挂上了 Apache 2.0","cohere-north-micro-vision-open-vlm","2026-08-18T13:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4244f57a-3afa-465c-aa67-793df6eba5cc","LFM2.5-VL-3B 开源：3.1B 参数让手机读懂屏幕、框住物体、自己调工具","liquid-ai-lfm2-5-vl-3b-edge-vlm","2026-08-14T13:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00"]