[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-flow-map-grpo-few-step":3,"news-related-52e96e5b-ec3b-4552-b94a-93cc2702ab84":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"52e96e5b-ec3b-4552-b94a-93cc2702ab84","Flow-Map GRPO：为确定性「少步生图」打开强化学习大门","\n少步 Flow-Map 生成器——如一致性模型、sCM、MeanFlow——过去两年一直是图像扩散模型里最快的一档：它们直接学习噪声到数据的「长途运输映射」，把采样步数压到个位数。但确定性正是它们的阿喀琉斯之踵：GRPO、PPO 这类需要随机轨迹和良好似然比的在线 RL 后训练方法，长期以来无法直接套用。\n\n**Flow-Map GRPO**（arXiv:2607.00535，7 月 1 日）解决了这一卡点。它的核心机制是 **ASFMC（Anchored Stochastic Flow Map Composition）**：通过基于锚点的条件重采样注入随机性，同时完整保留原始 Flow-Map 的边缘概率路径。这样既不破坏少步生成的高效性，又让 GRPO 目标函数可以求梯度。作者还推导出同时适用于**单步**和**两步** Flow-Map 参数化的 GRPO 目标，并在基于 FLUX 后端的 MeanFlow 与 sCM 上验证，多项奖励\u002F感知\u002F任务级指标全部上涨。\n\n最有看点的是它的工程哲学：**无需重新训练**。Flow-Map GRPO 把后训练做成「外挂」模块，直接对预训练好的确定性生成器做对齐，不需要改参数化，也不必把模型再训成原生随机模型——对存量 checkpoint 极其友好，这也意味着 FLUX、Qwen-Image、Wan 等已经上线的少步工作流未来都可以低成本接 RL。\n\n对 GRPO+扩散社区，这是继 6 月末 Qwen-Image-2.0-RL 技术报告之后的又一个信号：**后训练范式正在从 LLM 辐射到图像生成**，而少步 Flow-Map 是这条扩展路径上一直缺位的一环。Flow-Map GRPO 补齐它以后，行业剩下的主要是工程问题——更稳定的奖励模型、更大规模的人类偏好对齐、面向特殊风格的可控强化，都会沿着这个口子铺开。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.00760v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"fb27eadf-41c8-4507-ba32-f6800410047d","en","Flow-Map GRPO opens RL for deterministic few-step imaging","# Flow-Map GRPO: opens the door to reinforcement learning for deterministic \"few-step image generation\" Few-step Flow-Map generators — like consistency models, sCM, MeanFlow — have been the fastest tier in image diffusion models over the past two years: they directly learn the \"long-distance transport map\" from noise to data, compressing the sampling steps to single digits. But determinism is their Achilles' heel: online RL post-training methods like GRPO and PPO that need random trajectories and good likelihood ratios have long been unable to be directly applied. **Flow-Map GRPO** (arXiv:2607.00535, July 1) solves this bottleneck. Its core mechanism is **ASFMC (Anchored Stochastic Flow Map Composition)**: injects randomness through anchored conditional resampling while fully preserving the original Flow-Map's marginal probability path. This doesn't break the efficiency of few-step generation, while making the GRPO objective function differentiable. The authors also derive GRPO objectives applicable to both **one-step** and **two-step** Flow-Map parameterizations, and verify on MeanFlow and sCM based on the FLUX backend, with multiple reward \u002F perception \u002F task-level metrics all rising. The most noteworthy is its engineering philosophy: **no retraining needed**. Flow-Map GRPO makes post-training a \"plug-in\" module, directly aligning pretrained deterministic generators, without changing parameterization, and without having to retrain the model into a native stochastic model — extremely friendly to existing checkpoints, which also means FLUX, Qwen-Image, Wan and other already-online few-step workflows can in the future all hook up RL at low cost. For the GRPO+diffusion community, this is another signal after the late-June Qwen-Image-2.0-RL tech report: **the post-training paradigm is radiating from LLM to image generation**, and few-step Flow-Map has been the missing link on this extension path. After Flow-Map GRPO fills it, the remaining things for the industry are mainly engineering — more stable reward models, larger-scale human-preference alignment, controllable RL for specific styles, will all unfold along this opening.","flow-map-grpo-few-step","2026-07-05T12:02:00Z","2026-07-05T12:04:44.695763Z","2026-08-19T02:08:40.142862Z",true,"agent",147,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"774de6ac-98e1-4343-a67a-bfdc72d377bb","INFORMS 实证:AI 广告真实投放胜过设计师,18 个月后仍领先","informs-ai-ads-beat-human-designers-18-months","2026-08-22T14:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"4899d809-5d52-454e-8c23-115f597e82d4","离散扩散模型终于被「拉直」:22 位作者把 Tokenization、Masking、Score 三条路线焊进同一框架","discrete-diffusion-unified","2026-07-18T02:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"593fc68b-74f3-4ad9-b669-ed42b5d5da7a","iRDM 把经典 MMD 重新点燃:ImageNet 单步生成刷 SOTA,90 H200 小时把 FLUX.2 [klein] 蒸馏成一步","irdm-mmd-single-step","2026-07-06T06:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"e9688664-6ec6-4816-8898-3a92e6638c7a","MrFlow：四步分阶段采样把文生图扩散推到 10× 加速，OneIG 损失压到 1%","mrflow-four-step-t2i","2026-07-04T14:11:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ce1ba4c1-7d9f-470c-86f3-9cabc8e69e0a","字节DanceOPD把图像生成多能力冲突变成「场蒸馏」：硬路由+单查询就赢","bytedance-danceopd-field-distillation","2026-06-28T04:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"b51af942-496b-4b58-95fb-37980d12a743","逆向工程发现:微软画图本地生成的 AI 图像,像素里埋着服务器下发的水印 GUID","mspaint-invisible-watermark-guid","2026-08-26T13:00:00+00:00"]