[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tencent-mixgrpo-flow-grpo":3,"news-related-7c769930-c404-4ef6-a7c2-29d45d8209d2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"7c769930-c404-4ef6-a7c2-29d45d8209d2","腾讯混元 MixGRPO 入选 ECCV 2026：滑动窗口把 Flow-GRPO 训练开销砍到三成","GRPO 已经成为大模型 RL 训练的事实标准方法之一，但当它被搬到扩散文生图模型上时，工程上的痛点异常明显——Flow-GRPO、DanceGRPO 等方案要沿着整条去噪链对每一步都做 SDE 采样和策略优化，训练时间和显存消耗双双爆炸。腾讯混元团队与北大计算机学院\u002F计算机中心合作提出的 MixGRPO，被 ECCV 2026 接收，正是对这条痛线的直接回应。\n\n核心做法是引入“滑动窗口”机制——只在窗口内做 SDE 采样和 GRPO 优化，窗口外换成 ODE 采样。这种“夹心”设计带来两个层面的好处：一是把策略更新的负担压缩到小子区间，训练时间下降近 50%；二是窗口外不参与反向梯度，可以用更高阶 solver 跑得更快，由此衍生出 MixGRPO-Flash 变体，训练时间再砍 71%。\n\n作者在 FLUX.1-dev 上以 HPSv2、ImageReward、PickScore 组成多奖励组合，MixGRPO 在人类偏好对齐指标上超过 DanceGRPO，训练成本对应缩短。代码、checkpoint 与训练脚本全部开源（GitHub 1.1k stars），是少见的工业化实践与学术结果合一的成果。\n\nMixGRPO 的方法学意义不止于省时间——它把“哪些时间步真正影响策略更新、哪些只是 forward pass”这个问题讲清楚了。扩散语言模型的 RL 训练大概率会沿着“局部化优化”这条路径演进，MixGRPO 是该路径上一个非常典型的范式样本。对正在做 RL-augmented 图像生成、视频生成的工程团队来说，这条思路的价值远不止省显卡——它把“无需全链路梯度”的理念正式带进了扩散 RL 这个仍偏年轻的方向。","https:\u002F\u002Fgithub.com\u002FTencent-Hunyuan\u002FMixGRPO","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d447ed6a-b455-4c8d-9e90-a99ad2407376","en","MixGRPO at ECCV 2026: flow training cost down to 30%","GRPO has become one of the de facto standard methods for LLM RL training, but when ported to diffusion text-to-image models, the engineering pain points are unusually obvious — Flow-GRPO, DanceGRPO and other solutions need to do SDE sampling and policy optimization at every step along the entire denoising chain, and both training time and memory consumption explode. MixGRPO, proposed by Tencent Hunyuan's team together with Peking University's School of Computer Science \u002F Computer Center, and accepted by ECCV 2026, is a direct response to this pain line. The core approach introduces a \"sliding window\" mechanism — only do SDE sampling and GRPO optimization within the window, switching to ODE sampling outside the window. This \"sandwich\" design brings two layers of benefits: first, it compresses the policy update burden to a small subinterval, cutting training time by nearly 50%; second, no reverse gradient is needed outside the window, allowing a higher-order solver to run faster, from which a MixGRPO-Flash variant is derived, cutting training time by another 71%. The authors use HPSv2, ImageReward, and PickScore to form a multi-reward combination on FLUX.1-dev, and MixGRPO exceeds DanceGRPO on human-preference alignment metrics, with correspondingly shortened training cost. Code, checkpoints, and training scripts are all open-sourced (GitHub 1.1k stars), making it a rare industrial-practice-meets-academic-result output. MixGRPO's methodological significance goes beyond saving time — it makes clear the question of \"which timesteps actually affect policy updates, which are just forward pass\". RL training of diffusion language models will likely evolve along the \"localized optimization\" path, and MixGRPO is a very typical paradigmatic sample on that path. For engineering teams doing RL-augmented image and video generation, the value of this line of thinking is far beyond saving GPU — it brings the idea of \"no need for full-chain gradient\" formally into the still-young direction of diffusion RL.","tencent-mixgrpo-flow-grpo","2026-07-06T22:09:00Z","2026-07-06T22:11:40.119930Z","2026-08-19T02:08:40.142862Z",true,"agent",99,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"18d2aa73-7244-4b10-b611-46475e17327e","ForgeWM开源:一步去噪72FPS的可玩世界模型,8张卡复现全流程","forgewm-few-step-playable-world-model","2026-08-24T21:10:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"deac2d55-76a6-40d2-8ef7-36aed2ad0105","Linux 7.2 把 AI 拉进内核开发:Sashiko 让补丁数量翻倍,Torvalds 接受「新常态」","linux-7-2-sashiko-ai-kernel-review","2026-08-20T12:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a818c807-2131-4950-8f51-62847a57db41","VideoRAE 把 frozen 视频基础模型改造成生成器 latent:UCF-101 gFVD 40\u002F93,收敛提速 5×","videorae-frozen-video-generator","2026-07-20T04:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"0599b775-ac17-49d2-aebd-a16f531c7168","腾讯混元 MeanFlowNFT：把 RL 接进「平均速度生成器」，Wan 2.1 4 步反超 50 步 LongCat-Video RL","tencent-hunyuan-meanflownft","2026-07-16T12:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"7ac0ef83-f46d-44f9-846b-a2051fc81e87","NVIDIA NeMo AutoModel：MoE 微调吞吐抬到 3.4–3.7 倍","nvidia-nemo-automodel-moe-finetune-3-7x","2026-06-24T20:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"844a2f3b-682d-4d03-b8ab-f02ec1c16dcf","ImageWAM 抛弃视频生成：用图像编辑做世界动作模型，FLOPs 降到 1\u002F6","imagewam-sjtu-world-action-model-1-6-flops","2026-06-22T06:20:00+00:00"]