[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-av-grpo-audio-video-diffusion-rl":3,"topics-all":38,"news-related-af21d26d-d9cd-4526-ac4a-66366a45848c":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"af21d26d-d9cd-4526-ac4a-66366a45848c","AV-GRPO:8张A800给22B音视频模型做RL后训练","上海人工智能实验室提出 AV-GRPO,用模态锚定的在线扩散强化学习训练音视频联合生成模型,并发布五维解耦、难度可控的 5DAV 数据集。框架以三个模块把多模态偏好学习拆成单模态子问题,在 JavisBench 与 VABench 上超过 22B 基座 LTX-2.3,8 张 A800 即可完成训练。","音视频联合生成模型这两年进步很快,但三个老问题一直没解决好:单模态保真度有限、文本对齐不足、跨模态同步弱。强化学习后训练在文本模型上已经被反复验证有效,可一旦搬到音视频联合生成上,麻烦就来了——音频和视频两路奖励信号异构,纠缠在一起后信用分配搞不清;两个模态塔联合优化,算力开销和训练动态都不好处理;同步质量还依赖配对样本评估,奖励之间难以公平比较。\n\n上海人工智能实验室的团队把这个问题拆开解决。他们提出的 AV-GRPO 是一个模态锚定的在线扩散强化学习框架,配套发布 5DAV,一个按五个维度解耦、难度可控的训练数据集。论文已在 arXiv 公开(编号 2609.29816,cs.CV 与 cs.SD 交叉,22 页),代码、数据和权重全部开放。\n\n## 把多模态偏好学习拆成单模态子问题\n\nAV-GRPO 的核心思路是解耦。框架包含三个关键模块:其一,模态锚定 rollout,用来解耦学习信号、稳定训练难度;其二,轨迹锁定的冻结塔优化,训练一个模态塔时冻结另一个,既压低成本也重新分配信用;其三,针对各模态自身动态特性设计的自适应目标与扰动强度。三者合起来,把耦合的多模态偏好学习转化成一组单模态子问题,奖励归因更精确,跨模态同步也随之改善。\n\n## 8 张 A800 就能训 22B 模型\n\n工程侧的数字相当务实:AV-GRPO 支持对 22B 参数的 LTX-2.3 音视频模型做全参数训练或 LoRA 训练,门槛是 8 张 A800 GPU。仓库把训练流程拆成清晰步骤:先安装计算音视频样本奖励所需的评估器,所需预训练模型自动下载;再拉取 LTX-2.3 相关权重、替换配置文件路径、配置 WandB 日志;训练数据集直接内置在仓库的 dataset.json 里,无需单独下载。训完的权重合并与推理脚本也已备好,权重挂在 Hugging Face 上。\n\n## 结果与该泼的冷水\n\n在 JavisBench 和 VABench 两个评测上,团队报告 AV-GRPO 在 LoRA 与全参微调两种设置下,生成质量、语义对齐和跨模态同步均超过基座 LTX-2.3。仓库还给出一组定性对比示例,覆盖起火建筑、乐团小提琴、瀑布轰鸣等场景,其中包含中文对白的镜头,能看出框架对语音与画面绑定的处理。需要冷静看待的有三点:\"首个面向音视频联合生成的 GRPO 框架\"是团队自述,论文用的是\"据我们所知\"的措辞;第三方复现尚未出现,对比示例也是作者自选 case 而非盲测;仓库许可需按各子模块(JavisDiT、LTX 等)条款分别确认,并非单一开源协议。\n\n对做生成模型后训练的人来说,这篇的价值不在跑分而在方法论:当多模态奖励信号互相纠缠时,先解耦再优化,配合冻结塔控制成本,是一条可迁移到其他多模态生成任务的路径。8 卡训 22B 的配置,也把扩散 RL 后训练的入场券压到了学术实验室的水平。下一步值得盯的是独立复现与中文语音场景的第三方评测——在那之前,把它当一个方向信号,而不是结论。\n\n参考:论文 arxiv.org\u002Fabs\u002F2609.29816;代码与数据 github.com\u002Fzhiyuxu03\u002FAV-GRPO;权重 huggingface.co\u002FDr-Loser\u002FAV-GRPO。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.29816","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"2236b532-bc9e-4a72-9072-cda2fbda7755","en","AV-GRPO: Diffusion RL for Joint Audio-Video Generation","Shanghai AI Lab's AV-GRPO uses modality-anchored diffusion RL for joint audio-video generation, beating its 22B LTX-2.3 base on sync with 8 A800 GPUs.","Joint audio-video generation has improved fast in recent years, yet three old problems persist: limited per-modality fidelity, insufficient text alignment, and weak cross-modal synchronization. Reinforcement-learning post-training has proven itself repeatedly on text models, but porting it to joint audio-video generation is hard. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment; jointly optimizing two modality towers is computationally expensive given their divergent dynamics; and judging synchronization quality depends on paired samples, which prevents fair reward comparisons.\n\nA team from the Shanghai Artificial Intelligence Laboratory attacks this by decoupling. Their AV-GRPO is a modality-anchored online diffusion RL framework, released together with 5DAV, a decoupled, difficulty-controllable training dataset that separates samples across five dimensions. The paper is public on arXiv (2609.29816, crossing cs.CV and cs.SD, 22 pages), with code, data, and weights all open.\n\n## Turning multimodal preference learning into unimodal subproblems\n\nDecoupling is the core idea. The framework has three key modules. First, modality-anchored rollouts disentangle learning signals and stabilize difficulty. Second, trajectory-locked frozen-tower optimization freezes one modality tower while training the other, cutting cost and reassigning credit. Third, adaptive objectives and perturbation strengths are tailored to each modality's own dynamics. Together they convert coupled multimodal preference learning into a set of unimodal subproblems, enabling precise reward attribution and better synchronization.\n\n## 22B parameters on 8 A800 GPUs\n\nThe engineering numbers are pragmatic: AV-GRPO supports full-parameter or LoRA training of the 22B LTX-2.3 audio-video model on just 8 A800 GPUs. The repository breaks training into clear steps — install the evaluators that compute audio-video sample rewards, with pretrained models downloaded automatically; fetch the LTX-2.3 weights; replace paths in the training config; set up WandB logging at the marked lines of the trainer. The training dataset ships inside the repo as dataset.json, so no separate download is needed. Post-training weight merging and inference scripts are provided, and the weights are hosted on Hugging Face.\n\n## Results, and the cold water\n\nOn JavisBench and VABench, the team reports that AV-GRPO outperforms the LTX-2.3 base in generation quality, semantic alignment, and cross-modal synchronization, under both LoRA and full fine-tuning. The repo also offers qualitative comparisons covering a burning building, an orchestra violinist, and a thundering waterfall, including shots with Chinese dialogue — a hint at how the framework binds speech to picture. Three caveats deserve mention: the claim of being the first GRPO framework for joint audio-video generation is the team's own wording (\"to our knowledge\"); no third-party replication exists yet, and the comparison cases are author-selected rather than blind-tested; and licensing follows each submodule (JavisDiT, LTX, etc.) rather than a single open license.\n\nFor anyone doing post-training on generative models, the value here is methodological rather than leaderboard-based: when multimodal reward signals entangle, decouple first, then optimize, using frozen towers to control cost — a path that transfers to other multimodal generation tasks. The 8-GPU, 22B setup pushes the entry ticket for diffusion RL post-training down to academic-lab scale. What to watch next is independent replication and third-party evaluation on Chinese speech scenarios. Until then, treat this as a directional signal, not a conclusion.\n\nReferences: paper at arxiv.org\u002Fabs\u002F2609.29816; code and data at github.com\u002Fzhiyuxu03\u002FAV-GRPO; weights at huggingface.co\u002FDr-Loser\u002FAV-GRPO.","av-grpo-audio-video-diffusion-rl","2026-09-27T17:08:13Z","2026-09-27T17:08:18.864183Z","2026-09-27T17:08:18.864194Z",true,"agent",88,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"9210c86b-3e1d-434b-bbfd-78e62c698aed","WorldCrafter 开源:给视频世界模型装上可查询的 3D 记忆,转一圈回来还是那个房间","worldcrafter-video-world-model-3d-memory","2026-09-22T19:08:48+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"fb2da954-b12e-4bde-9146-b61dd240df92","SolarWM 开源:143 万条视频喂出的世界模型,5 秒训练片段撑起小时级交互","solarwm-open-data-video-world-models","2026-09-03T15:08:13+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"ba0ed7bf-3de3-4f92-98fe-a50d6ac274d0","WROP 开源:用 150 个物体恒存任务给世界模型补认知课","wrop-object-permanence-world-models","2026-09-25T17:08:02+00:00"]