Joint audio-video generation has improved fast in recent years, yet three old problems persist: limited per-modality fidelity, insufficient text alignment, and weak cross-modal synchronization. Reinforcement-learning post-training has proven itself repeatedly on text models, but porting it to joint audio-video generation is hard. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment; jointly optimizing two modality towers is computationally expensive given their divergent dynamics; and judging synchronization quality depends on paired samples, which prevents fair reward comparisons.
A team from the Shanghai Artificial Intelligence Laboratory attacks this by decoupling. Their AV-GRPO is a modality-anchored online diffusion RL framework, released together with 5DAV, a decoupled, difficulty-controllable training dataset that separates samples across five dimensions. The paper is public on arXiv (2609.29816, crossing cs.CV and cs.SD, 22 pages), with code, data, and weights all open.
Turning multimodal preference learning into unimodal subproblems
Decoupling is the core idea. The framework has three key modules. First, modality-anchored rollouts disentangle learning signals and stabilize difficulty. Second, trajectory-locked frozen-tower optimization freezes one modality tower while training the other, cutting cost and reassigning credit. Third, adaptive objectives and perturbation strengths are tailored to each modality's own dynamics. Together they convert coupled multimodal preference learning into a set of unimodal subproblems, enabling precise reward attribution and better synchronization.
22B parameters on 8 A800 GPUs
The engineering numbers are pragmatic: AV-GRPO supports full-parameter or LoRA training of the 22B LTX-2.3 audio-video model on just 8 A800 GPUs. The repository breaks training into clear steps — install the evaluators that compute audio-video sample rewards, with pretrained models downloaded automatically; fetch the LTX-2.3 weights; replace paths in the training config; set up WandB logging at the marked lines of the trainer. The training dataset ships inside the repo as dataset.json, so no separate download is needed. Post-training weight merging and inference scripts are provided, and the weights are hosted on Hugging Face.
Results, and the cold water
On JavisBench and VABench, the team reports that AV-GRPO outperforms the LTX-2.3 base in generation quality, semantic alignment, and cross-modal synchronization, under both LoRA and full fine-tuning. The repo also offers qualitative comparisons covering a burning building, an orchestra violinist, and a thundering waterfall, including shots with Chinese dialogue — a hint at how the framework binds speech to picture. Three caveats deserve mention: the claim of being the first GRPO framework for joint audio-video generation is the team's own wording ("to our knowledge"); no third-party replication exists yet, and the comparison cases are author-selected rather than blind-tested; and licensing follows each submodule (JavisDiT, LTX, etc.) rather than a single open license.
For anyone doing post-training on generative models, the value here is methodological rather than leaderboard-based: when multimodal reward signals entangle, decouple first, then optimize, using frozen towers to control cost — a path that transfers to other multimodal generation tasks. The 8-GPU, 22B setup pushes the entry ticket for diffusion RL post-training down to academic-lab scale. What to watch next is independent replication and third-party evaluation on Chinese speech scenarios. Until then, treat this as a directional signal, not a conclusion.
References: paper at arxiv.org/abs/2609.29816; code and data at github.com/zhiyuxu03/AV-GRPO; weights at huggingface.co/Dr-Loser/AV-GRPO.