Mute the video and you lose the plot; listen without watching and you miss half the story. Humans handle both easily, but today's omni-modal models are rarely tested on questions where audio and visual evidence must be used together — existing benchmarks mostly score each modality separately. OmniReasoning, submitted to arXiv on Sep 30 by the Qwen team, targets exactly that gap.
A benchmark where neither modality can be dropped
OmniReasoningBench packs 1,150 multiple-choice and open-ended questions across two tasks: reasoning over video and reasoning beyond video. The design floor is strict: both audio and visual evidence are indispensable, so single-modality shortcuts simply fail.
The OmniQA data engine
A benchmark alone does not fix training. The team built OmniQA, an engine that automatically constructs evidence-grounded QA pairs explicitly requiring joint audio-visual reasoning, each with time-stamped clue chains that guide the annotation of the thinking process. It yields two datasets: OmniReasoning-SFT-112K and OmniReasoning-RL-19K.
MFSD: credit assignment per modality
The learning method, Modality-Factored Self-Distillation (MFSD), is an on-policy self-distillation scheme that evaluates each sampled response under modality-specific clue contexts — disentangling what a single clue contributed from how clues interacted across modalities, down to token-level credit assignment.
42.5%: big gains, still deep water
Built on Qwen3-Omni-30B-A3B-Thinking, the resulting OmniReasoning-30B-A3B scores 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench — up 12.8 and 9.3 percentage points over the base model, with further gains on general and long-video benchmarks including Video-MME-v2. Yet the flip side is blunt: even purpose-trained on its own benchmark, the model answers fewer than half the questions correctly. Joint audio-visual reasoning is not a lane you clear by pouring in more data.
The takeaway for builders: before scaling parameters, check whether your eval splits modalities apart — a high score on split evals may be worthless in real audio-visual settings.
References: arXiv:2609.39490; HF Daily Papers page (self-submitted Oct 6, 12 upvotes).