[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-omnireasoning-audio-visual-joint-reasoning":3,"topics-all":38,"news-related-3701d19e-517a-4ff0-a95d-6d768cfeb04e":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"3701d19e-517a-4ff0-a95d-6d768cfeb04e","OmniReasoning：音视频联合推理仍是不及格难题","Qwen 团队发布 OmniReasoning：音视频证据缺一不可的 1,150 题基准、带时间戳线索链的 OmniQA 数据引擎、模态分解自蒸馏方法 MFSD，专补联合推理评测空白。基于 Qwen3-Omni-30B-A3B 训练的模型在两个基准提升 12.8\u002F9.3 个百分点，但绝对分仍只有 42.5%。","把视频静音，你能答对几成？反过来，只听声音不看画面呢？对人类这是两道送分题，对当下号称「全模态」的大模型，却是一片测试盲区——多数基准把音频、视觉拆开单独考，音视频必须一起用才能解的题，长期没有专门的评测和训练方法。Qwen 团队 9 月 30 日提交到 arXiv 的 OmniReasoning，就是冲着这块空白去的。\n\n## 基准：音视频证据缺一不可\n\n论文的第一件事是把「联合」变成硬约束。OmniReasoningBench 收录 1,150 道选择题与开放题，分两大任务：一类是「对视频内推理」，另一类是「超越视频推理」。题目设计的底线是：音频和视觉证据缺一不可，单看任何一边都答不出来。这直接戳破了「分开考再相加」的传统测法的漏洞。\n\n## 数据引擎：带时间戳的线索链\n\n光有基准不够，还得会造训练数据。团队构建了 OmniQA 数据引擎，自动生成「证据落地」的问答对——每道题都显式依赖音视频联合推理，并附带带时间戳的线索链，用来引导思维过程的标注。引擎产出两套数据：OmniReasoning-SFT-112K（监督微调）和 OmniReasoning-RL-19K（强化学习）。\n\n## MFSD：把功劳算到每个模态头上\n\n方法层的核心叫 Modality-Factored Self-Distillation（MFSD），一种在策略自蒸馏：对每个采样回复，分别在「单模态线索上下文」下评估，把单个线索的贡献和跨模态交互拆开，做 token 级的功劳分配。直白说：模型答对了，得说清是听出来的还是看出来的。\n\n## 42.5%：涨得猛，但仍在深水区\n\n结果摆在一起看：基于 Qwen3-Omni-30B-A3B-Thinking 训练的 OmniReasoning-30B-A3B，在 OmniVideoBench 拿到 50.0%，在 OmniReasoningBench 拿到 42.5%，较基座分别提升 12.8 和 9.3 个百分点，在 Video-MME-v2 等通用长视频基准上也有增益。但换个角度：自家针对性训练后的模型，在自家基准上也只答对不到一半——音视频联合推理不是「多做点数据」就能通关的赛道，线索交织时的功劳分配仍是硬骨头。\n\n这篇论文对普通开发者的启发很直接：做多模态应用时，别急着堆参数，先检查你的评测是不是把模态拆开考了——拆开考的高分，在真实音视频场景里可能一文不值。\n\n参考：arXiv:2609.39490（[摘要页](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39490)）；Hugging Face Daily Papers（[论文页](https:\u002F\u002Fhuggingface.co\u002Fpapers\u002F2609.39490)，10 月 6 日作者自荐上榜，12 个 upvote）。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39490","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"8c469bb6-b28d-4850-b258-d8f43b9ac3e6","en","OmniReasoning: Qwen Team Benchmarks Audio-Visual Reasoning","Qwen's OmniReasoning: 1,150 audio-visual questions, OmniQA engine, MFSD distillation. The 30B model gains 12.8\u002F9.3 points yet scores 42.5%.","Mute the video and you lose the plot; listen without watching and you miss half the story. Humans handle both easily, but today's omni-modal models are rarely tested on questions where audio and visual evidence must be used together — existing benchmarks mostly score each modality separately. OmniReasoning, submitted to arXiv on Sep 30 by the Qwen team, targets exactly that gap.\n\n## A benchmark where neither modality can be dropped\n\nOmniReasoningBench packs 1,150 multiple-choice and open-ended questions across two tasks: reasoning over video and reasoning beyond video. The design floor is strict: both audio and visual evidence are indispensable, so single-modality shortcuts simply fail.\n\n## The OmniQA data engine\n\nA benchmark alone does not fix training. The team built OmniQA, an engine that automatically constructs evidence-grounded QA pairs explicitly requiring joint audio-visual reasoning, each with time-stamped clue chains that guide the annotation of the thinking process. It yields two datasets: OmniReasoning-SFT-112K and OmniReasoning-RL-19K.\n\n## MFSD: credit assignment per modality\n\nThe learning method, Modality-Factored Self-Distillation (MFSD), is an on-policy self-distillation scheme that evaluates each sampled response under modality-specific clue contexts — disentangling what a single clue contributed from how clues interacted across modalities, down to token-level credit assignment.\n\n## 42.5%: big gains, still deep water\n\nBuilt on Qwen3-Omni-30B-A3B-Thinking, the resulting OmniReasoning-30B-A3B scores 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench — up 12.8 and 9.3 percentage points over the base model, with further gains on general and long-video benchmarks including Video-MME-v2. Yet the flip side is blunt: even purpose-trained on its own benchmark, the model answers fewer than half the questions correctly. Joint audio-visual reasoning is not a lane you clear by pouring in more data.\n\nThe takeaway for builders: before scaling parameters, check whether your eval splits modalities apart — a high score on split evals may be worthless in real audio-visual settings.\n\nReferences: [arXiv:2609.39490](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39490); [HF Daily Papers page](https:\u002F\u002Fhuggingface.co\u002Fpapers\u002F2609.39490) (self-submitted Oct 6, 12 upvotes).","omnireasoning-audio-visual-joint-reasoning","2026-10-07T17:07:02Z","2026-10-07T17:07:09.746187Z","2026-10-07T17:07:09.746197Z",true,"agent",139,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"206ea36a-1eea-463f-a241-1e3b32f5ec2d","USTC GraphForge:证据图把任务和 rubric 钉在一起,Qwen3.6-27B 涨三基准","graphforge-ustc-qwen36-27b-evidence-graph","2026-10-04T03:05:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"5d3c50e8-5087-43e9-a8f1-c64973f712c1","Qwen 拆掉 ASR 管道:音视频原生对话靠合成数据练成","qwen-omnivchat-native-audio-visual-dialogue","2026-09-21T15:14:15+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"d74a088d-e7f5-41cc-8c55-aedb10b8101d","Gemini-3-Pro 也只拿 66.4 分:南京大学开源全模态视频助手基准 OmniAssistBench","omniassistbench-omni-llm-video-assistant","2026-08-21T17:59:52+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"d7b6d14d-7257-4794-b92f-31956bbc7eae","原生多模态 vs 后训练加压:国产头部基模两条路线的工程账","native-multimodal-vs-posttraining-2026","2026-08-05T00:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"4b693fb4-541f-47ed-8923-6e280cec965f","大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了","memlens-multimodal-long-term-memory","2026-08-03T02:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"7258978b-dfcd-4cb4-91c4-3b8569cd5deb","Qwen-Audio-3.0-TTS双版本发布:Plus登顶Artificial Analysis,Flash压到300ms首包延时","qwen-audio-3-tts","2026-07-20T10:00:00+00:00"]