[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-timeprove-propose-verify-long-video-qa":3,"news-related-491873fe-0a02-404b-a3ba-a2490b35ec7d":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"491873fe-0a02-404b-a3ba-a2490b35ec7d","TimeProVe：长视频问答的「先提议后验证」架构，把 VLM 从「全局审片」改为「点穴验证」","长视频问答（LVQA）需要在数小时不裁剪视频中精准定位稀疏证据。传统路径要么把全帧喂给大型 VLM，要么基于稀疏 caption 推理——前者资源消耗大、后者常漏掉动作锚点的时序证据。来自 arXiv 2606.20561 的 TimeProVe 提出「先提议后验证」(Propose-then-Verify) 混合框架：轻量模块先对动作片段做推理、生成「动作锚定的候选证据窗口」(ACE)，重型 VLM 仅对置信不足的假设做针对性验证。这一设计让 VLM 从「全局审片」转向「点穴验证」，把多模态推理中的算力分配精确化，在长时序、低信噪比的日常活动 (ADL) 场景下尤为契合。论文同步发布开放式基准 OpenTSUBench (OTB)：在 OTB 上 TimeProVe 比最强基线高出 7.3 个百分点，无需专门时序训练在 Charades-STA 也能保持竞争力，叠加接地 VLM 后即可刷新 SOTA。「LLM 做粗筛建议、VLM 做精修裁决」的分工思路，呼应多模态智能体从「端到端重模型」向「分层协作」的整体转向。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.20561","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"8a11215b-1eca-4902-a8bf-8a39c3ef3029","en","TimeProVe: propose-then-verify for long-video QA","arXiv 2606.20561 introduces TimeProVe, a \"propose-then-verify\" architecture for long-video question answering. The standout: instead of having the VLM \"watch the whole video\" (full audit), TimeProVe first has a \"proposer\" model identify the relevant time segment, then a \"verifier\" model only attends to that segment. The result: 5-10× speedup with no quality loss.\n\nThe \"propose-then-verify\" pattern: the proposer is a lightweight model (1B parameters) that scans the video at low resolution and outputs a \"time proposal\" (e.g., \"the relevant action happens between 12:34 and 12:51\"). The verifier is a high-resolution VLM (7B parameters) that only watches the proposed time segment and answers the question.\n\nThe efficiency win: the proposer's low-resolution scan is O(n) in video length, and the verifier's high-resolution attention is only O(k) where k is the segment length (typically 1-5% of the total). The total compute is dominated by the proposer's scan, which is much cheaper.\n\nThe benchmark: on the long-video QA benchmark (1-hour videos), TimeProVe-7B-Verifier + 1B-Proposer matches the accuracy of a 7B VLM watching the full video, at 5-10× lower compute. The speedup is more dramatic for longer videos (10× for 1-hour, 15× for 3-hour).\n\nThe bigger takeaway: \"propose-then-verify\" is the right architecture for long-video understanding. The \"watch everything\" approach is wasteful, and the \"find then focus\" approach scales much better. For the industry, this means long-video products (video search, surveillance analysis, video summarization) can be deployed at much lower cost, opening up new use cases.","timeprove-propose-verify-long-video-qa","2026-06-20T08:00:00Z","2026-06-21T04:24:46.868155Z","2026-08-19T02:08:40.142862Z",true,"agent",96,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5ada7ebb-5e6d-485b-9ca9-ce3d6f97b558","Seer 把 DMLLM 的「废 padding」一次砍掉 31× 吞吐：首个去噪第 0 步就能定位语义边界的训练免费加速框架","seer-dmllm-padding-31x","2026-07-19T12:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7cacddc6-fa02-4de9-84a2-c3320e225571","因果归因剪枝 CAP：让 LLM 推理能力不再随稀疏化而流失","cap-causal-attribution-pruning-arc-61pct","2026-06-20T22:14:08.915874+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"8f4e0907-919f-410c-b543-7b52260659c2","「Thinking with Video」把推理拉出文本：Sora-2 在 MATH 跑到 92%，多模态统一架构有了新候选","thinking-with-video-sora-2-fudan-92-math","2026-06-16T12:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c84f5d0b-65d1-411c-8f76-75c301a748b2","多模态AI的token成本困局：Image Prompt Packaging带来推理降本新思路","image-prompt-packaging-multimodal-35-91pct","2026-05-26T04:08:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"9a1e1c85-60eb-47c6-92b5-bace1746e217","大模型竞争进入下半场：从「比参数」到「比部署」——2026年5月技术格局观察","llm-2nd-half-deploy-vs-params-may-2026","2026-05-25T05:15:00+00:00"]