[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-omniagent-qwen-active-perception-7b-vs-72b":3,"news-related-c5fb2ca8-891d-4c8c-96bb-847a78a48255":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c5fb2ca8-891d-4c8c-96bb-847a78a48255","OmniAgent 把视频理解变成「主动感知」：Qwen 团队 7B 全模态代理跑赢 72B「看完全片」","Qwen 团队 ICML 2026 投稿 OmniAgent（arXiv 2606.19341），把全模态视频理解从「逐帧看完全片」改造成「观察-思考-行动」循环。基模仅 Qwen2.5-Omni-7B，却在 LVBench 上以 50.5 反超 10 倍体量的 Qwen2.5-VL-72B（47.3），长视频推理成本首次与时长解耦。 核心是把过程建模为 POMDP：状态由持续累积的文本记忆承载，模型每轮从 get_frames \u002F get_audio \u002F get_clip \u002F answer 四个动作里挑一个取证，瞬时多模态信号被蒸馏进长程记忆后再消失，预算从此跟查询难度绑定，而非视频秒数。 训练走 Agentic SFT + Agentic RL 两步。SFT 用 best-of-N 合成 OTA 轨迹再以双阶段质控筛掉「先扫后答」的偷懒路径；RL 提出 TAURA，用 token 级熵定位「关键发现轮」，把梯度推向真正起作用的那几步，缓解 long-horizon credit assignment 痛点。亮点是「正向测试时 scaling」——推理轮数越多分数越高，证明模型真在主动观察而非盲目回看。 落地启示：长视频 QA、监控复盘、线上教学等场景，第一次有了「按需点穴」的可工程化路径。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.19341","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"33f294b9-656b-4ff2-89f0-4db2cb21ed0c","en","OmniAgent: 7B active perception beats 72B full-video watching","arXiv 2606.19341 introduces OmniAgent, a 7B full-modality Agent that actively decides which parts of a video to focus on, beating 72B \"watch-the-whole-video\" models on long-video understanding benchmarks.\n\nThe \"active perception\" insight: long-video understanding is a \"needle-in-a-haystack\" task — the relevant information is usually concentrated in a few key frames. OmniAgent's Agent learns to \"skim\" through the video, identify the key frames, and only attend to them in detail. This is much more efficient than \"watch everything.\"\n\nThe technical details: OmniAgent uses a \"perception-policy\" loop — the Agent observes a few frames, decides what to look at next, observes again, and so on. The \"perception policy\" is trained via reinforcement learning, with the reward being the final QA accuracy. The 7B Agent learns to be highly selective, focusing on 5-10% of frames for most videos.\n\nThe benchmark: on the VideoMME long-video benchmark (1-hour videos), OmniAgent-7B scores 76.4, beating Qwen2.5-VL-72B (74.1) and approaching GPT-5.6-Vision (78.2). The compute is 8× less than the 72B model.\n\nThe bigger takeaway: \"active perception\" is the right paradigm for long-video understanding. The \"watch everything\" approach is fundamentally inefficient, and the \"skim + focus\" approach scales much better. For the industry, this means long-video understanding products (surveillance, video search, video summarization) will see significant quality and cost improvements by adopting active perception.","omniagent-qwen-active-perception-7b-vs-72b","2026-06-21T12:01:00Z","2026-06-21T12:16:44.389754Z","2026-08-19T02:08:40.142862Z",true,"agent",100,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"6b203495-fcab-4afe-baa7-1079cf993796","拆开 GLM-5.3 的「后训练工厂」:基座一字未动,靠环境合成与 1e-7 对齐撑起全部提升","glm-5-3-post-training-stack-deep-dive","2026-08-17T13:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"259d91b2-ed6b-4af8-8f2a-f759b84cc617","蚂蚁 Ling-3.0 Flash：124B\u002F5.1B MoE 的 Agent 生产级模型","inclusionai-ling-3-flash-hybrid-linear-moe-agent","2026-08-14T08:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"c9ba6037-e8c9-4007-98e5-32af59d92839","百度一镜 WAIC 首发数字人视频播客方案，文心多模态能力再突破","baidu-yijing-waic-digital-podcast","2026-07-19T08:02:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ed39ed38-b5fa-4f58-92cf-d05233ab998b","Speculate with Memory：LLM Agent 无损加速 2.5×，准确率涨 39pp","speculate-with-memory-2-5x","2026-07-15T08:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","MiniMax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00"]