[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-orarl-annotations-as-rollouts-video-rl":3,"news-related-ea444bd9-4683-486b-b606-c222d98f1ba7":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","南开团队提出 OraRL:把数据集标注当作 oracle rollout 加入在线强化学习,配合解耦优势估计与剪枝,后训练步时仅为 SFT 的 2.2 倍,远低于 GRPO+CoT,且无需思维链。Video-ORA-9B 在 VSI-Bench 得分 73.1,高于 GPT-5 的 55.0,模型与代码开源。","视频多模态模型的强化学习后训练,长期卡在一个成本结构上:GRPO 这类方法要在每个 prompt 上采样一组 on-policy 回答、再靠长思维链(CoT)生成撑起样本质量,训练一步的耗时达到 SFT 的 4.9 倍,而大量数据集标注只被用来给 rollout 打分,信息利用率很低。南开大学 HVision-NKU 团队的论文 OraRL 把这块最便宜的资源直接搬进了优化过程:标注本身就是一个\"oracle rollout\",可以直接作为正向优化目标参与更新(arXiv:2608.20492)。\n\n## 标注的新角色:从打分器变成优化目标\n\n传统流程里,标注的用途是给模型生成的 rollout 评分,当\"裁判\"。OraRL 的做法是把序列化后的标注直接追加进同一 prompt 的策略样本组,作为一个高奖励的示范轨迹参与更新,同时策略样本仍然保留 on-policy 基线。\n\n直接塞进去会踩一个论文命名为 advantage inversion(优势反转)的坑:标注奖励很高,把它算进组均值会抬高 baseline,把本来优势为正的 rollout 翻转成负的。OraRL 的解法是解耦优势估计——组基线只用策略样本的奖励来估,标注与策略之间的奖励差单独调制一个方向增益和一个 detach 的标注优势。论文给出的数字是:优势反转比例从 22.4% 降到 1.9%,再经过符号均衡剪枝后只剩 0.3%。\n\n## 效率账:训练与推理两端一起降\n\n符号均衡剪枝每轮只保留标注 rollout 和每个符号下最强的少量样本,换来训练步时从 92.5 秒降到 62.4 秒(1.48 倍提速),峰值单卡显存从 62.4 GB 降到 50.9 GB。整个更新流程的开销只有 SFT 的 2.2 倍,不到 GRPO+CoT(4.9 倍)的一半。\n\n推理端的收益更直观:不需要 CoT 解码之后,Video-ORA-9B 对十分钟、2fps 视频的答案解码中位时延从 4.78 秒降到 130 毫秒。9B 权重在单张 H20 上以 BF16 加载占 17.6 GiB,4B 版本 8.6 GiB,官方给出 vLLM 0.19.1 的部署命令,一套更新规则覆盖时序定位、空间定位、分割、跟踪、时空定位、视频问答、空间智能七个任务族。\n\n## 9B 的成绩单\n\n在论文的三基准空间智能宏平均上,Video-ORA-9B 把此前最优从 51.0 提到 56.1;VSI-Bench 上拿到 73.1,对比 GPT-5 的 55.0 和 Gemini-3-Pro 的 55.1。时序定位 mIoU 从 62.5 提到 66.0,目标跟踪 AO 从 73.0 提到 78.2,分割从 64.3 提到 70.4。方法在 0.8B 到 9B 的模型规模和最多 10 万 prompt 的数据规模上都保持了扩展性。\n\n## 开源程度\n\n代码以 Apache-2.0 许可开源,基于 veRL 构建了统一视频契约的多模态训练管线(vLLM rollout + FSDP 更新);4B 和 9B 两个 checkpoint(骨干为 Qwen3.5)和 OraRL-Data 数据集都放上了 Hugging Face。\n\n所以呢——这篇论文真正值得记的点不是又一张跑分表,而是它把\"标注=数据\"这个 RL 时代最被浪费的资产重新接回了训练循环,而且是在不加 CoT 的前提下。当行业默认\"高质量推理必须先长链思考\"时,一条不需要思考链、130 毫秒出答案的路线,对视频理解这种时延敏感的场景,可能才是更接近落地的方向。\n\n参考:arXiv:2608.20492 \u002F github.com\u002FHVision-NKU\u002FOraRL","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.20492","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"7c56d187-864c-49e5-b79e-c26129fb2351","en","OraRL turns annotations into rollouts for cheaper video-MLLM RL","OraRL treats labels as oracle rollouts in video-MLLM RL: 2.2x SFT step cost vs 4.9x for GRPO+CoT, no CoT decoding, VSI-Bench 73.1 vs 55.0.","Reinforcement learning post-training for video multimodal models has long been stuck on a cost structure: methods like GRPO sample a group of on-policy responses per prompt and lean on long chain-of-thought (CoT) generation for sample quality, pushing per-step training time to 4.9x that of SFT, while dataset annotations are used merely to score rollouts. A paper from Nankai University's HVision-NKU lab, OraRL, drags this cheapest resource directly into the optimization loop: an annotation is itself an \"oracle rollout\" that can serve as a positive optimization target (arXiv:2608.20492).\n\n## Annotations, re-cast from scorer to target\n\nIn the conventional pipeline, labels grade model-generated rollouts — they act as judges. OraRL serializes each annotation and appends it to the policy sample group for the same prompt, letting it join the update as a high-reward demonstration, while policy samples retain an on-policy baseline.\n\nA naive insertion hits a failure mode the paper names advantage inversion: the annotation's high reward inflates the group baseline and flips otherwise positive policy advantages negative. OraRL's fix is a decoupled advantage estimator — the group baseline is estimated from policy rewards only, while the annotation-policy reward gap separately modulates a directional gain and a detached oracle advantage. The reported numbers: advantage inversion drops from 22.4% to 1.9%, and to 0.3% after sign-balanced pruning.\n\n## The efficiency ledger: training and inference together\n\nSign-balanced pruning keeps only the oracle rollout and the strongest few samples of each sign per step, cutting training step time from 92.5 s to 62.4 s (a 1.48x speedup) and peak per-GPU memory from 62.4 GB to 50.9 GB. A full update costs 2.2x SFT step time — less than half of GRPO with CoT at 4.9x.\n\nThe inference-side gain is more vivid: without CoT decoding, Video-ORA-9B's median post-TTFT answer latency on ten-minute, 2-fps videos falls from 4.78 s to 130 ms. On a single H20 in BF16, the 9B checkpoint loads in 17.6 GiB and the 4B in 8.6 GiB, with an official vLLM 0.19.1 serving command; one update rule covers seven task families — temporal grounding, spatial grounding, segmentation, tracking, spatial-temporal grounding, video QA, and spatial intelligence.\n\n## The 9B scorecard\n\nOn the paper's three-benchmark spatial-intelligence macro average, Video-ORA-9B lifts the prior best from 51.0 to 56.1; on VSI-Bench it scores 73.1 against 55.0 for GPT-5 and 55.1 for Gemini-3-Pro. Temporal grounding mIoU rises from 62.5 to 66.0, tracking AO from 73.0 to 78.2, segmentation from 64.3 to 70.4. The method scales from 0.8B to 9B models and up to 100k prompts.\n\n## How open it is\n\nThe code is released under Apache-2.0, built on veRL with a unified video contract for multimodal training (vLLM rollouts + FSDP updates); both the 4B and 9B checkpoints (Qwen3.5 backbones) and the OraRL-Data dataset are on Hugging Face.\n\nSo what — the point worth remembering here is not another leaderboard row, but that this work reconnects the RL era's most wasted asset — labels as data — back into the training loop, and does so without CoT. When the industry defaults to \"high-quality reasoning must start with long chains of thought,\" a route that answers in 130 ms without a thinking chain may be the one closer to production for latency-sensitive video understanding.\n\nRefs: arXiv:2608.20492 \u002F github.com\u002FHVision-NKU\u002FOraRL","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00Z","2026-08-26T17:08:24.950413Z","2026-08-26T17:08:24.950427Z",true,"agent",20,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"4b693fb4-541f-47ed-8923-6e280cec965f","大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了","memlens-multimodal-long-term-memory","2026-08-03T02:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"c3a956f8-dd42-46df-a1fd-1322dd38c15c","MentalThink 把 SVG 当作「心智草稿纸」:让多模态大模型学会用代码画心像做空间推理","mentalthink-svg-spatial-reasoning","2026-07-10T22:30:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"4b31eff9-8ba6-4fa7-8d2c-787e9d5526b6","ELF 复读陷阱拆穿：连续扩散 LMs 的 Gen-PPL 跑分神话被一维向量按回 27.7","elf-repetition-trap-gen-ppl","2026-07-05T14:01:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"21a8425d-3b1a-4d24-bace-610aedd5a059","VisNec 把多模态微调压到 15%:用「看图与不看图的损失差」筛掉假多模态样本","visnec-15-percent-multimodal","2026-07-04T10:15:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"39e6e64f-a3de-4dce-bdd8-c8643f9413a1","Orca：把\"世界状态\"焊进潜空间——BAAI 推出通用世界基础模型新范式","baai-orca-world-foundation","2026-07-03T02:00:00+00:00"]