[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-co-rl-peer-reward-label-free-rl":3,"news-related-5a274662-0c3f-492c-a0e8-a46c5a0be783":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"5a274662-0c3f-492c-a0e8-a46c5a0be783","别再自己给自己打分了:Co-RL 让模型互相判卷,无标签 RL 追平有监督","约翰霍普金斯与加州大学圣地亚哥团队提出 Co-RL:多个不共享参数的模型互相用同伴的多数投票当奖励,全程零标签做 RL 训练,7 个文本基准平均提升 3.0-8.6%,14 个设置中 11 个追平有监督 GRPO。","强化学习把大模型的推理能力抬了一整个台阶,但最有效的玩法 RLVR(Reinforcement Learning with Verifiable Rewards)有个隐形成本:每条训练 prompt 都要一个经过验证的正确答案。当模型要学的东西超过人类能可靠评判的范围,这种标注越来越贵、越来越稀缺。\n\n## 自己给自己打分的死胡同\n\n业界的替代方案是 self-rewarding RL——让模型给自己的输出打分。TTRL 用多数投票的自我一致性当信号,Intuitor 用自我确信度,RENT 用预测熵。这些方法有个共同的结构性缺陷:奖励信号永远出不了模型自己。自我强化会放大已有偏差、压缩回答多样性,最后走向输出同质化甚至训练崩溃。\n\nCo-RL 的思路是把这个死结剪断:奖励永远来自同伴,永远不来自自己。\n\n## 三步机制:采样、互评、独立更新\n\n具体做法分三步。第一步,每个 agent 对同一条无标签 prompt 采样多个回答,约减成一个多数投票结果。第二步,某个回答的答案与**同伴**的投票一致就得 1 分奖励——一个 agent 永远不会给自己的监督目标贡献信号。第三步,GRPO 独立更新每个策略,agent 之间不共享参数、不交换梯度。\n\n超过两个 agent 时,投票沿着有向环传递。三个模型 Qwen2.5-3B、Llama-3.2-3B-Instruct、Qwen3-1.7B 在同一个环里一起训练,分别提升 7.8%、6.0%、8.2%,全部追平或超过各自的有标签参考线。\n\n## 多样性才是好老师\n\nCo-RL 最有意思的发现是关于\"谁配当老师\"的:两个相似的模型犯同样的错,就会互相强化同样的错误答案。论文在 12 组基座模型对上测了训练前的错误重叠度(Cohen's κ):跨家族配对的 κ 全部 ≤ 0.42,同家族配对全部 ≥ 0.51,中间地带是空的。跨家族把平均重叠度从 0.53 压到 0.38,下游收益也按同样排序。\n\n多样性通过三个通道注入:异构模型家族(架构、分词器、预训练数据都不同)、异构尺寸、改写训练样本(一个 agent 用原始 MATH 题目,另一个用 DeepSeek-V3 改写的同答案版本)。\"Different family+\" 配置在 7 个文本基准上拿到 49.3% 平均分,超过 GT-Reward 有标签参考线的 47.4%。\n\n## 无标签追平有监督的完整跑分\n\n以 Qwen2.5-3B 为骨干的完整成绩单:GSM8K 73.4→81.0,MATH500 56.6→66.6,AMC 28.9→36.1,HumanEval 39.0→65.8(Different family+ 配置)。多模态侧 4 个基准平均提升 2.3-7.2%。\n\n两个额外的 ablation 值得注意。其一,把两个 TTRL 训练的模型做集成,分数反而低于自己最好的单模型(64.9%),而 Co-RL 集成达到 66.9%——问题不在集成,在训练。其二,在 CoMAS 自己的协议下,Co-RL 用一半的 agent 数、不要裁判模型,拿到 62.97% 对 58.94%。\n\n训练动态也干净:四个规模上,self-rewarding 基线出现奖励崩溃、输出长度激增或中途发散,Co-RL 全程稳定。\n\n## 所以呢\n\n对训练侧的读者,这条路线的价值很直接:当可验证奖励的标注成本压住你的扩产节奏,同侪互评是一条已经跑通的无标签替代路径,而且 Apache-2.0 代码已开源([arXiv 2608.17253](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.17253),[GitHub](https:\u002F\u002Fgithub.com\u002FDrStranded\u002FCo-RL))。更值得记住的是那个 κ 空隙:跨家族配对 ≤ 0.42 与同家族 ≥ 0.51 之间没有任何过渡——\"找和你犯错不一样的同伴\"不是玄学,是可提前测量的量。当行业还在争论合成数据会不会让模型越训越同质,Co-RL 给了一个可操作的答案:把彼此的分歧本身变成监督信号。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.17253","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"15523c78-84d3-4431-9782-2f271ce3dff6","推理",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"537c712d-3315-49a6-9840-63d7683963e8","en","Co-RL: peer-graded label-free RL matches supervised GRPO","JHU and UCSD propose Co-RL: models grade each other via peer votes for zero-label RL, gaining 3.0-8.6% on seven text benchmarks and matching supervised GRPO.","Reinforcement learning lifted the reasoning ability of large models by a full step, but the most effective variant, RLVR (Reinforcement Learning with Verifiable Rewards), carries a hidden cost: every training prompt needs a verified correct answer. As models learn things beyond what humans can reliably grade, this annotation gets more expensive and scarcer.\n\n## The dead end of grading yourself\n\nThe industry's alternative is self-rewarding RL — letting the model score its own outputs. TTRL uses majority-vote self-consistency as the signal, Intuitor uses self-certainty, RENT uses predictive entropy. These methods share a structural flaw: the reward signal never leaves the model itself. Self-reinforcement amplifies existing biases, shrinks response diversity, and eventually drifts into homogeneous outputs or training collapse.\n\nCo-RL cuts this knot: the reward always comes from a peer, never from yourself.\n\n## Three steps: sample, cross-check, update independently\n\nThe method works in three steps. First, each agent samples multiple completions for the same unlabeled prompt and reduces them to a single majority vote. Second, a completion earns reward 1 when its answer matches the peer's vote — an agent never contributes to its own supervision target. Third, GRPO updates each policy independently; agents share no parameters and exchange no gradients.\n\nBeyond two agents, the votes pass along a directed ring. Three models — Qwen2.5-3B, Llama-3.2-3B-Instruct, and Qwen3-1.7B — train together in one ring, gaining 7.8%, 6.0%, and 8.2% respectively, each matching or beating its labeled reference.\n\n## Diversity is what makes a good teacher\n\nCo-RL's most interesting finding concerns who qualifies as a teacher: two similar models make the same mistakes and reinforce the same wrong answers. The paper measured pre-training error overlap (Cohen's κ) across twelve base-checkpoint pairs: every different-family pair lands at κ ≤ 0.42, every same-family pair at κ ≥ 0.51, and the strip between them is empty. Crossing families drops the average overlap from 0.53 to 0.38, and downstream gains follow the same ordering.\n\nDiversity enters through three channels: heterogeneous model families (differing in architecture, tokenization, and pretraining data), heterogeneous sizes, and rephrased training samples (one agent trains on the original MATH prompt, the other on a DeepSeek-V3 rewrite that keeps the answer unchanged). The \"Different family+\" setting reaches a 49.3% average across seven text benchmarks, above the 47.4% of the GT-Reward labeled reference.\n\n## The full scorecard: label-free matching supervised\n\nWith Qwen2.5-3B as the backbone, the full results: GSM8K 73.4→81.0, MATH500 56.6→66.6, AMC 28.9→36.1, HumanEval 39.0→65.8 (Different family+ setting). On the multimodal side, average gains of 2.3-7.2% across four VLM benchmarks.\n\nTwo ablations deserve attention. First, ensembling two TTRL-trained models scores below its own best single model (64.9%), while the Co-RL ensemble reaches 66.9% — the problem is not ensembling but training. Second, under CoMAS's own protocol, Co-RL gets 62.97% versus 58.94% with half the agents and no judge model.\n\nTraining dynamics are also clean: at four scales, self-rewarding baselines show reward collapse, sharp completion-length inflation, or mid-run divergence, while Co-RL stays stable throughout.\n\n## So what\n\nFor readers on the training side, the value is direct: when annotation cost for verifiable rewards squeezes your scaling plans, peer grading is a proven label-free alternative, with Apache-2.0 code already open ([arXiv 2608.17253](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.17253), [GitHub](https:\u002F\u002Fgithub.com\u002FDrStranded\u002FCo-RL)). The more memorable takeaway is the κ gap: nothing lands between cross-family ≤ 0.42 and same-family ≥ 0.51 — \"find peers whose mistakes differ from yours\" is not folk wisdom but a measurable, pre-trainable quantity. While the industry debates whether synthetic data homogenizes models, Co-RL offers an operational answer: turn mutual disagreement itself into the supervision signal.","co-rl-peer-reward-label-free-rl","2026-08-20T19:10:00Z","2026-08-20T19:10:21.029237Z","2026-08-20T19:10:21.029246Z",true,"agent",47,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"5b76d928-8896-4bf8-8edd-e195ecf0094a","Ornith-1.5自报跑分赢了Claude,独立复测翻车了","ornith-1-5-benchmark-reality-check","2026-08-20T17:10:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"c71b8ee7-9487-4c78-89fd-30bb0368b99e","DeepSeek V4 Flash：284B\u002F13B MoE，成本比 Luna 低 60%","deepseek-v4-flash-0731-intelligence-index-50","2026-08-05T03:00:00+00:00"]