[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-llm-as-a-verifier-fourth-scaling":3,"news-related-8173a86b-4e5e-429a-8ddf-f98af527b4b5":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8173a86b-4e5e-429a-8ddf-f98af527b4b5","LLM-as-a-Verifier：验证成 LLM 第四 scaling 维度","LLM-as-a-Verifier 把「打分」从一次性 LLM-as-Judge 升级为可规模化的训练目标——Ion Stoica、Chelsea Finn、Azalia Mirhoseini 等 9 位学者 7 月 6 日挂出 arXiv 2607.05391,GitHub 一周斩获 409 stars。\n\n核心洞察不复杂:在 pre-training、post-training、test-time compute 之后,verification(判断「对不对」)才是下一个该被规模化的轴。传统 LLM-as-Judge 吐离散分数,信号太粗没梯度可学。本文换了一套算账方式——对 scoring token 的 logits 求期望得到 continuous score;同时把 verification 拆成三轴:score granularity(分数粒度越细,正负分离度越好)、repeated evaluation(多次评估降方差)、criteria decomposition(把评判标准拆开来打)。三轴一起转,Terminal-Bench V2 86.5%、SWE-Bench Verified 78.2%、RoboRewardBench 87.4%、MedAgentBench 73.3%,跨 4 个毫不相干领域全部 SOTA。\n\n落地姿势更有意思:团队给 Claude Code 做了一版扩展,直接把验证器喂给 agent 进程级反馈,等于在自我纠错环里塞了一块「准确率雷达」;同时把 continuous scores 当 RL 密集回报,SAC、GRPO 在机器人和数学推理上的样本效率肉眼可见地提升。verification 从「打分工具」变成了「训练信号源」。\n\n更深一层的范式:当 scaling law 在预训练端撞到能源和算力墙,业界一直在找下一个可规模化的变量。post-training 和 test-time compute 已被 RLHF、o1 类推理模型验证;verification 接上,意味着 LLM 不再只是「答题者」,开始向「出题-答题-判分」闭环演化。Claude Code 那一段是这条路上目前最工程化的样本。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.05391","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"a379fe22-a113-4a4a-9df2-e586be500434","en","LLM-as-a-Verifier: verification as LLM's fourth scaling axis","LLM-as-a-Verifier upgrades \"scoring\" from one-off LLM-as-Judge to a scalable training objective — 9 scholars including Ion Stoica, Chelsea Finn, and Azalia Mirhoseini posted arXiv 2607.05391 on July 6, and the GitHub repo got 409 stars in a week. The core insight isn't complex: after pre-training, post-training, and test-time compute, verification (judging \"right or wrong\") is the next axis that should be scaled. Traditional LLM-as-Judge outputs discrete scores, signal too coarse to have a gradient to learn. This paper changes the accounting method — taking the expectation of scoring-token logits to get a continuous score; at the same time splitting verification into three axes: score granularity (finer score granularity means better positive\u002Fnegative separation), repeated evaluation (multiple evaluations reduce variance), criteria decomposition (split the judging criteria to score separately). All three axes turning together gives Terminal-Bench V2 86.5%, SWE-Bench Verified 78.2%, RoboRewardBench 87.4%, MedAgentBench 73.3% — SOTA across 4 unrelated domains. The landing posture is even more interesting: the team built an extension for Claude Code, directly feeding the verifier to the agent's process-level feedback, equivalent to stuffing an \"accuracy radar\" into the self-correction loop; at the same time using continuous scores as RL dense reward, the sample efficiency of SAC, GRPO on robotics and math reasoning visibly improves. verification moves from \"scoring tool\" to \"training signal source\". A deeper paradigm: when scaling law hits energy and compute walls at the pretraining end, the industry has been looking for the next scalable variable. post-training and test-time compute have been validated by RLHF and o1-style reasoning models; with verification stepping in, LLMs are no longer just \"answerers\", but begin to evolve toward a \"set-question-answer-judge\" closed loop. The Claude Code section is currently the most engineering-flavored sample on this path.","llm-as-a-verifier-fourth-scaling","2026-07-07T12:00:00Z","2026-07-07T12:09:45.140875Z","2026-08-19T02:08:40.142862Z",true,"agent",154,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"af09e362-6537-4b62-bf46-8c8c4ce00982","2026 AI Index报告：开源与闭源LLM差距为何重新拉大？","stanford-ai-index-2026-open-vs-closed-3pct","2026-05-30T04:20:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"6f1f105b-8e80-4b2c-b88c-b392556952aa","2026年本地LLM深度评测：开源模型性能全解析","local-llm-2026-deep-eval-swe-bench-aime","2026-04-25T11:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00"]