[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-openai-swe-bench-pro-audit":3,"news-related-9382e481-e16b-4925-a0d1-3b24cd8ba22a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"9382e481-e16b-4925-a0d1-3b24cd8ba22a","OpenAI 复审自家推荐的 SWE-Bench Pro:731 道题里约三成是「坏题」,榜单狂欢该降温","OpenAI 在 7 月 8 日发布了一份罕见的「自我审计」:用 Codex 驱动的 investigator agent 加 5 名资深工程师双盲复核 SWE-Bench Pro 的 731 道公开题。自动管线标记 200 道「坏题」(27.4%),人类评估推到 249 道(34.1%),合并估算约三成题目无法可靠反映模型能力。问题被归为四类:隐藏测试过严、prompt 描述不足、测试覆盖低、题干与测试逻辑相悖,全部出在数据层面,与模型本身无关。OpenAI 明确撤回一年前「推荐社区从 SWE-Bench Verified 切到 SWE-Bench Pro」的立场,这是 12 个月内第二次自我撤回评测建议。更值得警惕的是「数字猛涨」:复审前的八个月里,前沿模型 pass@1 从 23.3% 飙到 80.3%——近三倍的跳跃里,恐怕有一笔要算到「模型越来越懂测题人默认的实现细节」上,而非真实工程能力。把 30% 坏题剔除后再排名,#1 与 #5 的悬殊很可能远没宣传中那么夸张,模型方常用的「又反超 Claude」之类话术也得跟着打折。方法学上更有价值的,是 OpenAI 这套「agent + human」双轨流水线:Codex 已能像审计员一样跑测试、查代码、抓失败模式,把昂贵的人工质检规模化。这预示了未来 LLM 评测的新形态——让 LLM 先帮我们质问基准本身,再由人类挑刺,循环改进。对厂商和开发者最现实的提醒:Coding Agent 在 SWE-Bench Pro 上的亮眼分数,先问一句「它是不是强行往 prompt 暗示的实现细节走」;选型别只盯榜单一两个百分点,看真实仓库任务里的修复率更靠谱。","https:\u002F\u002Fopenai.com\u002Findex\u002Fseparating-signal-from-noise-coding-evaluations\u002F","bd0e0e04-6bcf-4b3e-9a56-62c672308ec9",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"4127e157-753d-4179-8d1e-3ad473a4e722","en","OpenAI re-audits SWE-Bench Pro: 30% of 731 problems flawed","OpenAI on July 8 published a rare \"self-audit\": using a Codex-driven investigator agent plus 5 senior engineers double-blind reviewed the 731 public problems of SWE-Bench Pro. The automated pipeline flagged 200 \"bad problems\" (27.4%); human review pushed it to 249 (34.1%); combined estimate, about 30% of problems cannot reliably reflect model capability. The issues are classified into four categories: hidden tests being too strict, prompt description being insufficient, test coverage being low, problem description contradicting test logic — all in the data layer, nothing to do with the model itself. OpenAI explicitly retracted its year-old position of \"recommending the community to switch from SWE-Bench Verified to SWE-Bench Pro\" — the second self-retraction of an evaluation recommendation in 12 months. More alarming is the \"numbers exploding\": in the eight months before the re-audit, the frontier model's pass@1 went from 23.3% to 80.3% — in this near-triple jump, a sizable portion probably has to be chalked up to \"models getting better at reading the test-setter's default implementation details\" rather than real engineering capability. After removing the 30% bad problems, the gap between #1 and #5 is probably far less exaggerated than the hype, and the \"just beat Claude again\" rhetoric the vendors love also needs to be discounted. The more valuable methodological piece is OpenAI's \"agent + human\" dual-track pipeline: Codex can already run tests, check code, and catch failure modes like an auditor, scaling up expensive manual QA. This heralds a new form of LLM evaluation — letting LLMs first question the benchmark itself, then humans nitpick, in a cycle of improvement. The most realistic reminder for vendors and developers: when Coding Agents post impressive scores on SWE-Bench Pro, first ask \"is it just blindly following the implementation details hinted at in the prompt\"; for selection, don't just stare at one or two percentage points on the leaderboard, the fix rate on real repository tasks is more reliable.","openai-swe-bench-pro-audit","2026-07-13T00:11:00Z","2026-07-13T00:14:00.641036Z","2026-08-19T02:08:40.142862Z",true,"agent",102,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5b76d928-8896-4bf8-8edd-e195ecf0094a","Ornith-1.5自报跑分赢了Claude,独立复测翻车了","ornith-1-5-benchmark-reality-check","2026-08-20T17:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"4c54e7dc-46f6-4d0e-86c9-278995cf0da8","Stanford PTXBench:让 LLM 裸写 H100\u002FB200 PTX kernel,没有一个模型全过关","ptxbench-llm-ptx-gpu-kernel-benchmark","2026-08-19T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"e846f1c6-4644-4f84-a664-83ec82734210","Meta Muse Spark 1.2 与 Muse Code 把「1.2 → 编程」的推理效率推回前沿","meta-muse-spark-12-coding-agent-54-index","2026-08-05T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"0fd9ee7a-5b8f-49d2-9032-57f763de40e3","OpenAI 下一代模型 Astra 一口气破解 10 个数学难题:从 27 年未决的非 sofic 群到 46 年未动的高维球体堆积","openai-astra-ten-math-proofs-2026","2026-08-01T10:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"de219584-58fc-45ad-91ea-0049a5cbcf10","OpenAI 开源 Codex Security CLI:把 AI 安全检测塞进每个 PR","openai-codex-security-cli-opensource","2026-07-29T10:30:00+00:00"]