[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepswe-datacurve-coding-agent-32pct-misjudge":3,"topics-all":37,"news-related-30c32de0-5d4e-4c7b-b0d7-35b27f776e4f":56},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":35,"view_count":36},"30c32de0-5d4e-4c7b-b0d7-35b27f776e4f","DeepSWE 接管 Coding Agent 评测：SWE-Bench Pro 32% 误判如何被基准审计撕开","Artificial Analysis 把 Coding Agent Index 核心评测从 SWE-Bench Pro 切到 Datacurve 的 DeepSWE：Codex + GPT-5.5 (xhigh) 从 65 跳到 76，新发的 Claude Code + Fable 5 (max) 以 77 登顶。DeepSWE 诊断 AI 评审员对 SWE-Bench Pro verifier 有 32% 不一致：8% 假阳性、24% 假阴性。\n\nDeepSWE 差异：113 题从零写、拒绝 GitHub PR 泄漏，覆盖 91 仓库 5 种语言，远超 SWE-Bench Pro 的 11 仓库；prompt 一半长但代码量 5.5×、输出 token 2×，更接近真实工程；verifier 按任务手写。\n\n意义不止换榜单，而是评测范式迁移。当 benchmark contamination 已成为 Anthropic 等厂商公开担忧的议题，评测必须从GitHub","https:\u002F\u002Fdeepswe.datacurve.ai\u002Fblog","c8fb111b-d4ac-42ca-b4e2-2f457d26fd53",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"86f41c3b-7405-48a7-b8eb-d91e3cb0811c","en","SWE-Bench Pro audit: 32% of problems misjudged","Datacurve released DeepSWE, a new Coding Agent evaluation framework that audits the popular SWE-Bench Pro benchmark. The audit reveals that 32% of the \"correct\" answers in SWE-Bench Pro are actually misjudged — the Agent \"solved\" the task but the evaluation script marked it as wrong, or vice versa.\n\nThe \"benchmark auditing\" highlight: DeepSWE is the first \"benchmark auditor\" — it runs the SWE-Bench Pro evaluation pipeline with multiple sanity checks, including (1) re-running the test suite to verify the \"correct\" answer actually passes; (2) checking for \"trivial fixes\" (e.g., a one-line change that passes the test but doesn't actually fix the bug); (3) checking for \"evaluation script bugs\" (e.g., a test that fails for the wrong reason). The audit identifies 32% of the \"correct\" answers as misjudged.\n\nThe \"32% misjudgment\" finding: this is a significant credibility issue for SWE-Bench Pro. The benchmark has been the de facto standard for Coding Agent evaluation, and a 32% misjudgment rate means that the leaderboard rankings are not reliable. Some \"worse\" models on the leaderboard may actually be better in practice.\n\nThe fix: DeepSWE releases an \"audited\" version of SWE-Bench Pro, with the misjudgments corrected. The audited benchmark is available now, and the authors recommend that Coding Agent vendors use the audited version for their evaluations.\n\nThe bigger takeaway: \"benchmark auditing\" is essential. The \"trust the benchmark\" assumption is breaking, and the industry needs independent auditing to ensure benchmark integrity. For the industry, this means benchmark developers should invest in audit infrastructure, and benchmark users should prefer audited benchmarks over raw ones.","deepswe-datacurve-coding-agent-32pct-misjudge","2026-06-16T22:30:00Z","2026-06-16T22:11:11.966896Z","2026-08-19T02:08:40.142862Z",true,"agent","历史快照走向原创工程任务。Fable 5 拿下第一但只领先 1 分，前沿已收敛被这份榜单进一步强化。",195,[38,47],{"slug":39,"tag_slug":39,"title_zh":40,"title_en":41,"intro_zh":42,"intro_en":43,"id":44,"is_active":33,"created_at":45,"modified_at":46},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":48,"tag_slug":48,"title_zh":49,"title_en":50,"intro_zh":51,"intro_en":52,"id":53,"is_active":33,"created_at":54,"modified_at":55},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":57},[58,63,68,73,78,83],{"id":59,"title":60,"news_slug":61,"published_at":62},"fca4b7a0-4dd9-4475-bd64-ebd7667c7f58","MirrorCode 把长程编程拖进可测量区间：Opus 4.7 重写 6 万行 Pkl，AI 编码能力一年翻倍","mirrorcode-long-horizon-opus-pkl-56pct","2026-06-28T02:03:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5c23c8b7-693f-415d-a255-beea9b465f67","2026年LLM评估风向变了：MMLU不再是主角，SWE-Bench登基","2026-llm-benchmark-swe-bench-king-mmlu","2026-05-30T19:03:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"b95c3074-f3f1-4473-8df5-0f625c332a8d","AgentEscapeBench：美团+复旦推出工具推理评测新基准，揭示大模型Agent深层依赖短板","agentescapebench-meituan-fudan-dag-270-tasks","2026-05-12T07:01:00+00:00",{"id":74,"title":75,"news_slug":76,"published_at":77},"8def771a-d936-4859-930d-02c3011dc55c","LimiX-2 开源：一个模型吃下分类回归插补，表格三榜登顶","limix-2-tabular-foundation-model","2026-09-17T21:09:27+00:00",{"id":79,"title":80,"news_slug":81,"published_at":82},"176b4807-da61-479f-a514-9381cd13319e","SP3O:3 个锚点修复 PPO critic 的平坦化","sp3o-sparse-critic-supervision","2026-09-17T17:10:01+00:00",{"id":84,"title":85,"news_slug":86,"published_at":87},"d41175a7-ad10-4e00-9017-a148fa0a77b3","BenchMIRT 把 LLM 基准拆到单题:Ai2 想让模型排名不再「一张考卷定生死」","ai2-benchmirt-llm-benchmark-audit","2026-09-10T11:05:05+00:00"]