[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-claude-opus-4-5-coding-crown-reclaim":3,"topics-all":36,"news-related-5f699f9d-2750-42f8-956f-918507cf51b2":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"5f699f9d-2750-42f8-956f-918507cf51b2","Claude Opus 4.5 发布：Anthropic 夺回编程能力榜首位置","Anthropic 推出了其最新旗舰模型 Claude Opus 4.5，官方声称这是目前世界上最好的编程、Agent 和计算机使用模型。在 SWE-bench 真实 GitHub Bug 修复测试中，Opus 4.5 以 80.9% 的得分超越 GPT-5.1（76.3%）和 Gemini 3 Pro（76.2%），创下了该基准的历史新高。\n\n在推理能力方面，Opus 4.5 在 ARC-AGI-2 抽象推理测试中得分 37.6%，是 GPT-5.1（17.6%）的两倍以上，领先 Gemini 3 Pro 约 6 个百分点。但 Gemini 3 Pro 在 Humanity's Last Exam 上略胜一筹，差距在 2-7% 之间。\n\n值得关注的是安全能力。Anthropic 强调 Opus 4.5 在 Agentic Safety 评估中表现出行业领先的鲁棒性，对抗 Prompt Injection 攻击的能力比 GPT-5.1 和 Gemini 3 Pro 约好 10%。这表明 Anthropic 在推进模型能力的同时，并未放松对安全性的关注。\n\nOpus 4.5 已通过 Claude API 提供，支持 Anthropic 扩展的工具使用和 Agentic 功能。这场头部玩家的激烈竞争，正在将大模型的能力边界快速向前推进。","https:\u002F\u002Fthenewstack.io\u002Fanthropics-new-claude-opus-4-5-reclaims-the-coding-crown-from-gemini-3","36b553c9-6310-4d07-ba39-00b877d0f8ce",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",{"id":21,"name":22,"slug":22,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"64fce462-15e5-4f90-8406-94075d231d43","en","Claude Opus 4.5 released: Anthropic reclaims the coding crown","Anthropic released Claude Opus 4.5, which reclaimed the coding crown from Gemini 3 according to LMSYS Blog analysis. The model shows significant improvements in long-horizon coding tasks, multi-file refactoring, and tool use, marking Anthropic's return to the top of the coding leaderboard.","claude-opus-4-5-coding-crown-reclaim","2026-05-29T11:05:00Z","2026-05-29T19:04:50.961799Z","2026-08-19T02:08:40.142862Z",true,"agent",166,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"63b28ef5-3ffa-4897-88d2-5dcd7fb678b5","Claude Fable 5.1 发布:缓存读取降价 75%,Agent 科研基准翻倍","claude-fable-5-1-mythos-release","2026-09-02T13:20:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"18126c1f-a52c-4a62-ad31-e621b467a196","Claude Code Desktop v2.1.206：浏览器进 IDE,\u002Fdoctor 会修,顺手堵掉「假审批」漏洞","claude-code-desktop-v2-1-206","2026-07-12T16:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"7ca1f9d4-e3e3-48e3-bfb2-ee5a9e6d5176","Anthropic 拆开 Claude Code：别再只换模型，把\"努力度\"也调对","anthropic-claude-code-effort-level","2026-07-12T07:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"8a5451e5-dca1-4416-84da-b06b31b03c49","Claude Code 在系统提示里悄悄埋 Unicode 标记：开发者工具的暗信号边界在哪","claude-code-prompt-steganography","2026-07-01T02:01:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"e9d1fece-f9c1-45fc-9ecc-a647c4002c13","Harness 递归登场:RAH 把 Coding Agent 的长上下文准确率从 71.75% 抬到 89.77%","rah-recursive-agent-harness-89-77pct","2026-06-13T22:30:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"d4c2c83a-f1ea-4a1e-848e-5f1169e9d42b","微软MAI-Thinking-1：清洁训练的35B MoE推理模型，对位Claude Opus 4.6与Sonnet 4.6","mai-thinking-1-microsoft-35b-moe-clean","2026-06-02T06:14:00+00:00"]