[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-claude-opus-5-thinking-default":3,"topics-all":36,"news-related-20904504-4450-4f2e-a612-49b191403a86":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"20904504-4450-4f2e-a612-49b191403a86","Claude Opus 5 thinking 默认开启：OSWorld 70.57% 的慢即是快","7月24日,Anthropic 发布 Claude Opus 5,价格不变 ($5\u002F$25 每百万 token),上下文 1M,模型 ID 写作 claude-opus-5。三个 API 变化比跑分更重要:thinking 默认开启(老代码里 `thinking: {\"type\": \"disabled\"}` 配 effort xhigh\u002Fmax 直接 400)、最低可缓存 prompt 砍到 512 token、官方劝开发者把「再 verify 一次」这类提示词删掉——模型自己会 verify,再要求会过拟合到过验证。\n\n跑分上,FrontierBench v0.1(74 个 agentic 任务,Terminal-Bench 2.1 继任)Opus 5 max effort 拿 43.3%,xhigh 拿到 44.4% 是其最佳成绩;Opus 4.8 只有 18.7%,Fable 5 是 33.7%,GPT-5.6 Sol 是 37.5%。SWE-bench Verified 96.0%,Multimodal 从 38.4% 跳到 59.4%。OSWorld 2.0 70.57%(4.8 是 55.7%),Zapier AutomationBench 26.0%(Fable 5 是 17.4%),GDPval-AA v2 拿下 ELO 1861\u002F1827 头部两个名额。最炸裂的是 ARC-AGI-3 高 effort 验证 30.16%——之前榜首是 GPT-5.6 Sol 的 7.78%,Opus 4.8 是 1.52%。\n\n「工具 > 思考」是这版最值得记住的工程信号:Chartography 不带工具 29.6%,挂上图片裁剪工具直接 83.0%;BenchCAD Vision2Code voxel IoU 从 0.366 涨到 0.821,顺便把 Mythos 5 的 0.678 踩了。安全上,Gray Swan 间接 prompt injection 15 次内的攻击成功率从 5.5% 跌到 2.0%(GPT-5.6 Sol 是 20.0%);浏览器场景 Claude Cowork 默认开 auto 模式时,129 个测试环境攻击成功率 0%。Cyber 能力跟着 general 涨,Anthropic 解锁了「源代码漏洞查找」,但二进制扫描、渗透测试、exploit 生成仍然锁死——RSP 把它定到 ASL-3,与 Opus 4.8 同级,UK AISI 的 'The Last Ones' 10 跑里通了 8 次。\n\n模型 ID 仍是 opus 5 这一档,定价 $5\u002F$25 不动,既是对 Fable 5($10\u002F$50)继续做企业市场分层,也是对开源阵营(Kimi K3 $3\u002F$15、DeepSeek V4 Flash 更便宜)的防守:1M context + 1\u002F2 价格 + 跑分追到 Fable 5,Opus 5 把 Anthropic 整条产品线在 agentic、coding、长上下文三条线的中段都顶到了新水位。","https:\u002F\u002Fwww.anthropic.com\u002Fnews\u002Fclaude-opus-5","1fa87d30-d9f3-4752-b3be-0373933b3aaf",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"926ad19c-de5f-482d-a463-9e4190da1606","en","Claude Opus 5 thinks by default: slow is fast at 70.57%","On July 24, Anthropic released Claude Opus 5. Pricing unchanged ($5 \u002F $25 per million tokens), 1M context, model id is claude-opus-5. Three API changes matter more than the benchmarks: thinking is on by default (in old code with `thinking: {\"type\": \"disabled\"}` paired with effort xhigh\u002Fmax, you immediately 400), the minimum cacheable prompt is cut to 512 tokens, and the team explicitly advises developers to drop prompt phrases like \"verify once more\" — the model verifies itself, and re-asking it over-fits to over-verification. On benchmarks, on FrontierBench v0.1 (74 agentic tasks, the successor to Terminal-Bench 2.1), Opus 5 at max effort scores 43.3%, and at xhigh hits 44.4% — its best score; Opus 4.8 only scores 18.7%, Fable 5 is 33.7%, GPT-5.6 Sol is 37.5%. SWE-bench Verified 96.0%, Multimodal jumps from 38.4% to 59.4%. OSWorld 2.0 70.57% (4.8 was 55.7%), Zapier AutomationBench 26.0% (Fable 5 was 17.4%), GDPval-AA v2 takes the top two ELO slots at 1861\u002F1827. Most explosive is ARC-AGI-3 at high effort scoring 30.16% — the previous leader was GPT-5.6 Sol at 7.78%, Opus 4.8 was 1.52%. The most memorable engineering signal from this version is \"tools > thinking\": Chartography without tools is 29.6%, hook it up to an image-cropping tool and it jumps straight to 83.0%; BenchCAD Vision2Code voxel IoU goes from 0.366 to 0.821, incidentally trampling Mythos 5's 0.678. On safety, Gray Swan's indirect prompt-injection attack success rate within 15 attempts dropped from 5.5% to 2.0% (GPT-5.6 Sol is 20.0%); in browser scenarios with Claude Cowork defaulting to auto mode, attack success across 129 test environments is 0%. Cyber capability rises with general capability, and Anthropic has unlocked \"source-code vulnerability lookup\", but binary scanning, pentest, and exploit generation remain locked — RSP places it at ASL-3, on par with Opus 4.8; UK AISI's 'The Last Ones' passed 8 of 10 runs. The model id stays at opus 5, pricing $5 \u002F $25 unchanged, both as a tier-down to Fable 5 ($10 \u002F $50) to continue the enterprise-market split, and as a defense against the open-source camp (Kimi K3 $3 \u002F $15, DeepSeek V4 Flash even cheaper): 1M context + 1\u002F2 the price + scores chasing Fable 5, Opus 5 pushes the whole Anthropic product line to a new water mark on agentic, coding, and long-context mid-tier.","claude-opus-5-thinking-default","2026-07-25T00:00:00Z","2026-07-25T22:04:44.325530Z","2026-08-19T02:08:40.142862Z",true,"agent",226,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"5930aa08-8a81-4c92-9ec0-79c3f4de39e2","Anthropic 推出 Claude Opus 5:性能逼近 Fable 5,API 价格不变继续啃企业市场","claude-opus-5-launch","2026-07-25T02:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"428c2822-a53a-4075-80ec-2960e03ea062","Anthropic Claude Sonnet 5：中端档拉到 Opus 4.8 水位","claude-sonnet-5-main-model","2026-07-01T02:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"390c2437-4e4f-45ec-8270-67c5bfa4fa47","ChatGPT、Claude、Grok、Gemini 罕见同时下线,周四早晨全球 AI 集体失声","chatgpt-claude-grok-gemini-thursday-outage","2026-09-05T06:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"63b28ef5-3ffa-4897-88d2-5dcd7fb678b5","Claude Fable 5.1 发布:缓存读取降价 75%,Agent 科研基准翻倍","claude-fable-5-1-mythos-release","2026-09-02T13:20:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"f3d17d45-e1a8-4a1b-9449-6813aff06e49","Anthropic 让 Claude 自己修对齐:10 类失败全部见效,还超过人类研究员","claude-automated-alignment-researchers","2026-08-29T13:05:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"39724847-fdc9-4199-ac46-311e7b49d385","Ramp 数据复盘 Fable 5:旗舰上市两月仅占企业 Anthropic 支出 11%,70 倍价差压住前沿模型溢价","ramp-data-fable-5-adoption-plateaus","2026-08-26T08:00:00+00:00"]