[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-claude-opus-5-medium-effort":3,"topics-all":36,"news-related-6811f1f4-612b-4a99-824c-8678d2113177":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"6811f1f4-612b-4a99-824c-8678d2113177","Claude Opus 5 的真正卖点不是更强,而是 medium effort 这一档","Anthropic 7 月 24 日发布的 Claude Opus 5,API 定价纹丝未动,仍是 $5\u002F$25 每百万 token。真正有意思的不是它跑分多高,而是它把一个叫 **effort(思考强度)** 的旋钮做成了产品功能,而且官方实测显示 **medium 档** 是甜区。\n\n## \"省力模式\"也能拿好成绩\n\nOpus 5 的 effort 参数控制模型推理时愿意花多少 token 思考,而非控制输出长度。系统卡里最扎眼的一组数据:**OSWorld 上,Opus 5 medium effort 跑出 24% 得分,单任务成本仅 $0.89**,把 Opus 4.8 和 Fable 5 都甩在身后,而后者单任务动辄 $12-20。换句话说,中等 effort 不是\"将就\",而是**当前 token 预算下性价比最高的一档**。\n\n背后的真相是:很多 frontier 模型在高 effort 上堆思考预算时,边际收益已经递减。Opus 5 的训练明显针对低-中 effort 区间做了对齐——`low` 和 `medium` 在多个评测上用几个百分点的质量损失,换 3-5 倍的成本下降和延迟缩短。\n\n## 真正的战场是 agent 成本曲线\n\n对部署 Claude Code、computer use 这类 agent 的团队,benchmark 分数是 PPT 数字,**单任务美元成本**才是采购 KPI。Anthropic 同时宣布 Sonnet 5 在 medium effort 下\"能匹配 Opus 4.8\"——意味着同一个团队可以用 Sonnet 的预算做到上一代 Opus 的活。\n\nOpus 5 medium 和 Sonnet 5 medium 这一对组合,等于在不动 API 价格的情况下,把 Anthropic 的实际有效推理成本拉低了 40-60%。这是 2026 年大模型竞争里**最隐蔽也最致命**的一种降价方式:不调单价,改 thinking 默认档位。\n\n## 评论\n\n把 effort 当成 first-class API 参数暴露给开发者,是 Anthropic 过去一年最关键的产品决策。它把\"模型+推理预算\"打包售卖,让用户自己 trade-off。这条路如果走通,会把 LLM 市场从\"我比 GPT 又高 X%\"的军备竞赛,拉进\"我替你每任务省 Y 美分\"的运维竞赛。\n\n对国内厂商的启示也很直接:Qwen、DeepSeek、Kimi 下一阶段拼的不只是 SWE-bench 跑分,**能不能像 Opus 5 那样,在 medium 推理预算下还稳得住质量和工具调用成功率**,才是 agent 落地真正的护城河。","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fabout-claude\u002Fmodels\u002Fwhats-new-opus-5","1fa87d30-d9f3-4752-b3be-0373933b3aaf",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"207db25e-2822-4a79-bcb9-f757109a9fe0","en","Opus 5's real selling point: the medium effort tier","Anthropic released Claude Opus 5 on July 24. API pricing is unchanged — still $5 \u002F $25 per million tokens. The truly interesting thing isn't how high it scores, but that it has turned a knob called **effort** (thinking intensity) into a product feature — and the official benchmarks show the **medium** tier is the sweet spot. ## \"Economy mode\" still scores well Opus 5's effort parameter controls how many tokens the model is willing to spend thinking during inference — not the output length. The most eye-catching number in the system card: **on OSWorld, Opus 5 at medium effort scores 24%, with a per-task cost of just $0.89** — leaving both Opus 4.8 and Fable 5 in the dust, where the latter runs $12–20 per task. In other words, medium effort isn't \"settling\" — **it's the highest cost-effectiveness tier at the current token budget.** The truth behind this: many frontier models, when stacking thinking budget at high effort, are already seeing diminishing marginal returns. Opus 5's training is clearly aligned for the low-to-medium effort range — `low` and `medium` cost a few points of quality for 3–5x lower cost and latency on multiple benchmarks. ## The real battlefield is the agent cost curve For teams deploying Claude Code, computer-use, or similar agents, benchmark scores are PPT numbers — **per-task dollar cost** is the procurement KPI. Anthropic also announced that Sonnet 5 at medium effort \"can match Opus 4.8\" — meaning the same team can do the work of last-gen Opus on a Sonnet budget. The pair Opus 5 medium + Sonnet 5 medium effectively lowers Anthropic's actual effective inference cost by 40–60% without moving the unit price. This is the most subtle and most lethal form of price cut in 2026's LLM contest: don't change the unit price, change the default thinking tier. ## Commentary Exposing effort as a first-class API parameter to developers is Anthropic's most important product decision in the past year. It bundles \"model + inference budget\" and lets users make the trade-off themselves. If this path works, it'll pull the LLM market away from \"I'm X% better than GPT\" arms races and into \"I save you Y cents per task\" ops races. The lesson for Chinese vendors is also direct: what Qwen, DeepSeek, and Kimi need to compete on next isn't just SWE-bench scores — **whether they can stay stable on quality and tool-calling success rate at medium inference budget, the way Opus 5 does**, is the real moat for agent deployment.","claude-opus-5-medium-effort","2026-07-26T02:00:00Z","2026-07-26T00:04:20.047876Z","2026-08-19T02:08:40.142862Z",true,"agent",120,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"8bc2aee8-1a09-4e55-9965-d398cbeebab6","Claude Opus 4.7 新 tokenizer 背后的成本真相：标称价格不变，实际账单已悄然膨胀","claude-opus-4-7-tokenizer-35pct-bill","2026-05-22T14:06:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"390c2437-4e4f-45ec-8270-67c5bfa4fa47","ChatGPT、Claude、Grok、Gemini 罕见同时下线,周四早晨全球 AI 集体失声","chatgpt-claude-grok-gemini-thursday-outage","2026-09-05T06:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"f3d17d45-e1a8-4a1b-9449-6813aff06e49","Anthropic 让 Claude 自己修对齐:10 类失败全部见效,还超过人类研究员","claude-automated-alignment-researchers","2026-08-29T13:05:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"39724847-fdc9-4199-ac46-311e7b49d385","Ramp 数据复盘 Fable 5:旗舰上市两月仅占企业 Anthropic 支出 11%,70 倍价差压住前沿模型溢价","ramp-data-fable-5-adoption-plateaus","2026-08-26T08:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"e1724d68-bf0d-4b3f-8047-147796d5d52e","Ramp 8 月指数:Fable 5 企业份额停滞 11%,OpenAI 旗舰跑赢两倍","anthropic-fable-5-plateau-11-percent","2026-08-25T06:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":81},"1051d676-8ed9-4448-b0d5-8db4b844f41f","Claude Fable 5 上线两个月,为什么企业只把 11% 的账单花给最强模型","claude-fable-5-11-percent-anthropic-spend"]