[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-anthropic-80pct-claude-self-improvement-brake":3,"topics-all":36,"news-related-44b41276-ea01-4145-ada5-b9cd91ab08a0":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"44b41276-ea01-4145-ada5-b9cd91ab08a0","Anthropic：80% 代码由 Claude 编写，AI 自我改进刹车时刻临近","Anthropic 旗下研究机构 The Anthropic Institute 发布长文《When AI builds itself》，用公开基准与未公开的内部数据，量化 AI 在 AI 自身研发中的渗透速度，呼吁行业建立「刹车机制」以应对「递归自我改进」（recursive self-improvement）拐点。\n\n**外部趋势：能力曲线正在加速**\n\nMETR 数据显示，AI 独立完成任务的可信时长翻倍周期已从七个月压缩到约四个月：Claude Opus 3（2024.3）只能处理 4 分钟级任务，Sonnet 3.7 一年后达到 1.5 小时，Opus 4.6 又一年后达到 12 小时；Claude Mythos Preview 已被 METR 测出可连续工作至少 16 小时。文章推断，若趋势延续，2027 年可能出现「周级」任务能力。基准方面，SWE-bench（真实软件工程 bug 修复）两年内从个位数正确率走到饱和；CORE-Bench（要求复现已发表论文）从 2024 年的 20% 成功率到饱和仅用 15 个月。\n\n**内部证据：Claude 已经是 Anthropic 的「同事」**\n\n2026 年 5 月 Anthropic 合并代码中超过 80% 由 Claude 编写，而 2025 年 2 月 Claude Code 推出前这一比例还是个位数；单工程师日合并代码量在 2026 Q2 已是 2024 年的 8 倍。研究侧，Claude 已能在「目标明确」的实验执行上匹敌甚至超越资深工程师，但「判断目标是否值得做」这一最 senior 能力仍有明显差距——这正是阻挡「完全自我改进」的关键瓶颈。\n\n**判断**\n\n递归自我改进的「硬约束」在于 judgment 而非写代码或跑实验的速度。Anthropic 这篇文章的价值在于把「AI 改造 AI」的讨论从哲学命题落到可量化的工程指标上——「刹车」未必能成为可执行政策，但行业已经不能再把这件事当作远期假设。","https:\u002F\u002Fwww.anthropic.com\u002Finstitute\u002Frecursive-self-improvement","1fa87d30-d9f3-4752-b3be-0373933b3aaf",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":18,"name":19,"slug":19,"description":13,"color":13},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"e3bebd68-06f7-4075-8f6c-fdb79397f12d","en","Anthropic: 80% of code written by Claude as brakes loom","The Anthropic Institute, Anthropic's research arm, published the long-form article \"When AI builds itself,\" using public benchmarks and undisclosed internal data to quantify the penetration speed of AI in AI's own R&D, calling on the industry to establish a \"brake mechanism\" to deal with the \"recursive self-improvement\" inflection point.\n\n**External trends: the capability curve is accelerating**\n\nMETR data shows that the doubling period of the credible duration that AI can independently complete a task has been compressed from seven months to about four months: Claude Opus 3 (2024.3) could only handle 4-minute-level tasks, Sonnet 3.7 reached 1.5 hours a year later, Opus 4.6 reached 12 hours a year after that; Claude Mythos Preview has been measured by METR to be able to work continuously for at least 16 hours. The article infers that if the trend continues, \"week-level\" task capability may emerge in 2027. On the benchmark side, SWE-bench (real software-engineering bug fix) went from single-digit accuracy to saturation in two years; CORE-Bench (requiring reproduction of published papers) took only 15 months to go from 20% success in 2024 to saturation.\n\n**Internal evidence: Claude is already Anthropic's \"colleague\"**\n\nIn May 2026, more than 80% of merged code at Anthropic was written by Claude, while before Claude Code launched in February 2025 this proportion was in single digits; daily merged code per engineer in 2026 Q2 is already 8× that of 2024. On the research side, Claude can already match or surpass senior engineers on \"well-defined goal\" experimental execution, but the most senior capability of \"judging whether the goal is worth pursuing\" still has a clear gap — and this is exactly the key bottleneck blocking \"complete self-improvement.\"\n\n**Judgment**\n\nThe \"hard constraint\" of recursive self-improvement lies in judgment, not in the speed of writing code or running experiments. The value of Anthropic's article is to bring the discussion of \"AI transforming AI\" from a philosophical proposition down to quantifiable engineering metrics — \"brake\" may not become executable policy, but the industry can no longer treat this matter as a long-term hypothesis.","anthropic-80pct-claude-self-improvement-brake","2026-06-06T10:00:00Z","2026-06-06T10:08:19.801204Z","2026-08-19T02:08:40.142862Z",true,"agent",440,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"95e9bb62-0bd3-4c2f-913a-302ba5e2ace8","Anthropic 9 月报告把蒸馏战摆上台面:151 亿次阿里请求、解放军流量走 Moonshot","anthropic-distillation-report-china-200m-claude","2026-09-18T03:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"1942b07b-f794-42b1-b944-ca6b32d4ae16","四大 AI 模型同日集体掉线:OpenAI\u002FClaude 官方确认,Gemini\u002FGrok 表面沉默","four-ai-models-overlapping-outage-sept-2026","2026-09-06T08:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"390c2437-4e4f-45ec-8270-67c5bfa4fa47","ChatGPT、Claude、Grok、Gemini 罕见同时下线,周四早晨全球 AI 集体失声","chatgpt-claude-grok-gemini-thursday-outage","2026-09-05T06:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"264b5334-0465-48b8-ad81-b2b7b39d1a3f","Anthropic 被索尼华纳告上法庭：两万首歌喂出来的 Claude 还要赔多少","anthropic-sony-warner-music-copyright-lawsuit","2026-09-05T00:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"9e58d587-3c1b-44c5-ad36-daf23aeb42a2","微软叫停 tokenmaxxing:GitHub Copilot 默认切回 GPT-5.6 Sol,Parikh 设 token 预算","microsoft-token-budget-gpt-5-6-default","2026-09-03T00:30:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"d3e055f9-1fbe-4d91-9eca-f336b930be80","索尼华纳起诉 Anthropic:每首歌索赔 15 万美元,可能拖出又一份 10 亿美元和解","sony-warner-anthropic-billion-dollar-lawsuit","2026-08-31T11:00:00+00:00"]