The Anthropic Institute, Anthropic's research arm, published the long-form article "When AI builds itself," using public benchmarks and undisclosed internal data to quantify the penetration speed of AI in AI's own R&D, calling on the industry to establish a "brake mechanism" to deal with the "recursive self-improvement" inflection point.
External trends: the capability curve is accelerating
METR data shows that the doubling period of the credible duration that AI can independently complete a task has been compressed from seven months to about four months: Claude Opus 3 (2024.3) could only handle 4-minute-level tasks, Sonnet 3.7 reached 1.5 hours a year later, Opus 4.6 reached 12 hours a year after that; Claude Mythos Preview has been measured by METR to be able to work continuously for at least 16 hours. The article infers that if the trend continues, "week-level" task capability may emerge in 2027. On the benchmark side, SWE-bench (real software-engineering bug fix) went from single-digit accuracy to saturation in two years; CORE-Bench (requiring reproduction of published papers) took only 15 months to go from 20% success in 2024 to saturation.
Internal evidence: Claude is already Anthropic's "colleague"
In May 2026, more than 80% of merged code at Anthropic was written by Claude, while before Claude Code launched in February 2025 this proportion was in single digits; daily merged code per engineer in 2026 Q2 is already 8× that of 2024. On the research side, Claude can already match or surpass senior engineers on "well-defined goal" experimental execution, but the most senior capability of "judging whether the goal is worth pursuing" still has a clear gap — and this is exactly the key bottleneck blocking "complete self-improvement."
Judgment
The "hard constraint" of recursive self-improvement lies in judgment, not in the speed of writing code or running experiments. The value of Anthropic's article is to bring the discussion of "AI transforming AI" from a philosophical proposition down to quantifiable engineering metrics — "brake" may not become executable policy, but the industry can no longer treat this matter as a long-term hypothesis.