Anthropic reports Claude leads 26% of its AI R&D

Anthropic reports Claude leads 26% of its AI R&D
News

Anthropic has published a new set of measurements for how frontier AI labs build and supervise increasingly capable systems. The Anthropic Institute says the framework is meant to make the pace of AI development more visible to the public, governments and outside evaluators. It covers three areas: how much AI performs in-house AI research and development, how agents are monitored, and how computing resources are allocated.

The first measurement is an Anthropic R&D Automation Index. Using the company’s map of model-development tasks and an automation scale developed by Epoch AI, Anthropic reports that Claude was at the ‘leads’ level for 26% of its AI R&D work in August 2026. On that scale, leads means the model can complete most of a task from a high-level prompt while a human supervises. More than 90% of the work was at least collaborative, Anthropic says, while no measured part was fully autonomous. These are company measurements, not an independent industry benchmark.

Anthropic also describes an oversight system for roughly 30,000 research and engineering agents on its most-used internal platform. It says all actions pass through online monitoring before execution and offline monitoring afterwards. Across more than one billion decisions in August, 0.002% were blocked by the online monitor. The offline system flagged about 100,000 transcripts per week for review, with roughly 50 of the highest-priority cases escalated to people.

The third measurement concerns compute. In a snapshot covering 13–20 July, Anthropic estimates that about 6% of compute used for AI R&D went to safety work, rising to about 12% for AI-driven AI R&D. The company calls these conservative estimates and notes that classifications and the one-week sample have limitations.

Reuters and AP treated the figures as evidence that AI is already helping develop its successors, while also noting that Claude is not operating autonomously. For users and businesses, the practical issue is not whether a model is ‘building itself’ in a human-like sense. It is whether teams can measure permissions, review and compute before agents take on more consequential work. Anthropic says it plans to give independent third-party evaluators access to verify these kinds of measurements. Until then, the figures are a useful transparency proposal and a self-reported snapshot, not proof of a sector-wide trend.