Anthropic Discloses R&D Automation Levels, Agent Oversight Data and Safety Compute Metrics
Key Takeaways
- •Anthropic's new R&D Automation Index rates every internal AI R&D task on Epoch AI's AL0–AL5 automation scale, and as of August 2026 AI leads about 26% of the company's R&D work while Claude remains non-autonomous in all measured areas.
- •More than 90% of Anthropic's R&D tasks rank at or above the level where AI collaborates with humans, figures the company uses to gauge the industry's distance from recursive self-improvement.
- •Roughly 30,000 AI agents conduct research and engineering on Anthropic's most-used internal platform, and of more than a billion agent decisions analyzed in August, about 0.002% were blocked, with around 100,000 transcripts flagged weekly and only about 50 escalated to human review.
- •In a one-week compute snapshot, approximately 6% of Anthropic's AI R&D compute went to safety work, rising to about 12% within AI-driven R&D, estimates the company describes as deliberately conservative.
- •Anthropic plans to embed independent third-party evaluators with access comparable to its internal risk teams, but acknowledges that no shared methodology yet exists for comparing automation metrics across labs.

Anthropic has released a set of internal measurements intended to give the public, regulators, and third-party observers clearer visibility into how quickly frontier AI development is advancing — including how much of the work of building new models is now being performed by AI itself.
The centerpiece of the release is the Anthropic R&D Automation Index, a prototype metric that catalogues every kind of AI research and development task at the company and rates each one on an automation scale developed by Epoch AI. The scale runs from AL0, meaning no AI involvement, to AL5, meaning fully autonomous operation with no human in the loop. Within that framework, AL3 denotes tasks in which AI collaborates under close human, while AL4 covers tasks that AI completes end-to-end from a high-level prompt with human supervision. The distance between those rungs and AL5 is what gives the index its significance: it measures, task by task, how far the company's R&D operation sits from the point at which AI could carry the work with no human in the loop.
As of August 2026, Claude is not operating fully autonomously in any measured area of AI R&D. Even so, AI leads approximately 26% of Anthropic's R&D work, and more than 90% of tasks rank at or above the level at which AI collaborates with humans. Anthropic frames these figures as a way to track how close the industry is to recursive self-improvement — a scenario in which a model autonomously builds its own successor — and notes that the numbers would likely shift under any coordinated effort to slow frontier development. Framed that way, the release converts an abstract concern into a stated, repeatable baseline against which future disclosures can be measured.
The company announced the disclosure in a post on X:
AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public. Today, we're sharing three measurements that help track development: 1. How much AI R&D is done by AI. 2. How…
— Anthropic (@AnthropicAI) September 17, 2026
Together, the measurements cover how much AI R&D is performed by AI, how AI agents are supervised in practice, and how compute is allocated across the company's work.
The company paired the index with new data on the oversight of AI agents. Roughly 30,000 agents were conducting research and engineering work on Anthropic's most-used internal platform, and every one of their actions passes through real-time online monitors before execution, as well as offline monitors that review activity afterward. Across more than a billion agent decisions analyzed during August, about 0.002% — roughly one in 47,000 — were blocked. Offline monitors flag one to two transcripts per thousand, with around 100,000 flagged weekly and only about 50 escalated to human review. The numbers trace a narrowing supervision funnel — from more than a billion decisions analyzed in August to roughly 100,000 transcripts flagged weekly and about 50 escalated for human review — showing how automated screening is layered beneath human judgment as agents take on real research and engineering work.
Compute Allocation Emerges as a Verifiable Pacing Lever
The third measurement addresses compute, one of the most verifiable inputs to AI development and therefore a potential lever for any future industry-wide pacing regime. In a one-week snapshot of Anthropic's total compute usage, approximately 6% of the compute devoted to AI R&D went to safety work; within AI-driven R&D specifically, the figure was about 12%. The company describes these as deliberately conservative estimates, noting that compute is an imperfect proxy for safety investment, since much safety research is design-intensive rather than compute-intensive. The percentages are also the kind of hard, countable figure that outside parties — or a future pacing agreement — could track over time.
Anthropic positions the release as complementary to its Responsible Scaling Policy risk reports and its Advanced AI Framework policy proposal, which outlines transparency obligations that governments could impose on labs. The company also announced plans to embed independent third-party evaluators from multiple organizations inside Anthropic, granting them access to internal systems and data comparable to what its own risk assessment teams receive.
Cross-lab comparison remains a challenge. Anthropic acknowledges that there is no shared methodology for automation measurement and that using its own models to evaluate its systems creates potential blind spots. It points to third-party verification, or evaluation by other developers' models with safeguards for competitively sensitive information, as a path toward standardized reporting that governments and the public could rely on. Until such a methodology exists, a figure like the 26% automation rate has no agreed counterpart at other developers — making the planned embedding of independent evaluators, and any movement toward common measurement standards, the concrete markers of whether this transparency effort can extend beyond a single lab.
Source: Metaverse Post