NEWS · COMPANIES · #627
Anthropic says Claude 'leads' 26% of research work but definition and scoring are fuzzy
Anthropic published internal metrics saying Claude accounts for 26% of model-development work at AL4 (labelled "AI leads") as of August 2026, with over 90% at least AL3; the company computed levels using agents that collected internal logs and a Claude model that assigned scores. Anthropic also reported monitoring figures (about 30,000 concurrent agents, a real-time monitor blocking 0.002% of >1B actions in August) and that ~6% of research compute went to safety work, while acknowledging ambiguity in level boundaries and limits of self-scoring.
KEY POINTS
- Anthropic published internal metrics saying Claude accounts for 26% of model-development work at AL4 (labelled "AI leads") as of August 2026, with over 90% at least AL3; the company computed levels using agents that collected internal logs and a Claude model that assigned scores.
- Anthropic also reported monitoring figures (about 30,000 concurrent agents, a real-time monitor blocking 0.002% of >1B actions in August) and that ~6% of research compute went to safety work, while acknowledging ambiguity in level boundaries and limits of self-scoring.
- The disclosure illustrates both progress in automating research and the limits of self-reported autonomy metrics—definitions, measurement methods, and self-scoring affect trust, comparisons, and policy debates about slowing frontier AI.
WHY IT MATTERS
The disclosure illustrates both progress in automating research and the limits of self-reported autonomy metrics—definitions, measurement methods, and self-scoring affect trust, comparisons, and policy debates about slowing frontier AI.