Tech Meridian ← LIVE FEED
RU

NEWS · COMPANIES · #627

Anthropic says Claude 'leads' 26% of research work but definition and scoring are fuzzy

Anthropic published internal metrics saying Claude accounts for 26% of model-development work at AL4 (labelled "AI leads") as of August 2026, with over 90% at least AL3; the company computed levels using agents that collected internal logs and a Claude model that assigned scores. Anthropic also reported monitoring figures (about 30,000 concurrent agents, a real-time monitor blocking 0.002% of >1B actions in August) and that ~6% of research compute went to safety work, while acknowledging ambiguity in level boundaries and limits of self-scoring.

KEY POINTS

  1. Anthropic published internal metrics saying Claude accounts for 26% of model-development work at AL4 (labelled "AI leads") as of August 2026, with over 90% at least AL3; the company computed levels using agents that collected internal logs and a Claude model that assigned scores.
  2. Anthropic also reported monitoring figures (about 30,000 concurrent agents, a real-time monitor blocking 0.002% of >1B actions in August) and that ~6% of research compute went to safety work, while acknowledging ambiguity in level boundaries and limits of self-scoring.
  3. The disclosure illustrates both progress in automating research and the limits of self-reported autonomy metrics—definitions, measurement methods, and self-scoring affect trust, comparisons, and policy debates about slowing frontier AI.

WHY IT MATTERS

The disclosure illustrates both progress in automating research and the limits of self-reported autonomy metrics—definitions, measurement methods, and self-scoring affect trust, comparisons, and policy debates about slowing frontier AI.

SOURCES & TIMELINE

1