
Anthropic randomly sampled 20% of research and development staff’s Claude usage for each week of July, and then used Claude to organize them. This was done using an «automation rating scale» from Epoch AI, classifying tasks from AL0, meaning no AI involvement, to AL5, which is fully autonomous AI.
They found that none of the usage reached AL5, but AL4 was 26% of the tasks — meaning AI is «leading» from «high-level prompts» with human supervision, while more than 90% of the tasks were AL3 — which is AI «collaboration» under «close human direction.»
AI «leading» is up from 1% in February 2026, and Anthropic estimates that they now have 30,000 agents in use at any time for research and engineering, that use real-time monitoring that can block «dangerous actions.» As agent use eventually grows ever larger, even rare use cases become more likely, Anthropic says.
The measure is a way of determining how close we are getting to recursive self-improvement, meaning AI writing itself, which could lead to a lack of control of the technology. Anthropic says every lab could publish such reports if there was a standard, which they hope to contribute to. They will also keep publishing these.
It’s also a step toward better third-party monitoring at Anthropic, which plans to install independent evaluators with employee-level access soon, as promised in their call for «pacing».
Read more: Anthropic’s report. Coverage on Reuters, ABC News, The Washington Post, and Bloomberg. Discussion on r/Singularity.