What it means that Anthropic measures Claude leading 26% of its R&D — September 18, 2026
· CompaniesAutomation
Flash Radar edition: The Anthropic Institute publishes its R&D automation index for the first time. Claude "leads" 26% of the work (AL4 on the Epoch AI scale) up from less than 1% in February, with 30,000 agents monitored at 100%. Which step of the AI-First ladder your company occupies and how to measure it this week.
Flash edition. Anthropic has published for the first time how much of its own research Claude performs: it leads 26% of the R&D work. And it has measured this using a levels scale. The rest of the day's news is in this morning's Radar.
What happened
On Thursday, September 17, the Anthropic Institute published “Measurements for understanding the pace of AI development inside frontier labs”, featuring three unprecedented internal metrics. The first is an R&D automation index built on Epoch AI's AL0-AL5 scale: from AL0 (no AI intervention) to AL5 (autonomous, no human). As of August 2026, Claude “leads” (AL4) 26% of the work—completing tasks from end-to-end based on general instructions, with human oversight—compared to less than 1% in February; over 90% is at the “collaborates” level or above. The company itself emphasizes that Claude does not operate fully autonomously in any measured segment. The second metric: approximately 30,000 agents working simultaneously, with 100% of their actions passing through a monitor, 0.002% blocked (1 in 47,000), and ~50 cases per week escalated for human review. The third: only 6% of R&D compute was dedicated to safety. This is confirmed by Engadget.
Why it matters
It's not just a headline about “AI building itself.” It's that the laboratory that has gone the furthest is publishing its progress with a specific number and admitting it hasn't reached the final step. Look at what made the jump from 1% to 26% in six months possible: not just a better model, but a layer of supervision that monitors 100% of what agents do and escalates what matters to a human. That is the order, and it's the same in your company: first you measure and monitor, then you delegate.
For your company
Three things this week, without spending a cent. One: Take your five AI processes and assign each a label—assists, collaborates, leads—and today's date. You'll have your own index and know which step you're truly on, beyond the buzzwords. Two: Copy the three monitoring metrics: what percentage of your agent's actions are logged (coverage), how long it takes for someone to review them (latency), and how many cases end up in human hands (escalation). If the answer to the first is not 100%, that's where the work lies. Three: Ask your provider for those same three numbers during the next renewal. Anthropic has just set the precedent that any developer can publish them.
FAQ
Does this mean Claude is already improving itself?
No. Anthropic explicitly states that no measured segment reaches AL5 and that the gap remains in choosing the goal, not in executing it. Full autonomy is the scenario they are monitoring for, not the current state.
Is this comparable with other companies?
Not yet: the scale belongs to Epoch AI, but the measurement is internal and only Anthropic has published it. That is precisely the stated reason for publishing it: so that others measure the same way.
What can I take away if my company doesn't do programming?
The scale and the three metrics apply to invoices, budgets, or customer service just as much as to code: what matters is how much of what your AI does is recorded.