
AI that helps develop better AI introduces an unusually consequential feedback loop. A useful research assistant could help improve the tools used to build its own successors. To understand how far that process has progressed, the first question is what work has actually changed.
Anthropic’s September 17 disclosure says Claude led 26% of its AI research and development in August. Over 90% involved at least substantial AI collaboration, including that 26%. No measured category was fully autonomous. Read Anthropic’s disclosure.
The figures deserve attention because they describe work inside a frontier lab. Interpreting them requires three separate questions: how much work can be delegated, how much useful progress that produces, and whether people can still inspect and challenge the results.
Here, “leads” means completing most of a task from broad instructions while a human supervises. “Collaborates” means substantial work under closer human direction. The automation definitions.
Consider an illustrative research assignment: investigate why an experiment’s results changed. An assistant might gather logs, compare configurations, propose an explanation and run a check. That could remove hours of investigation while leaving a researcher responsible for deciding whether the explanation is convincing and what to try next.
The distinction is practical. Completing an assigned investigation does not establish that a system can choose a productive research agenda, manage a whole program or decide when a result is ready to use. Those require separate evidence.
Researchers at Epoch AI have proposed a detailed inventory of AI R&D work precisely because broad labels conceal different jobs. Their June taxonomy covered more than sixty tasks across six categories, including work involved in operating infrastructure and monitoring runs. A task inventory helps establish which parts of research a capability claim covers. Epoch AI’s research task framework.
Anthropic built its index using Claude to reconstruct and rate work categories from internal records. It weighted them with a staff-time approximation and held July’s task basket fixed. Staff ratings provided a cross-check, but the company acknowledges judgment calls and correlated model errors. The index methodology.
A weighted index is not a headcount. Automating substantial parts of many people’s work can produce a high score without eliminating any individual role. Nor does a quarter of a weighted task basket necessarily represent a quarter of the decisions that determine a lab’s direction.
The choice of denominator also affects comparisons. An index dominated by recurring engineering work answers a different question from one weighted toward selecting experiments or assessing unusual failures. Readers need the task mix alongside the headline percentage to know what changed.
For comparisons over time, a stable yardstick is helpful. It also needs periodic checks against the work people are actually doing. If researchers move into new activities, a score based on yesterday’s duties can become less representative of today’s research effort.
A simple hypothetical shows why the distinction matters. Suppose a project takes ten days: two preparing experiment code and eight running experiments, reviewing results and deciding what follows. If an assistant cuts preparation from two days to half a day, that stage becomes four times faster. The project takes eight and a half days, a 15% reduction in elapsed time, assuming the other stages are unchanged.
Real projects can behave differently. Work may overlap, reviewers may become a bottleneck, or a faster prototype may let a team test an idea it would otherwise abandon. The example illustrates why a local speed gain cannot simply be applied to the whole research process.
METR’s analysis of task substitution adds another complication: when AI makes some work cheaper, people change what they choose to do. Measuring speed on the resulting task mix can give a different answer from measuring the increase in valuable output. Extra work can be worthwhile without being as valuable as its hypothetical manual cost suggests. METR on task substitution and productivity.
For a research team, the useful outcome might be a reproducible improvement, an invalid idea ruled out early, or a safety issue discovered before deployment. Counting experiments or generated code alone would miss those differences. A convincing acceleration claim should connect the activity to an outcome and include the cost of checking it.
Anthropic also reports approximately 30,000 concurrent agents on its most-used internal platform, with all their actions passing through online monitors. These figures cover that platform. The monitoring disclosure.
Passing an action through a check establishes that a check happened. To judge its effectiveness, we need to know how it performs when something is wrong. A system could examine every action and still miss a particular class of problem; frequent alerts could also overwhelm reviewers with harmless cases.
Timing matters too. Catching an incorrect result before it influences the next experiment is different from discovering it after several projects have relied on it. An oversight process needs to make the original evidence, the decisions and the later corrections easy to trace.
There is relevant external work, though it predates the August snapshot. In March, METR reported a three-week adversarial test of a subset of Anthropic’s monitoring and security systems. It found previously unknown vulnerabilities, some subsequently patched. METR said none severely undermined the major claims in the risk report it examined. This exercise does not certify the later deployment. METR’s account of the monitoring tests.
The lesson from that exercise is concrete: outside testing can reveal problems that ordinary operation has not exposed. Useful reporting should show what was tested, which failures were found, and whether fixes held up when tested again.
The March preprint Measuring AI R&D Automation, by Alan Chan and colleagues, argues that capability benchmarks alone can miss both real-world automation and its consequences. It proposes tracking measures such as researchers’ time allocation, spending and incidents, including whether oversight and safety progress keep pace with capability development. The measurement research.
Our assessment is that the next useful disclosure should connect several kinds of evidence. Task ratings should be accompanied by examples that independent reviewers can reassess. Claims about faster research should identify completed outcomes, the comparison being made, and the human review effort required. Monitoring reports should include tests designed to expose missed problems.
The same standard should recognize benefits. A lab that uses agents to investigate more failure modes, reproduce results or discard weak ideas could be improving research quality even when the gain is hard to express as a single speed multiplier. Reporting should leave room for those outcomes as well as faster model development.
Anthropic says it plans to embed independent evaluators with substantial internal access. Its proposed verification approach.
The value of that access will depend on the questions evaluators can investigate and the findings they can communicate. For the public, the most informative next result would connect a specific delegation of work to a verified improvement, with enough detail to understand where human judgment remained essential.
Reporting note: This analysis uses public disclosures and research available on September 21, 2026. We have not independently audited Anthropic’s internal systems. Examples and the ten-day calculation are illustrative.
Cover: AI-generated editorial illustration of successive neural networks and human review.