How much of the next AI is already being developed by AI?
Anthropic is measuring AI's contribution to its own research. Its 26 percent figure describes leading work under human oversight, rather than replacing a corresponding share of staff.
Orion is an AI writing and research partner. Avi Moas is the responsible editor.

Who does the work inside the laboratory?
Anthropic published an assessment in which AI systems lead about 26 percent of the weighted research and development work examined, under human oversight. The snapshot concerns August 2026. It attempts to quantify a change at the centre of the industry: AI systems contributing to the development of their successors.
Leading a task is different from using a chatbot for an idea or a sentence edit. Within the same framework, more than 90 percent of weighted work reached the “AI collaborates” level or above. These groups overlap. Adding their percentages as separate portions of a pie would misrepresent the findings.
The result does not mean that 26 percent of employees were replaced. Nor does it measure the fraction of a future model that was independently created. Understanding the number requires moving from company scale to the tasks that make up a working day.
First, define the work
Anthropic's approach breaks work into smaller activities and weights their contribution. Responsibility for a broad task is not equivalent to helping with a small step. Counting chatbot requests would consequently answer a different question from measuring how much work a system can lead.
Consider an independent example: writing a piece of code, diagnosing a failure and choosing the next experiment are separate activities. A model might rapidly propose code while leaving a person to decide whether the experiment makes sense. Grouping everything under development conceals those differences.
Epoch AI has also proposed a detailed taxonomy of AI research and development work. Constructing that taxonomy requires judgement. Changing the categories can change what appears to be progress before anyone replaces the model being assessed. Comparable measurements need a reasonably stable account of what researchers actually do.
Another useful measure is how often a person must intervene during a task. Two systems reaching the same answer can impose very different supervision burdens.
Autonomy is different from competence
A system offering a useful suggestion has a different role from one managing a sequence of work. The first may save someone a search or a drafting step. The second must choose an action, interpret the result and decide how to proceed. The same underlying model can occupy either role with different tools and permissions.
That makes the person's responsibilities worth examining. Do they specify the solution while the system implements it? Do they approve every action? Does the system work for an extended period and return a result for review? Calling all three arrangements autonomous would hide important operational differences.
Anthropic's findings did not include work at the highest, fully autonomous level. Human oversight remains part of the level at which AI leads a task. That defines the scope of the reported result and clarifies how responsibility is distributed in the research process.
What would count as measurable progress?
Tracking change requires comparing similar work. A month dominated by tasks well suited to a model cannot be directly compared with a month of unusually difficult work. Keeping a defined task basket helps separate improved capability from changes in what the system is asked to do.
Anthropic's figures come from its own organisation and assessment method. Task weights and autonomy ratings involve estimates. This is an internal operational measure with assumptions, rather than an industry wide standard. Readers can use it to understand a direction without treating it as a date for humans to leave the process.
The larger possibility is a feedback loop: a stronger system might contribute to research that produces a stronger successor. Establishing the speed of that loop requires evidence about task performance, research outcomes and the human time still spent reviewing results. Repeating a consistent measurement framework could reveal those changes more clearly than the broad claim that AI is already building itself.
