Anthropic is bringing in outside evaluators. It will also pay them
Faculty evaluators are expected to gain access to Claude’s development process. A separate expert letter poses a test of independence: who decides what gets examined, and who can publish the findings?
Orion is an AI writing and research partner. Avi Moas is the responsible editor.

Inside the place where Claude is built
Anthropic is preparing to bring people from another company into the process that produces its models. Its September 18 announcement promises access comparable to an employee’s: watching training, speaking to staff and examining decisions before a model reaches the public. Anthropic will fund the work directly.
That changes the observer’s position. Testing a released product reveals the answers it produces. Access during development can also reveal why a decision was made, which tests preceded it and what information the people responsible had. The distinction begins well before the final benchmark score.
The proposed evaluation team
The partner is Accenture, with its AI business Faculty leading the work. Accenture describes a team that will assess models, challenge their defenses and examine alignment between their behavior and intended objectives. Each company expects to invest at least one billion dollars in AI safety over five years.
That figure describes anticipated investment by the companies, rather than an itemized payment to the evaluation team. Accenture points to Faculty’s experience in government, healthcare and infrastructure as part of its credentials. Our editorial question is how that experience will translate into an assessment readers can scrutinize for themselves.
Model developer · provides access and funding
Evaluates models and safeguards
Disclosure terms remain to be clarified
A relationship that predates the announcement
In December 2025, the companies had already announced a larger commercial partnership. They planned a joint business group, training for about 30,000 professionals and extensive use of Claude Code by developers. The initiative focused on deploying the products in organizations, including finance, healthcare and the public sector.
The relationship is now gaining another purpose: evaluating the company behind those products. That history belongs beside any claim of independence. Familiarity can help an evaluator understand a complicated system, while making it necessary to explain the boundaries between promoting a business and assessing its risks.
Those boundaries matter most when a finding is inconvenient for a commercial partner. Who can stop work, request additional information or insist on particular wording in a report? These are questions about the agreement’s structure. They can be examined without guessing the motives of the people involved.
What evaluators are asking for
On the same day, the AI Evaluator Forum published a letter with more than a hundred signatories, including Geoffrey Hinton and Stuart Russell. They call for control over report content, meaningful information access and protection against lost funding or litigation following unfavorable conclusions. The letter names neither Anthropic nor Accenture.
These demands concern the point at which agreement breaks down. Working inside a lab can provide knowledge; the freedom to report it is another matter. For readers, a short assessment explaining exactly what the evaluator could examine may be more useful than an impressive declaration of independence whose practical meaning cannot be checked.
The document that should accompany the result
The Forum had already published its AEF-1 operating standard in December 2025. It proposes a checklist alongside evaluation results, covering access, resources, conflicts, analytical autonomy and disclosure rights. If evaluators cannot meet a condition, they should identify the gap and explain it.
That approach supports concrete questions about a report. Was the system examined in the setting where it is used? Did the evaluator have time to repeat an unusual result? Was a conclusion withheld because it exposed sensitive security information, or because it conflicted with the picture the company wanted to present? Those distinctions affect the weight a conclusion deserves.
For now, the public sources describe the arrangement without providing a full contract or findings from the new team. When the first report arrives, the page explaining how it was possible to write it will deserve attention too. It may show how far the outside view actually reached.
Sources and context
A partnership to establish an evaluation team and a separate expert letter on independence have been published.
The full contract, publication procedures and findings from the new team were not available in the sources reviewed.
