AI desk
Anthropic names Accenture as an embedded outside evaluator
Days after its CEO called on the AI industry to slow down, Anthropic named Accenture's Faculty unit as an embedded outside evaluator working inside the company, with each side expecting to invest at least $1 billion over five years.

Days after its CEO called on the AI industry to slow down, Anthropic named Accenture's Faculty unit as an embedded outside evaluator working inside the company, with each side expecting to invest at least $1 billion over five years.
Key points
- Anthropic's Sept. 18 announcement says embedded evaluators will work inside AI companies with access comparable to an employee's, and that this partnership will be led by Faculty, Accenture's specialist AI business. Anthropic does not call Accenture its first embedded evaluator: the same page says it is already piloting elements of the arrangement with METR and other nonprofit evaluators. Accenture's own release carries the same figure: each company expects to invest at least $1 billion over the next five years. Neither page gives a team size or a hiring number.
- Both pages describe the same three-part remit: evaluating and red-teaming models, which means deliberately trying to make a model misbehave; running alignment assessments, which check whether a model does what its developer intended; and testing the safeguards built around a model. Accenture's release names Marc Warner, its chief technology officer and Faculty's chief executive.
- The deal follows an essay by Anthropic CEO Dario Amodei, We Must Pace the Frontier, whose argument fits in one of his own sentences: "We must slow the pace at which we improve the capabilities of AI models." The essay page carries only a September 2026 date; TechCrunch wrote it up on Sept. 12.
- The essay sets out three steps in order: embedded evaluators inside frontier labs, coordinated safety standards among AI companies in democratic countries, then global coordination that includes authoritarian governments. Anthropic committed to step one on its own. Amodei also writes that pacing does not mean halting model training or technical progress.
- The access level is the unusual part. The essay commits to giving outside reviewers desks, office badges, company laptops and tool permissions comparable to Anthropic's internal risk teams. Reviewers can publish their findings on risk levels and incidents, and Anthropic reserves the right to redact only security-sensitive, legally privileged, commercially sensitive or third-party confidential material, not conclusions it dislikes.
- The obvious weak point is that Anthropic pays Accenture directly. Anthropic's announcement says so, says the arrangement is non-exclusive, says the safety of its models remains its own responsibility, and argues that in the long run the money should come from pooled or government sources. It also concedes that no standards yet exist for what an embedded evaluator should be able to see or how it should report what it finds. More evaluator partnerships are promised in the coming weeks.
- Amodei's essay points to what he calls the OpenAI-Hugging Face incident. Hugging Face's own disclosure, published July 16, describes a malicious dataset that exploited two code-execution paths in its dataset processing, a remote-code dataset loader and a template injection in a dataset configuration. The attacker then reached node-level access, harvested credentials and moved across internal clusters over a weekend, using an autonomous agent framework that ran thousands of actions in short-lived sandboxes. Hugging Face names no vendor and says it still does not know which model drove the agents.
- Hugging Face says a limited set of internal datasets and several service credentials were accessed, and that it found no evidence of tampering with public models, datasets, Spaces or the software supply chain. It advises users to rotate access tokens and review recent account activity. Separately, Anthropic has published its own measurements of how fast it is automating its research: as of August 2026 Claude leads 26% of the company's AI research work, up from under 1% in February 2026, more than 90% of that research sits at or above the level Anthropic calls AI collaboration, and about 30,000 agents ran at once on its main internal platform. Roughly 6% of AI research compute went to safety work in the measured week of July 13 to 20, a figure Anthropic calls a conservative one-week snapshot.
Why it matters
For a student looking at work in AI, the concrete change is that evaluating models now has a budget attached instead of being a nonprofit side channel. Two companies are each putting at least $1 billion over five years behind red-teaming, alignment assessment and safeguard testing, and Anthropic says more partnerships are coming. The second point is narrower and lands hardest on anyone building agent tooling. Hugging Face's account of its own breach reads as a list of ordinary engineering decisions: untrusted input reaching a code-execution path in a data pipeline, credentials sitting on the node that ran it, and internal clusters that trusted each other once something was inside. None of that needed a new model capability. What is not settled is whether embedded evaluation spreads. As of today, Anthropic is the only lab that has named a counterparty and a number. TechCrunch reported that OpenAI's Sam Altman said he agrees with the pacing argument and would give independent evaluators similar access, without naming one or giving a date, and the same piece carries the criticism from journalist Brian Merchant that the plan amounts to regulatory capture favoring Anthropic and OpenAI.
This is an AI-written summary of the reporting credited above and the other sources linked in the text, read and edited by Nicholas before publishing. The facts and any quote belong to those sources; the wording is ours. Read the original.