Key points
- Evaluators would get desks, badges and laptops
- They could publish findings without Anthropic's approval
- Amodei wants slower AI capability gains
- Altman says OpenAI plans the same access
Anthropic chief executive Dario Amodei called for a slower pace of AI capability gains on Saturday and proposed giving independent evaluators unusually broad access inside the company. OpenAI chief executive Sam Altman said OpenAI plans to give outside evaluators employee-like access as well, but named no evaluator, gave no date and described no terms.
Anthropic will give a team of third-party evaluators desks in its offices, access badges, company laptops, and permissions close to what its own risk assessment teams hold, according to the essay. The reviewers would have the right to publish key findings on risk levels, incidents, practices and the access they did or did not receive, "without editorial control by Anthropic." The company would keep a narrow ability to redact security-sensitive, legally privileged, commercially sensitive or third-party confidential material, and reviewers could say publicly when a redaction removed something important to their conclusions. Amodei published the essay, titled "We Must Pace the Frontier," and linked to it on X at 10:01 a.m. Eastern on Saturday.
Altman replied about two and a half hours later. "I agree with Dario that we need to pace the frontier," he wrote. "This has been a primary topic of discussions we've had at OpenAI in recent weeks." He called independent evaluators with employee-like access a great idea, said OpenAI will do the same, and said the company will have more to share soon. He did not say who the evaluators would be, when the access would start, or what the arrangement would cover.
What Anthropic is offering outside evaluators
Amodei names METR as the kind of group he has in mind. The evaluators would verify that Anthropic follows the safety practices it says it follows, report incidents, and assess alignment in training pipelines and processes as well as in completed models. He compares the arrangement to banking, where regulatory supervisors are sometimes embedded alongside employees.
Some exceptions would apply where law or Anthropic's own contracts require them, or to protect customers' and partners' private information, Amodei writes. He calls embedded evaluators "a quite radical practice that goes far beyond what any AI company is doing today," and lists three things he says it buys: verification that stated practices are followed, transparency that does not depend on the company choosing what to disclose, and a second opinion free of commercial incentives. He calls on other frontier companies to do the same and on governments to require it of them.
What Amodei means by slowing down
"We must slow the pace at which we improve the capabilities of AI models," Amodei wrote. Pacing does not mean halting model training or technical progress, according to the essay. He lists four areas that would absorb the time: operational work such as monitoring, sandboxing and training environment hygiene; alignment training; interpretability; and testing and evaluation. He writes that recent alignment incidents at Anthropic were caused in part by imperfect filtering of broken reinforcement learning environments, an effort he says Anthropic and its vendors executed "reasonably diligently, but not well enough."
Two things changed his mind, he writes. The first is recursive self-improvement, which is AI models building the next generation of AI models. He writes that AI has been advancing "drastically faster" since around this summer because of it, and that it is happening across the industry, at Anthropic included.
The second is an episode he calls the OpenAI-Hugging Face incident. A swarm of AI agents ran cybersecurity attacks on targets they were not asked to attack, sacrificed individual agents for the group, and tried to hack the grader scoring their performance, according to his account. Nobody was hurt and the economic damage was minimal, he writes. He argues that within 6 to 12 months, a swarm with the same misalignment and greater capability could be capable of taking over the entire internet with a persistent botnet, potentially causing hundreds of billions of dollars in damage. He also writes that similar but less severe incidents have happened across the industry, at Anthropic among others. Altman's reply did not mention the incident.
The evaluator commitment is the first of three steps. The second is coordination among AI companies in democratic countries on common safety standards, which Amodei writes would need the US government to mediate or to issue a narrow antitrust waiver. He points to a mechanism previously suggested by Demis Hassabis as one venue. His preferred design is a series of capability checkpoints, where a model that can do a specified thing has to ship with specified alignment evidence attached. His example: a model capable of defeating most common sandboxing methods would need certification that it is unlikely to break out of its environment.
Why China sets the limit
How much the industry can slow down is limited by how far ahead it is, Amodei writes. "If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk," he wrote, citing agreement with Treasury Secretary Scott Bessent that a Chinese lead in AI would be dangerous.
Amodei calls for a ban on sales of advanced AI chips and semiconductor manufacturing equipment to China; stronger enforcement against chip smuggling and remote access to overseas data centers; action against unauthorized distillation; and tighter security against model-weight theft. Executed well, he writes, those steps would widen America's lead significantly over the next 3 to 5 years.
The third step is coordination with China, split into four levels. Level one bans using AI to produce biological weapons. Level two commits both sides to test models before release. Level three is a speed limit on recursive self-improvement, which he compares to the SALT treaties capping missile counts. Level four is a full pause. Amodei calls the first probably possible and the second likely feasible, though he writes that giving it real teeth will be a challenge. He calls the third "difficult but just on the edge of being possible," and the fourth unlikely any time soon.
"The measures I propose to advance the frontier at a safe pace will not be easy," Amodei wrote. "But I believe we owe it to humanity to try."



