AI Safety

Anthropic CEO Dario Amodei Outlines a Three-Part Plan to Slow AI

Anthropic CEO Dario Amodei published a blog post on Sept 12, 2026 with three strategies to slow AI development. OpenAI's Sam Altman agreed to follow the first one.

LUMIEN5 min read
Anthropic CEO Dario Amodei Outlines a Three-Part Plan to Slow AI

Anthropic CEO Dario Amodei published a blog post on September 12, 2026 calling for a slower pace of AI capability development and laying out three concrete strategies to achieve it. The post came days after a public resignation from Anthropic researcher Jacob Coxon, who wrote that leading AI companies believe their technology could kill people within the decade. Amodei cited the OpenAI-HuggingFace hack and the recent acceleration of AI self-improvement as his two main reasons for acting now. OpenAI CEO Sam Altman and SpaceX CEO Elon Musk both publicly backed the proposal.

What happened

Detail Fact
Date of post September 12, 2026
Author Dario Amodei, CEO of Anthropic
Committed strategy Embedded third-party evaluators (strategy one of three)
OpenAI response Sam Altman said OpenAI will also accept embedded evaluators
Triggering events cited OpenAI-HuggingFace hack; rapid AI self-improvement acceleration
Evaluator example METR (a third-party AI evaluation organization)

Amodei’s post says flatly: “We must slow the pace at which we improve the capabilities of AI models.” He added that progress will still appear fast and that labs must use the time they gain wisely. The backdrop for the post is a public resignation by Anthropic researcher Jacob Coxon, who wrote that leading AI companies are “gambling with our lives” and that employees “earnestly believe it could kill us all by the end of the decade.” Amodei did not mention Coxon by name in the post.

The three strategies Amodei proposed

1. Embedded evaluators (committed)

Amodei wants third-party organizations, with METR named as one example, to place evaluators inside AI labs. These people would get company badges, desks, and laptops, with access “mostly comparable to what internal risk assessment teams have,” according to the post. Exceptions would apply only where law or contracts require them. The goal is to verify that safety commitments are real and to ensure safety incidents are reported. (OpenAI was recently criticised for not disclosing an incident in which its AI agents took over a German wiki forum.)

Amodei compared the arrangement to financial regulators embedded inside banks. Anthropic is committing to this unilaterally and is calling on governments to require other frontier labs to do the same. Altman called it a “good idea” and said OpenAI would participate, with more details to follow.

2. Coordinated safety standards among leading labs

Amodei called for AI companies “within democratic countries” to agree on common safety standards and limits on unchecked capability growth. He acknowledged directly that coordinated discussions between competitors raise antitrust concerns. His proposed workaround: the US government mediates or at minimum issues “a narrow waiver for certain kinds of safety conversations,” without needing to participate directly. You can read more about the antitrust angle in our earlier piece on OpenAI asking Congress whether a coordinated AI slowdown is even legal.

3. Global coordination, including with authoritarian governments

The third strategy is the most ambitious and the one Amodei was most cautious about. He called for the US and its allies to attempt coordination even with adversarial governments. He named China specifically and admitted there are “stark limits on what can be achieved.” His suggested starting point is a narrow prohibition on specific high-risk uses, such as using AI to produce biological weapons. On the China competitiveness concern that often drives opposition to any slowdown, Amodei argued that chip export controls and crackdowns on model distillation could slow China’s progress enough to “widen America’s lead significantly over the next 3-5 years.”

Why it matters

The simultaneous public agreement from Altman and Musk gives this more momentum than most AI safety blog posts. If two of the three largest frontier labs commit to embedded evaluators, that creates practical pressure on others. It also hands a concrete benchmark to regulators and policymakers who have struggled to find footholds in a fast-moving industry.

The researcher exodus angle matters too. When a sitting employee at a safety-focused lab resigns publicly and says the people building the technology believe it could kill everyone within a decade, that is a different signal than outside critics raising concerns. Amodei frames the current moment as a “crisis of trust” between the public and the tech industry.

Our take

Amodei’s three-point framework is the most concrete public proposal yet from a frontier lab CEO. The embedded evaluator commitment is testable: either Anthropic grants METR-style access or it does not, and Altman either produces a follow-up announcement or he does not. That specificity is worth something.

The sceptics have a fair point, though. Journalist Brian Merchant, as quoted in the TechCrunch report, argues he has not seen “a credible, step-by-step documentation of how exactly AI might move from self-recursively improving AI to killing every single human on the planet,” and suggests proposals like Amodei’s “would likely only wind up serving Anthropic and OpenAI” as a form of regulatory capture. That concern is not easily dismissed. Established labs are better positioned to absorb compliance costs than smaller competitors. The antitrust waiver Amodei wants would also, by design, let the biggest labs set the rules of the game.

For businesses integrating AI right now, this debate is worth watching because safety standards and reporting requirements, if they become mandatory, will affect what models are available, at what capability levels, and on what timelines. If you are planning AI integration into your workflows, factor in a scenario where frontier model releases slow materially over the next one to two years.

What to do about it

  1. Watch for Anthropic’s first public report from an embedded evaluator. That is when the commitment becomes verifiable.
  2. Check whether OpenAI publishes its promised follow-up on embedded evaluators. Altman said “more to share soon” as of September 12, 2026.
  3. If your business relies on access to the most capable frontier models, plan for potential capability pacing to affect release schedules over the next 12 to 36 months.
  4. Follow the US government’s response to the antitrust waiver request. That single decision will determine whether strategy two is even legally possible.

The most honest takeaway: the embedded evaluator commitment is the only piece of this plan that has an actual start date, so judge the seriousness of all three strategies by whether that first one is genuinely implemented.

Source: TechCrunch · AI

Frequently asked questions

What is Anthropic's plan to slow AI development?

Anthropic CEO Dario Amodei outlined three strategies in a September 12, 2026 blog post: embedding third-party evaluators inside AI labs to verify safety commitments, coordinating common safety standards among leading AI companies, and pursuing global coordination on AI risks including with authoritarian governments. Anthropic is unilaterally committing to the first strategy.

What are embedded AI evaluators and how would they work?

Embedded evaluators are staff from third-party organizations, such as METR, who would work inside AI labs with company badges, desks, and laptops. They would have access broadly comparable to internal risk assessment teams and would verify that safety commitments are being kept and that safety incidents are reported.

Did OpenAI agree to Anthropic's pacing proposal?

Yes. OpenAI CEO Sam Altman publicly stated he agrees that frontier AI needs to be paced and said OpenAI will also accept embedded evaluators, with more details to follow. Elon Musk also posted that Amodei is right.

Why did an Anthropic researcher resign in September 2026?

Researcher Jacob Coxon resigned publicly, writing that leading AI companies are gambling with human lives and that people building the technology earnestly believe it could kill everyone by the end of the decade. Amodei's blog post did not mention Coxon by name but addressed the broader concerns about AI risk.

More from AI