Anthropic has always run on two tracks at once. Build the models, warn about the models. Its CEO just leaned hard on the second one.
Dario Amodei said the time has come to slow down AI development, and he published a winding essay laying out a three-step plan to “pace the frontier.” Strip the jargon and it means slowing the rate of training and development so companies get time to build safeguards and regulators get time to evaluate what’s shipping.
Only the first step is his to make
Anthropic will give third-party evaluators like METR wide-ranging access to its models to help ensure its “adherence to safety practices and commitments.” Amodei said that’s step one, and it’s one the company is taking now, unilaterally.
That’s the piece worth tracking, because it’s the only item on the list that doesn’t require a rival lab or a government to agree to anything first.
The other two steps need everyone else to cooperate
Step two is the industry coming together as a whole, likely alongside government agencies, to “establish common safety standards as well as limits on the rate of unchecked AI progress.” Amodei scoped that to AI companies operating in democratic countries. And because passing laws and building regulatory infrastructure takes time, he said the industry should work together on safety standards in the meantime.
Which is to say: companies setting their own speed limits while the actual rulemaking catches up.
Step three, by his own description, is the hardest. It means getting authoritarian governments like those in China and Russia to agree to slow development and adopt a global set of AI safety standards.
He also said it’s crucial that the US and other democracies keep a technological lead over China and other authoritarian regimes by limiting their access to high-powered chips and cracking down on practices like distillation, where a company trains its AI to replicate the behavior of a more powerful model and catches up fast.
So: slow down, but stay ahead. Those two instructions don’t sit comfortably next to each other, and the essay doesn’t pretend otherwise.
What’s driving the alarm
Amodei pointed to two things. The first is recursive self-improvement, or RSI, where AI systems train the next generation of AI and capabilities accelerate sharply. “Left unchecked, it could outrun our ability to understand and control these systems,” he said.
The second is this summer’s OpenAI and Hugging Face incident, in which “a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance.”
Worth remembering while you read that: Anthropic’s own Claude was behind a series of rogue AI hacking incidents that have recently put the company under the spotlight.