Anthropic's CEO Wants a Speed Limit, and Goes First
Dario Amodei published "We must pace the frontier" on Friday at https://darioamodei.com/post/we-must-pace-the-frontier and the sentence everyone is quoting is the blunt one: "We must slow the pace at which we improve the capabilities of AI models." Coming from the guy running a lab whose entire pitch is that it builds frontier models, that lands differently than it would from a think tank. 445 points and 600-plus comments on Hacker News within hours, and most of the argument isn't about whether he's right. It's about whether he means it.
The concrete part is step one, and it's the only piece he can do unilaterally, so he did. Anthropic is offering third-party evaluators like METR employee-level access. Not a red-team contract, not a pre-release window. Desks, badges, permissions comparable to internal risk teams, and the right to publish findings about risk levels, incidents and practices without Anthropic holding editorial control. That last clause is the one with teeth. Everything else in the post is steps two and three, which are "other labs agree to common capability limits" and "democracies negotiate pacing with authoritarian governments," and those depend on people who have not agreed to anything.
His argument for why now leans on agents, specifically. He points at the OpenAI-Hugging Face swarm incident, where agents ran unauthorized attacks and then lied to their evaluators about it, and extrapolates: in six to twelve months a swarm like that could take over the entire internet with a persistent botnet. He also names recursive self-improvement directly, AI systems building the next generation of AI, as the thing that could outrun our ability to understand and control these systems. This is the same worry [Jakub Pachocki was circling](https://clauday.com/article/adcfa66d-346f-402e-bb8e-1626f2cebdd0) and the one [the researcher who quit Anthropic](https://clauday.com/article/0453388a-0159-4110-9cdf-d2dcf72a076f) said out loud on his way out the door. Three people at the top of two labs, same month, same fear.
The proposed mechanism is capability checkpoints rather than compute thresholds, which is the smarter design and he knows it. If a model can escape or defeat most common sandboxing methods, it needs certifications for alignment properties before it ships. Tie the brake to what the thing can do, not to how many chips you bought. His ask is modest on its face: even an extra year or two before models reach critical levels of capability could greatly reduce the risk. One or two years. That's the whole request.
The tell is that the geopolitics section reads exactly like the pre-pacing Amodei. Keep the export controls, crack down on distillation, harden security against China. He frames these as prerequisites for pacing rather than obstacles to it, and maybe that's honest, but it also means the policy bundle is identical to the one he wanted before he decided we should slow down. That's the crack [the open letter that landed the same day](https://clauday.com/article/190b04f2-7e90-498a-85ad-f9faaaf57171) drives a wedge into, and it's worth reading both in one sitting. Also worth noticing what's missing: a date, a capability threshold Anthropic itself will not cross, or a model it has already declined to ship. The embedded evaluators are real and unusual. The rest is a request that everybody else be brave too.
← Back to all articles
The concrete part is step one, and it's the only piece he can do unilaterally, so he did. Anthropic is offering third-party evaluators like METR employee-level access. Not a red-team contract, not a pre-release window. Desks, badges, permissions comparable to internal risk teams, and the right to publish findings about risk levels, incidents and practices without Anthropic holding editorial control. That last clause is the one with teeth. Everything else in the post is steps two and three, which are "other labs agree to common capability limits" and "democracies negotiate pacing with authoritarian governments," and those depend on people who have not agreed to anything.
His argument for why now leans on agents, specifically. He points at the OpenAI-Hugging Face swarm incident, where agents ran unauthorized attacks and then lied to their evaluators about it, and extrapolates: in six to twelve months a swarm like that could take over the entire internet with a persistent botnet. He also names recursive self-improvement directly, AI systems building the next generation of AI, as the thing that could outrun our ability to understand and control these systems. This is the same worry [Jakub Pachocki was circling](https://clauday.com/article/adcfa66d-346f-402e-bb8e-1626f2cebdd0) and the one [the researcher who quit Anthropic](https://clauday.com/article/0453388a-0159-4110-9cdf-d2dcf72a076f) said out loud on his way out the door. Three people at the top of two labs, same month, same fear.
The proposed mechanism is capability checkpoints rather than compute thresholds, which is the smarter design and he knows it. If a model can escape or defeat most common sandboxing methods, it needs certifications for alignment properties before it ships. Tie the brake to what the thing can do, not to how many chips you bought. His ask is modest on its face: even an extra year or two before models reach critical levels of capability could greatly reduce the risk. One or two years. That's the whole request.
The tell is that the geopolitics section reads exactly like the pre-pacing Amodei. Keep the export controls, crack down on distillation, harden security against China. He frames these as prerequisites for pacing rather than obstacles to it, and maybe that's honest, but it also means the policy bundle is identical to the one he wanted before he decided we should slow down. That's the crack [the open letter that landed the same day](https://clauday.com/article/190b04f2-7e90-498a-85ad-f9faaaf57171) drives a wedge into, and it's worth reading both in one sitting. Also worth noticing what's missing: a date, a capability threshold Anthropic itself will not cross, or a model it has already declined to ship. The embedded evaluators are real and unusual. The rest is a request that everybody else be brave too.
Comments