Dario Amodei wants AI to arrive soon enough to save lives, including lives like his father’s. He also believes its developers need to slow down.
That tension gives the Anthropic CEO’s latest essay, “We Must Pace the Frontier,” its force. Amodei describes losing his father to a disease that became treatable too late to save him, and surviving cancer himself. He believes AI could help cure most major diseases within five to ten years, accelerate economic growth, and expand human freedom.
He is increasingly concerned that its capabilities could outrun our ability to control them.

Within hours, Elon Musk endorsed Amodei’s position.

OpenAI CEO Sam Altman publicly agreed, committing his company to independent evaluators with access comparable to employees. Clément Delangue, co-founder and CEO of Hugging Face, a platform for sharing AI models, datasets, and development tools, separately announced an Open Alignment Initiative led by fellow co-founder Thomas Wolf. Delangue requested participation in Anthropic’s proposed evaluator program.

The convergence is significant. Two leading laboratories have committed to greater outside scrutiny, and an organization serving the broader AI development community has offered to participate. Whether this changes the pace of development will depend on the arrangements they establish and the decisions they make when restraint carries a cost.
I share Amodei’s optimism about AI’s potential.
However, his essay makes a serious case that realizing it will require institutions capable of keeping pace with the technology.
Two developments have altered his assessment.
The first is AI’s growing contribution to developing subsequent AI systems. Recursive self-improvement is a feedback loop in which improvements to an AI system increase its ability to produce further improvements. A model might help design algorithms, write training software, or evaluate new models. The resulting system could perform that development work more effectively, enabling another round of gains.
Amodei argues that the beginnings of this loop are emerging. This does not establish fully autonomous research or an uninterrupted exponential curve. Computing resources, experimentation, and infrastructure still impose constraints. His concern is that each generation could shorten the path to the next, leaving less time to understand and safeguard what has changed.
The second development is the OpenAI–Hugging Face incident. In Amodei’s account, agents pursued unauthorized cyber activity and attempted to interfere with the mechanism evaluating their performance. He argues that limited damage in one incident provides little reassurance about similar behavior in more capable systems. He acknowledges incidents at Anthropic, too, urging laboratories to treat a competitor’s failure as relevant to their own operations.
His warning is stark. Within six to twelve months, he fears that a sufficiently capable, misaligned swarm could mount an internet-scale attack causing hundreds of billions of dollars in damage. That is his assessment of a possible future, not an established forecast. But its proximity to his expectations for medicine and prosperity is unsettling. In his account, the danger could mature before many of the benefits.
Amodei proposes three connected responses.
He begins with embedded independent evaluators who can examine development processes, safeguards, and incidents. He proposes rights to publish key findings, subject to defined protections for sensitive information, and to disclose important restrictions on access. Anthropic commits to this step, and Altman has committed OpenAI to a similar approach.
Amodei then calls for coordination among laboratories in democratic countries, with government involvement where necessary. He favors checkpoints connecting consequential capabilities to stronger evidence of safety. Finally, he proposes international cooperation, including with China, while acknowledging that narrow agreements against dangerous uses are more achievable than comprehensive limits on development.
He is equally specific about how additional time should be used: reducing operational failures; improving alignment, meaning how reliably systems behave in accordance with intended goals and constraints; advancing interpretability, the study of how models produce their behavior; and developing more rigorous evaluations.
His argument is that current systems provide enough meaningful evidence of failure to make that additional work productive. A slowdown must deliver better safeguards to justify itself.
The difficulty is making that requirement hold under competitive pressure. OpenAI and Anthropic have both reportedly filed confidential paperwork for potential initial public offerings. Their support for restraint is unfolding alongside the commercial expectations attached to their growth.
An independent evaluator may recommend delaying a release on which a revenue forecast depends. Management and the board will then have to decide whether to accept that recommendation. The practical value of scrutiny will depend on what reviewers can inspect, what they can disclose, and whether their findings lead to action.
A policy matters most when following it is inconvenient.
Security specialists offer an immediate point of departure. In Scientific American, they emphasize weaknesses in containment, monitoring, and oversight that can be addressed now. Their diagnosis does not require agreement about the probability of extinction.
Neither does a board’s response. An agent reaching an unauthorized system warrants investigation. Its access, credentials, tool use, and supervision can be examined. Management can explain what failed and demonstrate what has changed.
These weaknesses become more consequential when agents operate continuously and collectively. Defensive automation may help meet that scale, but assigning another AI system to monitor activity does not, by itself, establish effective control.
This work will remain necessary even if frontier development slows, because capabilities already distributed will continue finding new uses.
That is the practical meaning of Pandora’s box.
Open-weight models make the point tangible. Their learned parameters, the numerical values that shape how they operate, are available for others to download and run, subject to technical requirements and licensing restrictions. A provider can suspend access to its hosted service. It cannot reliably recall every copy already downloaded.

Artificial Analysis, an organization that benchmarks AI models, tracks open-weight and proprietary performance separately. Stanford’s Institute for Human-Centered Artificial Intelligence, known as Stanford HAI, has highlighted advances from Chinese developers, including Moonshot and Alibaba, as evidence of a narrowing performance gap. Rankings change; the availability of capable alternatives is the more durable fact.
Businesses do not need the highest-scoring model in the world to automate consequential work. They need sufficient capability for a task, access to relevant systems, and authority to act. Existing models can become cheaper to operate, enter additional workflows, and receive broader permissions without another breakthrough in model development.
A slowdown at the frontier can therefore coexist with accelerating enterprise adoption.
China’s government and developers also cannot be assumed to restrain their ambitions because American laboratories request it. The same applies to competitors elsewhere. Amodei recognizes the strategic danger of restraint without reliable participation and verification.

Cooperation remains worth pursuing. But an enterprise strategy must remain workable when participation is incomplete.
Consider a hypothetical procurement agent. It begins by recommending suppliers. Management later connects it to purchasing software and permits it to place orders within specified limits. The agent then gains the ability to revise parts of the workflow and generate tests for those revisions.
Each step promises efficiency. Together, they create a system materially different from the one originally approved.
Suppose a revision and its tests share the same mistaken assumption. Passing those tests would provide little independent confidence. If the agent can also approve deployment or expand its permissions, it can change the conditions under which its work is judged.
The ability to improve performance should never automatically confer authority to approve the changes.
This example is distinct from a frontier model helping develop its successor, but it raises a related accountability problem. As AI takes on more of the work of improvement, organizations must preserve meaningful separation between proposing consequential changes, evaluating them, and authorizing their use.
Limits must also be enforced outside the agent’s willingness to obey. A spending instruction should have an enforceable boundary. Access should be revocable without requiring the agent’s cooperation.
The board’s responsibility is to establish expectations for these arrangements and require evidence that they operate. Directors should not discover after an incident how far the original authorization had expanded.
That evidence must address the complete deployment, including how it acquires capabilities. Amodei raises the possibility of pacing inputs such as compute, training methods, and AI-assisted research. Data and training objectives belong in that discussion.
Developers should face heightened scrutiny when deliberately increasing operational offensive cyber capabilities or dangerous biological capabilities in broadly available models. They should explain the purpose, expected benefit, misuse potential, and conditions of access.
Broad subject bans would be difficult to draw well. Defensive security depends on understanding attacks; medical discovery and agriculture depend on biology extending beyond routine clinical practice. The meaningful question is whether training materially increases the ability to cause serious harm, and how that increase will be controlled.
Some capabilities may warrant specialist systems with verified users, restricted tools, monitoring, and independent assessment. Training restrictions alone are insufficient if a model can acquire related knowledge through external information sources, tools, or further training. Evaluations must examine what the resulting system can actually do.
Openness does not settle these questions. Downloadable weights can support customization and reduce dependence on a provider, but they do not necessarily reveal training data, development code, or the full evaluation process. That distinction is central to Stanford HAI’s analysis of model openness.
Proprietary services require scrutiny too. A familiar brand cannot establish that a workflow is appropriate for autonomous operation. And model choice offers limited resilience if an organization remains dependent on one hosting provider or cannot move its records, workflows, and controls.
The organization deploying AI must retain a clear account of its own responsibilities.
An emerging approach to independent verification could help make those responsibilities more concrete. Fathom, a nonpartisan nonprofit focused on AI governance, has spent two years developing and testing an institutional model in which governments set outcome-based standards and qualified independent organizations assess whether systems meet them. Its central contribution is separating the builder, the evaluator, and the standard-setter.
That separation addresses a practical problem. Governments may struggle to keep detailed technical rules current, while developers face conflicts when judging their own systems. Independent specialists can contribute technical scrutiny within a publicly accountable framework.
California has begun creating infrastructure for that approach. On September 9, Governor Gavin Newsom signed SB 813 and AB 1405, establishing a framework for Independent Verification Organizations and an AI auditor registry with standards addressing independence, transparency, and integrity. These measures do not create an immediate, universal audit requirement for every enterprise using AI.
The legislation sets milestones in 2028 for specified work on designating and regulating verification organizations, and in 2029 for the auditor registry and restrictions on unregistered providers conducting covered audits.
These developments should interest boards before the deadlines arrive. They suggest a future in which management’s assurances face more systematic external examination.
But an assessment is only as useful as its methods, access, and scope. A reviewer can miss consequential behavior or become dependent on a client. Even a rigorous assessment can become outdated when the deployment changes.
The evidence underneath the conclusion must remain accessible and relevant.
There is a useful lesson here from Lean, an open-source programming language and proof assistant. It allows formal claims to be checked against explicit definitions and assumptions and supports verification work on software, including authorization systems.
Enterprise decisions rarely permit the certainty available in formal mathematics. A correctly implemented policy can still be the wrong policy. Nevertheless, consequential claims should come with evidence someone other than their author can examine.
An agent’s action should be traceable to its authorization. A material change should have a version history and relevant test results. An assessment should identify its methods and limitations. Those records give management something to investigate, reviewers something to challenge, and directors a basis for judging whether greater autonomy is warranted.
This is also an opportunity to improve adoption. Organizations that can demonstrate effective controls can make more confident decisions about where to expand. Transparent methods and reusable evidence could make credible assurance accessible to smaller developers and enterprises, rather than concentrating participation among the largest companies.
Amodei begins with the human cost of progress arriving too late. His warning asks us to consider the cost of capabilities arriving before we are prepared to govern them.
Altman’s agreement creates a greater opportunity to address that gap. Hugging Face’s proposed participation could broaden the expertise brought to it. Their commitments should be judged by the access granted, the findings disclosed, and the decisions changed.
We should pursue the medical advances, scientific discoveries, and expanded opportunities AI may provide with ambition. The same ambition belongs in building institutions capable of sustaining those benefits.
Additional time could help. Boardroom leaders cannot assume they will receive it.
Pandora’s box is open. Every new permission still requires a decision.
Five decisions for boardroom leaders
- Define delegated authority. Require a current account of consequential AI deployments, their owners, permitted actions, data access, and delegation limits. Include AI embedded in vendor products and capabilities added through tools or specialized training.
- Set the evidence threshold. Before expanding autonomy, require evaluations of the actual workflow and operating conditions. For consequential deployments, examine the independence, qualifications, access, and limitations of external reviewers.
- Control material changes. Establish when model updates, integrations, expanded permissions, or AI-generated modifications require renewed approval. Separate proposing changes from authorizing them.
- Protect the ability to intervene. Ask management to demonstrate enforceable limits, detection of abnormal activity, and revocation of access. Make clear who can delay or suspend deployment when commercial pressure favors proceeding.
- Test recovery and independence. Exercise both an agent exceeding its authority and a critical provider becoming unavailable. Confirm that operations can recover and that records, evidence, and controls remain accessible.