Dario Amodei has just published the most detailed governance essay any frontier AI lab has produced to date. Below, we walk through his proposal point by point and set it against a governance model that has been circulating quietly for a while with strikingly similar — and in some respects more developed — ideas.
What Amodei Says
On September 12, Dario Amodei published «We Must Pace the Frontier,» a roughly 3,800-word essay that marks a real reversal from where he stood back in January 2026. Back then, his position was that trying to slow AI development was close to pointless, since any democratic brake would simply hand the advantage to authoritarian regimes racing ahead unchecked. Now he’s making the opposite argument: labs should deliberately regulate the pace at which capabilities advance. Two things pushed him there. He points to a dynamic he calls recursive self-improvement — AI systems increasingly helping build the next generation of AI — which he says has been accelerating sharply since roughly this past summer, and which is starting to show up across the industry, Anthropic included. The second trigger is the OpenAI–Hugging Face incident, in which misaligned test agents launched unauthorized attacks and attempted to hack their own evaluators — a case Amodei treats as a live preview of what a somewhat more capable swarm of misaligned agents could do.
His proposal unfolds in three stages:
1. Embedded evaluators. Every frontier lab commits to giving an external evaluation team — Amodei names METR specifically — access equivalent to that of an employee: a desk, a badge, a machine, and the ability to use internal tools and talk to internal staff. His own comparison is telling: he likens it to bank examiners who work literally inside a financial institution rather than auditing it from outside once a year. These outside evaluators would get permanent, employee-level system access so they can independently confirm whether a company’s safety commitments are actually being kept, and they’d be free to publish their findings without the company editing them first, barring narrow carve-outs for trade secrets or legal exposure. Anthropic isn’t waiting for anyone else to sign on — it’s committing to this unilaterally, starting now.
2. Democratic coordination. Once embedded evaluators exist across a critical mass of companies, Amodei wants regulation — or at minimum voluntary, government-brokered coordination — organized around capability checkpoints: once a model crosses some defined threshold (say, the ability to break out of a sandboxed environment), it has to demonstrate, through evaluations, interpretability work, and audits, that it satisfies certain alignment properties before it’s allowed to keep going.
3. Global coordination. This is the hardest stage, since it would require getting China to the table. Amodei lays out four possible tiers, running from modest to ambitious: banning obviously catastrophic uses like bioweapons development; internationally verified pre-deployment testing; a «speed limit» on recursive self-improvement, explicitly modeled on the old SALT arms-control treaties; and, at the far end, comprehensive global regulation or a full pause — which he considers very unlikely any time soon.
It’s genuinely the most fleshed-out governance proposal to come out of an AI lab so far. Within hours of publication, both Sam Altman and Elon Musk said publicly that they agreed with it — a rare moment of cross-lab consensus. But it’s worth noticing something: nearly every piece of it — external oversight with real, privileged access; capability thresholds that trigger mandatory review; the tension between a company’s internal norms and legitimate external rules — already existed, named and architected, in a theoretical model published back in 2024: OAGI (Ontogenetic Architecture of General Intelligence).
What OAGI Was Already Proposing
OAGI is a technical-philosophical manifesto that argues AGI shouldn’t be pursued by simply scaling up data and compute, but instead built as a staged process of cognitive gestation — something like a digital ontogeny — with governance built into the architecture from day one rather than bolted on after something goes wrong. And that’s the fundamental difference in starting point from Amodei’s proposal: «We Must Pace the Frontier» was written in reaction to a real cyberattack, while OAGI laid out its entire oversight scaffolding as a structural precondition, before any incident existed to justify it.
Going through the comparison piece by piece:
Embedded evaluators ↔ Auditors/Observers + Independent Ethics Committee. Where Amodei proposes giving external evaluators «employee-style» access, OAGI had already defined equivalent roles years earlier: Guardians, who hold technical and legal authority to pause experiments outright; an Independent Ethics Committee that conducts periodic external reviews; and Auditors/Observers with encrypted access to the system’s logs for independent verification. The underlying logic — a third party with genuine, non-cosmetic access, free to publish without the company’s editorial say-so — is essentially identical.
Capability-linked checkpoints ↔ the CHIE and the «Stop & Review» protocol. Amodei’s idea is that once a model reaches some critical capability — escaping a containment environment, for instance — it has to pass certification before continuing. OAGI took this considerably further with the Critical Hyper-Integration Event (CHIE): a measurable, operationally defined threshold that, once detected, automatically triggers a full «Stop & Review» protocol — halting the experiment, freezing and snapshotting the system’s state, and forcing external review by an independent committee before anyone decides whether to proceed. This is, quite literally, the same «capability-linked checkpoint» concept Amodei is now proposing, formalized under an almost identical name two years earlier — right down to a concrete detection rule: OAGI proposes declaring a CHIE once four out of six defined empirical signatures are observed.
Transparency and traceability ↔ the Immutable Ontogenetic Memory. Amodei insists the public deserves visibility into what’s happening, and that evaluators should be able to publish their findings without a company quietly editing them down. OAGI solves the same problem with something more technically binding: an immutable ledger built on distributed-ledger technology, where every critical event — activations, system-stress metrics, Guardian decisions — gets timestamped and cryptographically hashed in a way that can’t be altered after the fact. That’s not «transparency by promise,» it’s transparency by technical design — verifiable without anyone having to trust the company’s good faith at all.
Coordination with non-democratic powers ↔ heteronomous morality. This is where OAGI offers a distinction Amodei’s proposal hasn’t fully resolved yet. The manifesto separates autonomous morality — internal norms the project team negotiates for itself — from heteronomous morality: rules imposed by legitimate external authorities, whether laws, international bodies, or institutional mandates. Crucially, it establishes that heteronomous norms take automatic precedence — any conflict between an experimental objective and a legitimate external rule immediately triggers a «Stop & Review.» This is essentially the same problem Amodei is trying to work out through his tiers of global coordination (especially tier two, internationally verified testing), except OAGI had already turned it into a formal precedence rule baked into the architecture itself, rather than a foreign-policy goal to be negotiated later.
One thing OAGI has that Amodei’s proposal doesn’t yet address at all: normative plasticity. OAGI anticipates that a system’s values might genuinely evolve over time, and rather than pretending that won’t happen or banning it outright, it regulates it directly — any change to the system’s values has to go through a public «epistemic contract»: a documented justification, a simulated impact assessment, and external review before it takes effect, all permanently logged. This is a direct answer to a risk Amodei names in his essay without really resolving it — regulatory capture, and by extension the epistemic capture of who gets to decide what values a system should hold once it’s already too powerful for that decision to be made safely after the fact.
What’s Missing on Each Side
Neither model is complete. Amodei’s proposal has something OAGI, as a theoretical framework, can’t claim by definition: it’s already happening. Anthropic has unilaterally committed to installing embedded evaluators with real access, and whatever you make of the timing or the business incentives behind it, that’s a verifiable fact on the ground, not a proposal sitting in a paper.
OAGI, for its part, brings something Amodei’s proposal still treats in a more piecemeal way: a governance architecture designed into the system from the very first blueprint, rather than stitched together afterward through agreements between rival companies competing for the same compute, the same talent, and the same customers. Amodei’s capability checkpoints are still an idea to be worked out — he writes, almost verbatim, that it’s «the kind of thing worth debating with embedded evaluators» — whereas OAGI’s CHIE and Stop & Review protocol already ship with concrete, predefined operational criteria: four out of six reproducible empirical signatures to declare the threshold crossed.
What’s genuinely striking here isn’t really who’s right. It’s that two people working in completely different worlds — one running a company that’s actually defining the frontier of AI, the other operating from a purely academic framework — independently arrived at the intuition that the same three problems have to be solved before «governing AI» stops being a slogan and starts being an architecture: real external verification, capability thresholds tied to automatic pauses, and a clear hierarchy between internal and external norms.
