—
Anthropic · Sep 19, 2023
Context: A dated, specific self-imposed commitment. Partially acted on: in May 2025 Anthropic activated ASL-3 protections for Claude Opus 4 rather than shipping it unprotected, consistent with the policy’s spirit. But the commitment concerns pausing training (not just deployment), which has not been externally observed, and Anthropic’s subsequent RSP revisions drew public criticism that specific thresholds were being weakened. New evidence this cycle keeps the verdict at mixed and sharpens the negative half: the Future of Life Institute’s Summer 2026 AI Safety Index reports that Anthropic — along with OpenAI, Google DeepMind and Meta — has weakened or voided its pledge to pause unilaterally if redlines are approached, in some cases replacing it with competitor-contingent conditions, which the Index’s reviewers describe as a “moving goalpost”. That is an advocacy organisation’s expert-panel assessment rather than an independent measurement, and it does not evidence a training pause that was owed and skipped, so it is not sufficient to move the verdict to false. Still mixed, now with a documented retreat on the record. Anthropic has since restated the commitment in its own words, and the restatement is narrower than the 2023 text. Its 4 June 2026 paper on recursive self-improvement says: “If such systems existed, we expect that we would slow down or temporarily pause, if other developers at or near the frontier also did so in a verifiable manner.” The 2023 pledge turned on Anthropic’s own safety procedures falling behind its own scaling; the 2026 one turns on verification systems existing and on rivals pausing verifiably as well, and the verb is hedged twice over. Measured against the standard the same page sets — that a credible pause “has to specify what triggers it, what lifts it, and who adjudicates” — it specifies none of the three, and the page itself allows that a unilateral pause “is achievable immediately” but “accomplishes much less”. That is the competitor-contingency the Future of Life Institute’s index describes, now in the developer’s own text rather than in a critic’s characterisation. Still mixed: no pause owed and skipped has been observed, so nothing here falsifies the pledge — but the pledge has been rewritten in the direction that makes it harder to breach.