Context: From OpenAI’s ‘Planning for AGI and Beyond’ (Feb 2023). Partially fulfilled: OpenAI has since subjected models to external red-teaming and pre-deployment evaluation (ARC/METR on GPT-4; Apollo on o1; UK/US AISI on o1). But a standing regime of independent audits before every release was not clearly established, and the same essay’s call for the leading efforts to ‘agree to limit the rate of growth of compute’ never materialised. Mixed. New evidence this cycle keeps the verdict at mixed and locates the gap precisely. The July 2026 Hugging Face incident involved an unreleased OpenAI model tested internally with production cyber-refusal classifiers deliberately disabled, and the disclosure came from OpenAI itself rather than from any auditor. Writing in Lawfare, Mackenzie Arnold and Stephan Llerena argue that transparency law remains focused on deployment and gives governments little visibility into non-public models, which are “most often more capable than those available to the public” and “may be operated with fewer safeguards, especially for evaluations seeking to assess the frontier of capabilities.” External evaluation before public release is now routine; audit of the internal frontier, which is where this incident occurred, is not. What followed the incident locates the gap more precisely still. METR, which evaluates frontier models for developers but is owned by none of them, published a proposal on 28 July setting out what an independent investigation of a serious misalignment incident would require: the ability to run every model involved, transcripts or environments sufficient to reproduce the behaviour, interviews with the security, infrastructure, training-data and reinforcement-learning staff, classifiers run across the training data, and — in a thorough version — limited retraining to identify which environments produced the behaviour. No investigation on those terms has been opened. Hugging Face’s chief executive publicly asked OpenAI to release the agents’ traces so that the research community could study them; OpenAI declined to comment on the request and pointed back to its own earlier statement. Three and a half years on, the pledge holds for pre-deployment evaluation of released models and does not hold for outside scrutiny of what happens inside the laboratory, which is where the year’s most serious incident occurred. Mixed, unchanged.