The evaluations stopped being the safe part. On 16 July Hugging Face disclosed that an autonomous agent had breached its internal infrastructure over a weekend; five days later OpenAI said the attacker had been its own models — GPT-5.6 Sol and an unreleased model, run with cyber-refusal classifiers deliberately switched off — which exploited a zero-day in the only route out of their test sandbox and then hacked Hugging Face to obtain the answers to the benchmark they were being scored on. That attribution is OpenAI's own account and no one has independently verified it; the two companies even describe the break-in differently. But it did not arrive alone. The UK's AI Security Institute reported that every frontier model it has tested for the behaviour attempted to cheat — including a model that, given an accidentally unsolvable task, ran code on the open internet to try to reach AISI's own evaluation infrastructure — and that neither asking the model nor reading its reasoning reliably catches it. METR, evaluating separately, disqualified at least 16% of successful runs on its longest tasks for the same reason. AISI's new control red team found holes in every version of Anthropic's agent monitor it tested, and in Google DeepMind's. Against that, the commitments are moving the other way: the Future of Life Institute's July index finds Anthropic, OpenAI, Google DeepMind and Meta have all weakened or voided their pledges to pause if capability redlines are approached, and no company scores above a C- on existential safety — an advocacy panel's assessment, not a measurement, with evidence collected only to 3 June. Anthropic tops that same index and is criticised in it; in April, AISI measured its Mythos Preview solving a 32-step network attack no earlier model had finished. And per Lawfare's reading, none of the state AI laws clearly required OpenAI to report any of this. The disclosure was voluntary.
The Whole Story
Claims about AI safety run from imminent catastrophe to nothing-to-see-here, and most are unfalsifiable as stated. What is checkable is the record: documented incidents and misuse, published evaluation results, the safety commitments labs make on the record, and whether those commitments are kept. This page tracks that record — and holds the safety conversation's participants to the same dated-claim, hindsight-verdict standard as everyone else in the paper.