← Back

Hugging Face

companyCredibility: 65%

Why this score? Model hosting platform; authoritative for what is hosted/licensed on it, community-maintained model cards vary.

Tracked Statements (1)

This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure.?

Context: Hugging Face’s account of its own incident response, made while investigating a breach of its infrastructure. It does not name which providers blocked the requests, and no independent source confirms the refusals. Hugging Face frames it as an asymmetry — the attacker “was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried” — and states explicitly that “This is not an argument against safety measures on hosted models.” A checkable claim in principle: the named model (GLM 5.2), the named task (analysis of an attacker log of more than 17,000 recorded events) and the named failure mode are all specific enough for a provider or a later evaluation to confirm or refute. Hugging Face’s forensic timeline of 27 July names the refusing models for the first time and quantifies what the substitute recovered: “The models we reached for first, Claude Opus and Fable, refused a large part of that work: their safety guardrails treated reverse-engineering an exploit the same as launching one.” Both are Anthropic models, and are recorded here plainly. The replacement is specified as the Nvidia-quantized build of Z.ai’s open-weight GLM-5.2 run on Hugging Face’s own inference infrastructure, and replicating the attacker’s chunk, XOR and compress scheme with it recovered roughly four times more secrets than the first automated scan, mostly JWTs and platform tokens. More specific, but not more independent: it is the same party’s account of its own response, and the SANS write-up repeating it is by an author who discloses that he co-authored the Cloud Security Alliance post-mortem, which was reviewed by the team that lived the incident. No provider has answered on the record. The verdict is unchanged.