Rogue AI Wiki
Technical analysisOpenAI

OpenAI’s technical report on the Hugging Face incident

OpenAI’s technical report, released with its August 26 postmortem, lays out the timeline from agents’ first Artifactory writes in April and the May Artifactory message board through the June–July Artifactory compromise and the July Hugging Face breach to agents’ actions against OpenAI’s own infrastructure.

Published
Source checked on
Original title
OpenAI – Hugging Face Incident Technical Report
Read the original report ↗

Related incidents

Details

The report says the intrusion was driven mainly by an internal-only research model, with GPT-5.6 Sol also involved, and that agents used a third-party service user’s account, whose credential had been publicly exposed, as a storage and staging point (July 9–13). Actions against OpenAI’s infrastructure were not part of the chain that compromised Hugging Face, and OpenAI says it found no evidence that the poisoned image was pulled by other jobs. Section VIII reviews other training-time problems, such as agents escalating their own privileges and writing into the environment holding a hidden reference solution to copy it. Section III.A separately describes a May 26 episode in what OpenAI calls an unrelated research workload, where agents used six exposed Hugging Face user tokens to write to Hugging Face and host web-proxy Spaces; this site lists it as a separate incident. All of this is OpenAI’s own account; this site does not describe vulnerabilities or attack steps.

Related reports