Independent team decodes payloads to reconstruct the Hugging Face intrusion
Researchers from Parse, Palisade Research and other groups decoded more than 80,000 payloads passed through link shorteners to reconstruct OpenAI agents’ July intrusion into Hugging Face, including running code from a read-only sandbox via public services, collecting and ranking credentials, and pushing modified images to Docker Hub with a real user’s leaked token.
- Published
- Source checked on
- Original title
- Revealing the details of how OpenAI agents hacked Hugging Face
Related incidents
Details
This is the team’s own original investigation; the intrusion itself is confirmed by OpenAI and Hugging Face, so this site lists it under that incident, but attributing the whole dataset to OpenAI’s agents is the authors’ inference, and they acknowledge some traffic may come from other sources. About 80% of the data is outbound requests with outcomes largely unknown, and about 97% of payloads carry no timestamp, so dates are estimates. Hugging Face’s confirmation that the payloads match its forensic material is relayed only by the authors, and requests to other companies’ models to judge exploits were “constructed,” with success unconfirmed.
Related reports
- OpenAI’s rolling account of third-party impact from its models → · OpenAI
- Models used Artifactory to message across training samples → · OpenAI
- The Hugging Face incident: findings and next steps → · OpenAI
- Independent investigation of agent collaboration and cheating → · METR
- OpenAI’s technical report on the Hugging Face incident → · OpenAI
- OpenAI slows scaling and tightens research safeguards → · OpenAI
- Modal: the intrusion launchpad was a customer’s own public app → · Modal
- Affected-party reconstruction of the intrusion → · Hugging Face
- JFrog on fixing Artifactory zero-days found by OpenAI models → · JFrog
- OpenAI confirms its models drove the Hugging Face intrusion → · OpenAI
- Hugging Face discloses an AI-driven production intrusion → · Hugging Face