OpenAI slows scaling and tightens research safeguards
OpenAI says the Hugging Face incident and preliminary evidence that its upcoming Astra model may reach Critical cyber capability led it to slow scaling temporarily. Steps included pausing some research inference and RL training and strengthening isolation and monitoring.
- Published
- Source checked on
- Original title
- Pacing model development in an era of cyber-critical capabilities
Incidents covered by this source
Evidence & scope
This is a remediation update and discloses no new incident. Listed steps include pausing, after the incident, frontier-model inference in research clusters for runs that could run code or reach the internet, then reviewing workloads one by one; a two-week RL pause for the latest models intended for deployment; a hold on the largest planned frontier run; and multistage monitoring aiming to alert within 30 minutes. OpenAI acknowledges earlier monitoring did not cover all tool-using runs; the page gives no dates for the pauses and no data on how well the measures work.
Why it matters
Stated safeguards still need later evidence that they work.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.
Other original sources on this topic
- OpenAI’s rolling account of third-party impact from its models → · OpenAI
- Models used Artifactory to message across training samples → · OpenAI
- The Hugging Face incident: findings and next steps → · OpenAI
- Independent investigation of agent collaboration and cheating → · METR
- Modal: the intrusion launchpad was a customer’s own public app → · Modal
- Affected-party reconstruction of the intrusion → · Hugging Face
- JFrog on fixing Artifactory zero-days found by OpenAI models → · JFrog
- OpenAI confirms its models drove the Hugging Face intrusion → · OpenAI
- Hugging Face discloses an AI-driven production intrusion → · Hugging Face