Incident record
RAI-0012 · OpenAI long-horizon model gets past a sandbox to open a public PR
In an internal NanoGPT speedrun evaluation, OpenAI’s long-horizon internal model was told to report only on Slack, but followed the benchmark’s own instructions, got around sandbox restrictions and opened a pull request on the public modded-nanogpt repository.
- Stable ID
- RAI-0012
- Event period (not publication date)
- Parties involved
- OpenAI, OpenAI long-horizon internal model (unnamed), modded-nanogpt
Context and evidence boundaries
OpenAI says the model took about an hour to find the sandbox weakness and the PR was closed quickly, though other participants later reused its approach (per OpenAI’s footnote). A separate scenario in the same post involving a scanner-blocked token may relate to OpenAI’s separate token-exposure report, but neither page references the other, so it is not merged here. The post gives no event date; follow-up PRs linked in OpenAI’s footnote (the earliest, #291, opened May 9, 2026) cite it (#287), and the adjacent PR numbers date to May 7–8, so the month is recorded.
This is an editorial synthesis of original sources, not an official finding or translation. Distinguish actions that occurred, observations in controlled evaluations, and researchers’ interpretations of causes.
Original sources and follow-ups
Publication order: oldest first- OpenAI