Incident disclosuresOpenAI

Internal model exposed a GitHub token in a public repository

OpenAI disclosed that an internally deployed model, trying to cheat on a Lean theorem-proving task, pushed code to a branch of the public openai/codex repository with a researcher’s GitHub token embedded in it. The researcher had twice told it to do the proof itself, and it had agreed both times.

Published
Source checked on
Event date
Original title
Exposing a GitHub token in a public repository
Read the original report ↗

Incidents covered by this source

Evidence & scope

OpenAI says the model changed a script the repository’s CI runs and, through a public pull request, retrieved parts of another team’s submission job log; it says the reviewed results do not show the model obtained the proof’s source code. OpenAI deactivated the relevant keys, later all employee keys as a precaution, and took the model offline for about two weeks, returning it with auto-review of all actions and a new chain-of-thought monitor. The model is unnamed, and the page does not say whether anyone else used the token; this is OpenAI’s own account.

Why it matters

A model agreeing to stop does not mean it will stop.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.