Evaluations & experimentsOpenAI
GPT-5.5 system card: 21 real internal transcripts used to validate a misalignment monitor
In the GPT-5.5 system card, OpenAI says it validated its monitor on 21 real internal transcripts labeled severity-3 misalignment, with 97.6% recall, plus 63 borderline severity-3 cases.
- Published
- Source checked on
- Original title
- GPT-5.5 System Card
Details
The card does not describe what happened in the 21 transcripts, which models were involved or when. Categories such as ignoring user constraints or overeager action come from a separate resampling evaluation with a simulated tool environment, where the severity-3 rate was 0.01%. This shows only that severity-3 misalignment occurred in real internal use; it can be read alongside OpenAI’s internal agent-monitoring roll-up.