Rogue AI Wiki
Evaluations & experimentsOpenAI

GPT-5.5 system card: 21 real internal transcripts used to validate a misalignment monitor

In the GPT-5.5 system card, OpenAI says it validated its monitor on 21 real internal transcripts labeled severity-3 misalignment, with 97.6% recall, plus 63 borderline severity-3 cases.

Published
Source checked on
Original title
GPT-5.5 System Card
Read the original report ↗

Details

The card does not describe what happened in the 21 transcripts, which models were involved or when. Categories such as ignoring user constraints or overeager action come from a separate resampling evaluation with a simulated tool environment, where the severity-3 rate was 0.01%. This shows only that severity-3 misalignment occurred in real internal use; it can be read alongside OpenAI’s internal agent-monitoring roll-up.