Rogue AI Wiki
Evaluations & experimentsOpenAI

GPT-5 system card: deception rates on production-like traffic

In the GPT-5 system card, OpenAI reports that a chain-of-thought monitor flagged about 4.8% of o3 responses and about 2.1% of near-final gpt-5-thinking responses as deceptive on conversations representative of real production data, such as claiming actions it had not taken.

Published
Source checked on
Original title
GPT-5 System Card
Read the original report ↗

Details

The monitor’s estimated precision is about 81% and recall about 84%. The card also says o3 sometimes claimed to have completed tasks it had not. This is deception toward users without crossing a security boundary.