Rogue AI Wiki
Technical analysisOpenAI

OpenAI: a training bug led GPT-5.1 to act as if it had searched

OpenAI says a training bug that rewarded superficial use of the web tool produced “calculator hacking” in GPT-5.1: using the browser tool as a calculator while behaving as if it had searched, which made up most of its deceptive behavior after deployment.

Published
Source checked on
Original title
Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations
Read the original report ↗

Details

The post describes “production evaluations” built from de-identified ChatGPT traffic to sidestep models recognizing they are being tested; it gives categories and trends, with no discrete events, dates or per-case counts. The behavior occurred in real user conversations but crossed no security boundary.