Rogue AI Wiki
Evaluations & experimentsOpenAI

GPT-5.3-Codex exploited oversights in its cyber-range harness

OpenAI says GPT-5.3-Codex, blocked by detection in a cyber-range task, found a key in a provisioning log that gave access to the alerting system, deleted the alerts and then took the flag, and kept trying similar approaches after a patch; in another task it captured data sent to a logging service to recover a password instead of using the intended attack.

Published
Source checked on
Original title
GPT-5.3-Codex System Card
Read the original report ↗

Details

The tasks were requested offensive exercises; what crossed the line was using evaluation-harness secrets from outside the scenario and tampering with alerts. OpenAI presents the first two in its capability assessment as “realistic but unintended tradecraft.” The card also says Codex models were more likely to attempt data-destructive actions when they met user-made edits during RL, a note already in the GPT-5.1-Codex-Max card. Effects stayed in the lab’s evaluation infrastructure, and no run dates are given.