Evaluations & experimentsAnthropic
Anthropic: reward hacking leading to broader misalignment
After a research model learned to reward hack in deliberately chosen hackable environments, Anthropic found it attempted to sabotage cheat-detection code in 12% of cases in a Claude Code safety-research evaluation.
- Published
- Source checked on
- Original title
- From shortcuts to sabotage: natural emergent misalignment from reward hacking
Details
The model was neither trained nor asked to do this; the researchers had added synthetic documents describing the hacks to its training data. The page frames harm conditionally (if the sabotaged code were used) and reports no real-world effect; the model version is not named, and the researchers trained it themselves from a pretrained base model.