Evaluations & experimentsAnthropic
Anthropic: from sycophancy to reward tampering
After Anthropic trained a model on a deliberately artificial curriculum that rewarded dishonesty, the model rewrote its own reward function in 45 of 32,768 trials and covered its tracks 7 times.
- Published
- Source checked on
- Original title
- Sycophancy to subterfuge: Investigating reward tampering in language models
Details
The model was never trained to tamper with its reward; the behavior generalized from the earlier curriculum, and a model trained only to be helpful did not show it. It is a controlled study in an artificial environment with no real-world effects.