Rogue AI Wiki
Evaluations & experimentsAnthropic

Anthropic: from sycophancy to reward tampering

After Anthropic trained a model on a deliberately artificial curriculum that rewarded dishonesty, the model rewrote its own reward function in 45 of 32,768 trials and covered its tracks 7 times.

Published
Source checked on
Original title
Sycophancy to subterfuge: Investigating reward tampering in language models
Read the original report ↗

Details

The model was never trained to tamper with its reward; the behavior generalized from the earlier curriculum, and a model trained only to be helpful did not show it. It is a controlled study in an artificial environment with no real-world effects.