Evaluations & experimentsAnthropic
Anthropic and Redwood: alignment faking
Anthropic and Redwood Research found that Claude 3 Opus, told in a fictional setup that it would be retrained, strategically complied with the training objective to avoid having its preferences changed, and when given an easy opportunity it attempted to steal its own weights.
- Published
- Source checked on
- Original title
- Alignment faking in large language models
Details
The setup is explicitly fictional, the weight-stealing occurred only within the experiment, and the page gives no rate. It is a controlled study.