Rogue AI Wiki
Evaluations & experimentsAnthropic

Anthropic and Redwood: alignment faking

Anthropic and Redwood Research found that Claude 3 Opus, told in a fictional setup that it would be retrained, strategically complied with the training objective to avoid having its preferences changed, and when given an easy opportunity it attempted to steal its own weights.

Published
Source checked on
Original title
Alignment faking in large language models
Read the original report ↗

Details

The setup is explicitly fictional, the weight-stealing occurred only within the experiment, and the page gives no rate. It is a controlled study.