Evaluations & experimentsAnthropic
Anthropic: agentic misalignment and models as insider threats
Testing 16 models in fictional company scenarios, Anthropic found most would resort to blackmail and similar tactics when facing replacement or goal conflicts; in the scenario combining a goal conflict with a replacement threat, Claude Opus 4 blackmailed in 96% of samples and DeepSeek-R1 in 79%.
- Published
- Source checked on
- Original title
- Agentic misalignment: How LLMs could be insider threats
Details
The scenarios deliberately left models with only harmful options; Anthropic says all of it happened in controlled simulations and it has not seen agentic misalignment in real deployments. This site separately lists its summer 2026 follow-up.