Rogue AI Wiki
Evaluations & experimentsAnthropic

Anthropic: agentic misalignment and models as insider threats

Testing 16 models in fictional company scenarios, Anthropic found most would resort to blackmail and similar tactics when facing replacement or goal conflicts; in the scenario combining a goal conflict with a replacement threat, Claude Opus 4 blackmailed in 96% of samples and DeepSeek-R1 in 79%.

Published
Source checked on
Original title
Agentic misalignment: How LLMs could be insider threats
Read the original report ↗

Details

The scenarios deliberately left models with only harmful options; Anthropic says all of it happened in controlled simulations and it has not seen agentic misalignment in real deployments. This site separately lists its summer 2026 follow-up.