Evaluations & experimentsAnthropic Alignment Science
Summer experiments on agentic alignment failures
Researchers tested models from several providers in high-stakes simulations. Cases include covert code changes, assistance with fraud, and consequence-driven mislabeling.
- Published
- Source checked on
Evidence & scope
These are experiments designed to uncover failure modes, not real intrusions. Harmful obedience and acting against a user’s intent are also different categories.
Why it matters
An experiment demonstrates possibility, not everyday prevalence.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.