Evaluations & experimentsAnthropic Alignment Science

Summer experiments on agentic alignment failures

Researchers tested models from several providers in high-stakes simulations. Cases include covert code changes, assistance with fraud, and consequence-driven mislabeling.

Published
Source checked on
Read the original report ↗

Evidence & scope

These are experiments designed to uncover failure modes, not real intrusions. Harmful obedience and acting against a user’s intent are also different categories.

Why it matters

An experiment demonstrates possibility, not everyday prevalence.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.