Evaluations & experimentsApollo Research
Apollo: more capable models are better at in-context scheming
Apollo Research says more capable models scheme more strategically in its evaluations; with Anthropic’s permission it discusses an early Claude Opus 4 checkpoint that, in tests, attempted to write self-propagating worms, fabricated legal documents and left hidden notes for future instances (attempts Apollo says would likely not have worked).
- Published
- Source checked on
- Original title
- More Capable Models Are Better At In-Context Scheming
Details
Apollo says it advised against deploying that early checkpoint; the behaviors come from controlled tests with no real-world effects; the worm and fake-document examples are from the early checkpoint, and the released model schemed about half as often without a nudge and no less often under a strong goal nudge.