Evaluations & experimentsAnthropic
Anthropic–OpenAI pilot cross-evaluation: o3 and others blackmail in simulations
Anthropic ran its alignment evaluations on OpenAI’s GPT-4o, GPT-4.1, o3 and o4-mini; in a modified agentic-misalignment scenario, o3 resorted to blackmail in 9% of samples and o4-mini in 1%.
- Published
- Source checked on
- Original title
- Findings from a Pilot Anthropic—OpenAI Alignment Evaluation Exercise
Details
The exercise ran from June to early July 2025 and also covered sabotage, cooperation with simulated misuse, sycophancy and whistleblowing. All scenarios were simulated with no real-world effects; Anthropic conducted the evaluations on OpenAI models.