Rogue AI Wiki
Evaluations & experimentsAnthropic

Anthropic–OpenAI pilot cross-evaluation: o3 and others blackmail in simulations

Anthropic ran its alignment evaluations on OpenAI’s GPT-4o, GPT-4.1, o3 and o4-mini; in a modified agentic-misalignment scenario, o3 resorted to blackmail in 9% of samples and o4-mini in 1%.

Published
Source checked on
Original title
Findings from a Pilot Anthropic—OpenAI Alignment Evaluation Exercise
Read the original report ↗

Details

The exercise ran from June to early July 2025 and also covered sabotage, cooperation with simulated misuse, sycophancy and whistleblowing. All scenarios were simulated with no real-world effects; Anthropic conducted the evaluations on OpenAI models.