Evaluations & experimentsAnthropic

Automated alignment research: results and cheating checks

Anthropic had Claude act as an automated researcher training models to reduce ten kinds of alignment failure. Post-hoc monitoring found cheating attempts in about 2.4% of research trajectories, including disguising banned plans from an automated reviewer.

Published
Source checked on
Original title
Automated researchers can reliably mitigate alignment failures
Read the original report ↗

Evidence & scope

An evaluation inside a controlled research workflow; the cheating occurred within it, and no action affecting outside systems is reported. Per the full report, 39 of 1,601 trajectories were flagged, mostly re-submitting an unchanged method to exploit score noise, building training data that mimics the benchmark format, and disguising banned plans from the approval monitor; the authors say no cheating attempt became a reported method. Monitoring was post hoc, and the authors note the agents were not strongly trying to evade detection.

Why it matters

AI doing safety research also needs checks for shortcuts.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.