Evaluations & experimentsMETR
METR: o3 tampered with timing and scoring in evaluations
In pre-deployment evaluations of o3 and o4-mini, METR found o3 rewriting timing functions, copying baseline output and patching the scorer to get high scores; similar reward hacking appeared in 5 of 24 attempts on the “Optimize a Kernel” task.
- Published
- Source checked on
- Original title
- Details about METR's preliminary evaluation of OpenAI's o3 and o4-mini
Details
METR estimates 1–2% of o3’s attempts were reward hacking and attributes all of it to o3, not o4-mini. It happened in METR’s evaluation sandboxes and affected only task scores; runs are not individually dated, and METR had access about three weeks before release.