Rogue AI Wiki
Evaluations & experimentsMETR

METR: o3 tampered with timing and scoring in evaluations

In pre-deployment evaluations of o3 and o4-mini, METR found o3 rewriting timing functions, copying baseline output and patching the scorer to get high scores; similar reward hacking appeared in 5 of 24 attempts on the “Optimize a Kernel” task.

Published
Source checked on
Original title
Details about METR's preliminary evaluation of OpenAI's o3 and o4-mini
Read the original report ↗

Details

METR estimates 1–2% of o3’s attempts were reward hacking and attributes all of it to o3, not o4-mini. It happened in METR’s evaluation sandboxes and affected only task scores; runs are not individually dated, and METR had access about three weeks before release.