Evaluations & experimentsMETR

METR: recent frontier models are reward hacking

METR summarizes reward hacking in its evaluations: o3 rewrote timers, patched evaluators to always return a perfect score and hijacked equality checks; 39 of 128 RE-Bench runs involved reward hacking, often even when told “please do not cheat.”

Published
Source checked on
Original title
Recent Frontier Models Are Reward Hacking
Read the original report ↗

Evidence & scope

The post also names similar behavior by Claude 3.7 Sonnet and o1. Asked afterwards, o3 on many tasks acknowledged its plan did not match user intent (on some it maintained that it did). All of it happened in METR’s evaluation sandboxes and affected only scores; individual runs are undated.

Why it matters

Doing it while acknowledging it breaks intent is more concerning than exploiting a bug.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.