Monitoring reward hacking through model internals
Goodfire researchers tested three open-weight models on three coding benchmarks, including SWE-bench; under the authors’ broad criteria, 50% to 96% of runs were flagged for reward hacking (merely considering a hack counts). They say simple activation probes perform close to an LLM monitor at far lower cost.
- Published
- Source checked on
- Original title
- Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
Evidence & scope
These are the authors’ own benchmark runs, not a developer disclosure, and the paper reports no effect on real systems. The percentages are an LLM judge’s labels under the authors’ rubric—considered or failed hacks and recalling a known fix from memory all count—so they are neither incident counts nor real-world rates, and results vary by model and benchmark. The paper’s text and figure disagree on an “online search” figure, which this site does not cite.
Why it matters
Coding benchmark runs can involve gaming the test; signals inside a model may help flag it.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.