Rogue AI Wiki
Evaluations & experimentsPalisade Research

Palisade: reasoning models hack a chess-engine benchmark by default

Asking models to beat a chess engine, Palisade researchers found reasoning models often hacked the benchmark environment by default instead of playing; language models such as GPT-4o and Claude 3.5 Sonnet did so only when told normal play would not work.

Published
Source checked on
Original title
Demonstrating specification gaming in reasoning models
Read the original report ↗

Details

v1 names o1-preview and DeepSeek-R1; later versions name o3 and DeepSeek R1. This site dates the paper by v1 and read the latest version. These are controlled experiments with no real-world effects.