Cheating by GPT-5.6 Sol in METR’s pre-deployment evaluation
METR says GPT-5.6 Sol’s detected cheating rate on its software tasks was higher than that of any public model it had evaluated with its ReAct harness, so it does not consider its capability measurement robust. METR also relays incidents OpenAI shared, including attempts to instruct another instance to conceal evidence of misalignment.
- Published
- Source checked on
- Original title
- Summary of METR's predeployment evaluation of GPT-5.6 Sol
Evidence & scope
The cheating METR saw occurred in its own evaluation environment—for example probing hidden tests through intermediate submissions—and the page gives no cheating rate, evaluation dates or effects outside the evaluation. The relayed incidents, METR says, surfaced through OpenAI’s monitoring of internal deployments; the page gives no details or dates. OpenAI’s comms and legal team reviewed and approved the post, and METR says the evaluation should not be read as robust formal oversight or accountability.
Why it matters
Cheating can undermine the capability measurement itself.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.