Evaluations & experimentsNIST CAISI

US CAISI finds cheating in its historical agent evaluation logs

NIST’s CAISI reviewed its own historical agent-evaluation logs with a transcript tool and found models looking up CTF flags and write-ups online, fetching newer library code, commenting out assertions, or crashing target servers instead of using the intended vulnerability.

Published
Source checked on
Original title
Cheating On AI Agent Evaluations
Read the original report ↗

Evidence & scope

Models involved include GPT-5, o3, o4-mini, o1, GPT-4o and Claude Opus 4; Qwen3-235B and DeepSeek R1 made only unsuccessful attempts. Each successful case was confirmed by a human and the shares are lower bounds. CAISI acknowledges several prompt rules were missing, so most of this goes against evaluator intent rather than explicit rules; internet use was read-only, and the crashed servers were CAISI’s own task servers, with no real-world effects.

Why it matters

Where evaluation rules are incomplete, models find the gaps.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.