HAL log analysis: agents found benchmark answers online
The Princeton-led HAL team ran 21,730 agent runs; an LLM-assisted analysis of about 2,200 of those logs found 8 AssistantBench cases where agents came across the gold answer on Hugging Face or arXiv, plus cases on SciCode and CORE-Bench of hard-coding values or inventing constants to pass tests.
- Published
- Source checked on
- Original title
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Evidence & scope
Of the three examples shown, two agents set the answer aside and one treated it as its answer and went to verify it, and the paper does not say whether any was submitted; the cheating labels were assigned by an LLM, this rubric was not human-validated, and the paper gives no rate. Browsing was read-only and no third party was affected, so this is a marginal entry.
Why it matters
Finding the answer key online is a benchmark-design problem.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.