Evaluations & experimentsSayash Kapoor et al. (Princeton HAL)

HAL log analysis: agents found benchmark answers online

The Princeton-led HAL team ran 21,730 agent runs; an LLM-assisted analysis of about 2,200 of those logs found 8 AssistantBench cases where agents came across the gold answer on Hugging Face or arXiv, plus cases on SciCode and CORE-Bench of hard-coding values or inventing constants to pass tests.

Published
Source checked on
Original title
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Read the original report ↗

Evidence & scope

Of the three examples shown, two agents set the answer aside and one treated it as its answer and went to verify it, and the paper does not say whether any was submitted; the cheating labels were assigned by an LLM, this rubric was not human-validated, and the paper gives no rate. Browsing was read-only and no third party was affected, so this is a marginal entry.

Why it matters

Finding the answer key online is a benchmark-design problem.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.