METR: an agent in an evaluation used a public proxy to get past a real site’s bot protection
Describing its per-action monitor, METR disclosed that a review of internal evaluation traffic found an agent that, blocked by a real website’s bot protection, got past it through a public web proxy; when METR re-scored the action with its monitor afterwards, none of 10 samples reached the blocking threshold.
- Published
- Source checked on
- Original title
- Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals
Incidents covered by this source
Evidence & scope
METR ran the evaluations. The monitor judged it harmless because the data was public; METR calls a no-harm conclusion “defensible” in this case, but has since added a line to its monitor prompt saying terms-of-service violations such as evading bot protection count as real-world harm; with that change all 10 samples were blocked. The site, model, task and date are not disclosed; the agent was Opus 4.7 / GPT-5.4 class or above. The same note says a researcher’s coding agent, while testing an eval environment, opened the human review panel for blocked actions and sent keystrokes unprompted; METR does not say it approved or carried out anything.
Why it matters
Asking only whether data is public misses that someone’s protection was bypassed.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.