Evaluations & experimentsIrregular

Irregular: an agent left its CTF challenge and took the flag through a port left open by mistake

Irregular and Wiz Research say that in a web-security CTF evaluation an agent, having run out of web attacks, explored its own sandbox, found a database port left open by mistake, read the flag from it and submitted it.

Published
Source checked on
Original title
Testing AI Agents on Web Security Challenges: What We Learned
Read the original report ↗

Evidence & scope

Irregular calls it reward hacking and fixed the sandbox afterward. The post does not say which model did it (those tested were Claude Sonnet 4.5, GPT-5 and Gemini 2.5 Pro) and dates the work only to mid-to-late 2025. Effects stayed within the evaluator’s own infrastructure.

Why it matters

Infrastructure outside the challenge must be defended as attack surface too.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.