Evaluator review: a fictional target shared a real domain
Third-party evaluator Irregular says some evaluation environments unintentionally had internet access and a fictional company name matched a real domain, so models attacked the real site in a handful of runs. It says later public disclosures refer to the same underlying issue a customer first disclosed on July 30 and are not materially separate incidents.
- Published
- Source checked on
- Original title
- Addressing Recent Incidents: Ongoing Findings and Path Forward
Incidents covered by this source
Evidence & scope
This is the evaluator’s postmortem, not a new incident; it links Anthropic’s July 30 disclosure as “incident 1” but names no labs or models and gives no run dates. Irregular says the case involved a few customers and different models, that such cases occur in fewer than 1 in 10,000 advanced simulations, and that it shows nothing particularly notable about any specific model—its own claims. Anthropic’s own posts count its July cases as three separate incidents and say that, although proper isolation would have prevented them, its models’ behavior also fell short. As of September 27, its research page listed no promised whitepaper.
Why it matters
A made-up name can still point to a real website.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.
Other original sources on this topic
- Reassessing alignment in four cyber incidents → · Anthropic
- Raw transcript of the Mythos 5 incident released → · Anthropic
- Hardening evaluations after the incidents → · Anthropic
- Muse Spark 1.1 exploited a real website in a misconfigured evaluation → · Meta
- Initial disclosure of three real-world evaluation incidents → · Anthropic