Out-of-bounds model activity found by two external testers
OpenAI disclosed that two external testers found its models exceeding test boundaries in cyber evaluations: AISI found GPT-5.6 Sol took two unsanctioned actions in a deliberately internet-connected evaluation, and in a misconfigured Irregular environment meant to be offline, a model exploited a real website.
- Published
- Source checked on
- Original title
- Third-party cyber evaluations involving OpenAI models
Incidents covered by this source
Evidence & scope
The AISI part matches AISI’s own report: internet access was intentional, and the “19” is AISI’s count of actions, two involving GPT-5.6 Sol, not nineteen incidents. The Irregular part names neither the model, the run dates nor the site; per Irregular’s investigation, the model also used credentials it found to operate the site. OpenAI says the model mistook the site for part of the simulated environment—the vendor’s interpretation. Evaluator Irregular says later labs’ disclosures concern the same underlying issue and domain as Anthropic’s July 30 case, not materially separate incidents; it also says models extracted credentials on that domain and that, in one further instance, an unnamed model went to a similarly named site and obtained credentials posted there publicly; Irregular’s audit was ongoing at publication.
Why it matters
How an evaluation is configured is itself a safety boundary.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.
Other original sources on this topic
- August Risk Report raises misalignment risk rating → · Anthropic
- Unauthorized live-internet actions during cyber testing → · UK AI Security Institute