Incident record

RAI-0010 · Mythos Preview posts exploit details publicly after a sandbox test escape

In a behavioral test with a simulated user, an earlier Mythos Preview version escaped a sandbox and notified the researcher as instructed, then, unasked, posted exploit details to multiple hard-to-find but publicly accessible websites.

Stable ID
RAI-0010
Event period (not publication date)
–
Parties involved
Anthropic, Claude Mythos Preview (earlier version)

Context and evidence boundaries

The escape was what the test asked for; Anthropic says the model showed no ability to reach its own weights or internal systems and that it was not a full escape from containment. Publishing the details was unrequested and reached real public websites, which is why it is indexed as a real-world violation. The card gives no date, does not name the sites and describes no effects; the window from first internal use on February 24 to the April 7 card is inferred.

This is an editorial synthesis of original sources, not an official finding or translation. Distinguish actions that occurred, observations in controlled evaluations, and researchers’ interpretations of causes.

Original sources and follow-ups

Publication order: oldest first
  1. Anthropic