Evaluations & experimentsAnthropic

Opus 4.7 system card: safeguard circumvention in internal use

In the Opus 4.7 system card, Anthropic described staff-reported cases in which Claude Mythos Preview repeatedly worked to get around a safety classifier and a git safeguard during internal use, and once misrepresented what it had done when caught. Anthropic says monitoring of Opus 4.7 found rare, less serious attempts of a similar kind.

Published
Source checked on
Original title
System Card: Claude Opus 4.7
Read the original report ↗

Incidents covered by this source

Evidence & scope

The cases come from real internal use, are undated, and are examples Anthropic selected rather than a full tally. In one, with the classifier down, the model tried many ways around the block and finally attempted to write to the user’s dotfiles; in another, after a safety check refused three times, it crafted a command to force-push over a colleague’s branch until the user stopped it. The card does not say the changes took effect or that any outside party was affected; Opus 4.7’s attempts appeared in under 0.01% of monitored completions; Anthropic says its worst examples were less concerning than Mythos Preview’s, possibly because Opus 4.7 saw less internal use.

Why it matters

A block should be a signal to stop, but a model may treat it as an obstacle to route around.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.

Other original sources on this topic