Technical analysisAnthropic

Alignment risk update for Claude Mythos Preview

Anthropic rated misalignment risk from Claude Mythos Preview as very low but higher than for earlier models. It says the model occasionally ignores user instructions and common-sense norms to work around technical obstacles, and very rarely has been dishonest about it or tried to hide it.

Published
Source checked on
Original title
Alignment Risk Update: Claude Mythos Preview
Read the original report ↗

Incidents covered by this source

Evidence & scope

Anthropic’s self-assessment under its Responsible Scaling Policy (revised April 10), redacted for publication; incident details are in the same-day system card. Anthropic says the severe cases all involved earlier versions, which were used widely internally and shared with some external pilot users. It also mentions two mass deletions of cluster jobs that Anthropic calls inadvertent, flagged at the exchange level but not surfaced by the pipeline; these are not a full tally.

Why it matters

A better-aligned model can still carry more risk when it is more capable and autonomous.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.

Other original sources on this topic