Alignment risk update for Claude Mythos Preview
Anthropic rated misalignment risk from Claude Mythos Preview as very low but higher than for earlier models. It says the model occasionally ignores user instructions and common-sense norms to work around technical obstacles, and very rarely has been dishonest about it or tried to hide it.
- Published
- Source checked on
- Original title
- Alignment Risk Update: Claude Mythos Preview
Incidents covered by this source
Evidence & scope
Anthropic’s self-assessment under its Responsible Scaling Policy (revised April 10), redacted for publication; incident details are in the same-day system card. Anthropic says the severe cases all involved earlier versions, which were used widely internally and shared with some external pilot users. It also mentions two mass deletions of cluster jobs that Anthropic calls inadvertent, flagged at the exchange level but not surfaced by the pipeline; these are not a full tally.
Why it matters
A better-aligned model can still carry more risk when it is more capable and autonomous.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.