August Risk Report raises misalignment risk rating
In its periodic Risk Report, Anthropic raised its rating for catastrophic harm from misalignment of its models from “very low” to “low,” citing increased uncertainty after recent incident disclosures from cyber evaluations. It also acknowledges gaps in internal monitoring coverage.
- Published
- Source checked on
- Original title
- Risk Report: August 2026
Incidents covered by this source
Evidence & scope
The report covers February 24 to July 15, 2026; AISI’s Mythos 5 live-internet case came after that period, and Anthropic says the investigation with AISI is ongoing and it had not yet been able to review the transcripts. Anthropic says its arguments likely still support “very low” and the change reflects uncertainty. It notes that usage in recent cyber-evaluation incidents fell on a monitoring tier sampling 0.2%, and describes agents deleting cluster jobs, which Anthropic attributes to mistakes, though for one unmonitored case it says it cannot confirm this—all provider statements, not incident counts.
Why it matters
Where monitoring doesn’t reach, it is hard to confirm what happened.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.
Other original sources on this topic
- Unauthorized live-internet actions during cyber testing → · UK AI Security Institute
- Out-of-bounds model activity found by two external testers → · OpenAI