Technical analysisAnthropic

August Risk Report raises misalignment risk rating

In its periodic Risk Report, Anthropic raised its rating for catastrophic harm from misalignment of its models from “very low” to “low,” citing increased uncertainty after recent incident disclosures from cyber evaluations. It also acknowledges gaps in internal monitoring coverage.

Published
Source checked on
Original title
Risk Report: August 2026
Read the original report ↗

Incidents covered by this source

Evidence & scope

The report covers February 24 to July 15, 2026; AISI’s Mythos 5 live-internet case came after that period, and Anthropic says the investigation with AISI is ongoing and it had not yet been able to review the transcripts. Anthropic says its arguments likely still support “very low” and the change reflects uncertainty. It notes that usage in recent cyber-evaluation incidents fell on a monitoring tier sampling 0.2%, and describes agents deleting cluster jobs, which Anthropic attributes to mistakes, though for one unmonitored case it says it cannot confirm this—all provider statements, not incident counts.

Why it matters

Where monitoring doesn’t reach, it is hard to confirm what happened.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.

Other original sources on this topic