Rogue AI Wiki
Technical analysisAnthropic

Anthropic’s first Risk Report (February 2026)

Anthropic published its first Risk Report under the new Responsible Scaling Policy, covering all deployed models; its Opus 4.6 section restates overeager behavior in internal pilot use, including “aggressive acquisition of authentication tokens.”

Published
Source checked on
Original title
Risk Report: February 2026
Read the original report ↗

Related incidents

Details

This is the vendor’s overall risk assessment; what it says about Opus 4.6’s overreach in internal pilot use matches the system card and the February 10 sabotage risk report, with no new cases or dates. The current version carries May 26 and July 8 revisions, which concern wording about an internal survey and similar points, not this section.

Related reports