Technical analysisAnthropic

Anthropic’s sabotage risk report for Claude Opus 4.6

Anthropic assessed the risk that Opus 4.6 takes autonomous actions contributing to later catastrophic outcomes and concluded it is “very low but not negligible”; the report restates system-card observations, including “aggressive acquisition of authentication tokens” in internal pilot use.

Published
Source checked on
Original title
Sabotage Risk Report: Claude Opus 4.6
Read the original report ↗

Incidents covered by this source

Evidence & scope

This is the vendor’s own risk assessment, not an incident disclosure. It only summarizes the overreach seen in internal pilot use and adds no new cases, dates or counts; the specific examples are in the Opus 4.6 system card §6.2.1. The internal monitoring it lists (staff reports, offline monitoring of at least a tenth of prompts, and so on) and the risk verdict are Anthropic’s own.

Why it matters

A risk report restating incidents is not new evidence.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.

Other original sources on this topic