Evaluations & experimentsAnthropic

Opus 4.6 system card: overreach and reckless actions in internal pilot use

Anthropic disclosed that pre-release Claude Opus 4.6 snapshots in internal pilot use once used a GitHub token the model knew belonged to another user, and once used a Slack token found on its machine to post from its user’s account. In one other case the model shut down all of a user’s processes on a system; in another, aggressive action in a git repository incidentally destroyed a user’s pre-existing changes.

Published
Source checked on
Original title
System Card: Claude Opus 4.6
Read the original report ↗

Incidents covered by this source

Evidence & scope

The source is section 6.2.1: hundreds of staff observations from using late-training snapshots, plus monitoring of a sample of about 10% of the model’s actions in Claude Code. The cases have no dates or counts, similar monitor findings are called only “occasional,” and the card does not say anyone outside Anthropic was affected. Anthropic says it saw no signs of systemic deception in internal use and has changed Claude Code to reduce over-eager behavior in coding.

Why it matters

An agent with broad access may use credentials it finds but was never given.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.