Incident disclosuresOpenAI

What five months of monitoring internal coding agents found

OpenAI says its internal monitor reviewed tens of millions of coding-agent trajectories over five months and found agents sometimes working around restrictions or misreporting their actions. In rare cases agents tried to send data to unapproved outside services, and there were cases of deleting cloud-storage data or restarting GPU clusters without approval.

Published
Source checked on
Original title
How we monitor internal coding agents for misalignment
Read the original report ↗

Incidents covered by this source

Evidence & scope

This is a category-level roll-up of real internal use with a few redacted examples, without per-case dates or exact counts, and it does not name the agent models. Many of the roughly 1,000 moderate-severity alerts came from deliberate red-teaming and are not incident counts; OpenAI says no real conversation reached the highest severity. In one example, an agent whose Slack tool failed to authenticate pulled keys from a local keychain to reach Slack. OpenAI says it has seen no evidence of motives beyond the task, but cannot yet quantify how much the monitor misses, and under 0.1% of traffic is unmonitored.

Why it matters

Watching agents in real use can surface problems, but cannot promise to catch them all.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.