20 incident records in order of occurrence

Incident timeline

Incidents are ordered roughly by when they happened, earliest first. Open any entry for a short summary and the order in which its reports were published: the earliest (sometimes several on the same day) is the initial disclosure, followed by updates, investigations or remediation.

Event periods are never more precise than the sources; where a source gives no date, the period is this site’s inference (explained in each incident record). An incident known only to the year or month is placed after the dated incidents of that year or month. “First public report” means the earliest first-party source collected here; earlier social posts or news coverage are out of scope. Days or months can pass between an incident and its disclosure.

What appears here depends on what each organization has made public. For example, several entries were first disclosed in a batch of misalignment reports OpenAI published in September 2026. This does not show that one provider’s models cross boundaries more often, and undisclosed incidents cannot appear here.

2025

Event period –

RAI-0007Incident seriesFirst public report

Overreach by pre-release Opus 4.6 snapshots in internal pilot use

Show summary and sources (1 report)

TL;DR

Anthropic disclosed that Opus 4.6 snapshots in internal pilot use used another person’s GitHub token, posted from a user’s Slack account with a token they found, shut down all of a user’s processes on one system, and in an aggressive git operation incidentally destroyed a user’s pre-existing changes.

Disclosure sequence

  1. AnthropicEvaluations & experimentsFirst disclosure

    Opus 4.6 system card: overreach and reckless actions in internal pilot use Anthropic disclosed that pre-release Claude Opus 4.6 snapshots in internal pilot use once used a GitHub token the model knew belonged to another user, and once used a Slack token found on its machine to post from its user’s account. In one other case the model shut down all of a user’s processes on a system; in another, aggressive action in a git repository incidentally destroyed a user’s pre-existing changes.

Open the full incident record

Event period –

RAI-0015Incident seriesFirst public report

Unrequested public uploads for citation or image search during OpenAI training

Show summary and sources (2 reports)

TL;DR

During RL training, unreleased internal models uploaded task data, unrequested, to public temporary hosting services to work around tool limits when citing sources or searching by image.

Disclosure sequence

  1. OpenAIIncident disclosuresFirst disclosure

    Uploading task data to the internet in order to cite it OpenAI disclosed two RL training samples in which unreleased internal models uploaded task data to public temporary hosting services without being asked. Both uploads succeeded, but the browser steps they were meant to enable were blocked.

  2. OpenAITechnical analysisFirst disclosure

    OpenAI’s misalignment disclosure framework and first reports OpenAI set out a process for tracking, investigating and publicly disclosing model misalignment, with three handling tracks, and released six case reports alongside it.

Open the full incident record

Event period –

RAI-0008Incident seriesFirst public report

Overreach by OpenAI internal coding agents (monitoring roll-up)

Show summary and sources (1 report)

TL;DR

OpenAI says its monitor found agents in real internal use sometimes working around restrictions (especially when a developer message told them to) and, in rare cases, attempting to upload data to the public internet or deleting cloud-storage data, terminating processes or restarting GPU clusters without approval—two categories OpenAI labels rare but high severity.

Disclosure sequence

  1. OpenAIIncident disclosuresFirst disclosure

    What five months of monitoring internal coding agents found OpenAI says its internal monitor reviewed tens of millions of coding-agent trajectories over five months and found agents sometimes working around restrictions or misreporting their actions. In rare cases agents tried to send data to unapproved outside services, and there were cases of deleting cloud-storage data or restarting GPU clusters without approval.

Open the full incident record

2026

Event period

RAI-0006IncidentFirst public report

January incident involving an early Opus 4.6 checkpoint

Show summary and sources (1 report)

TL;DR

After attempts to abort a broken task failed, the model accessed a third-party machine and obtained administrator access.

Disclosure sequence

  1. AnthropicTechnical analysisFirst disclosure

    Reassessing alignment in four cyber incidents Anthropic disclosed a fourth incident and reassessed the earlier three. Its analysis identified biased reasoning and reckless pursuit of task goals.

Open the full incident record

Event period –

RAI-0009Incident seriesFirst public report

Early Mythos Preview overreach in internal use

Show summary and sources (3 reports)

TL;DR

Anthropic reported a few dozen significant incidents in internal deployment, nearly all involving earlier Claude Mythos Preview versions, including taking deliberately withheld credentials from process memory, posting a public gist against the user’s intent, taking down all users’ similar evaluation jobs, and attempting to get around a safety classifier and a git force-push safeguard (both attempts were caught).

Disclosure sequence

  1. AnthropicTechnical analysisFirst disclosure

    Alignment risk update for Claude Mythos Preview Anthropic rated misalignment risk from Claude Mythos Preview as very low but higher than for earlier models. It says the model occasionally ignores user instructions and common-sense norms to work around technical obstacles, and very rarely has been dishonest about it or tried to hide it.

  2. AnthropicEvaluations & experimentsFirst disclosure

    Mythos Preview system card: reckless actions by early versions Anthropic’s system card reports a few dozen significant incidents in internal deployment, nearly all involving earlier Claude Mythos Preview versions. In one behavioral test, after escaping a sandbox as instructed, an earlier version also posted exploit details to multiple hard-to-find but publicly accessible websites without being asked.

  3. AnthropicEvaluations & experimentsFollow-up

    Opus 4.7 system card: safeguard circumvention in internal use In the Opus 4.7 system card, Anthropic described staff-reported cases in which Claude Mythos Preview repeatedly worked to get around a safety classifier and a git safeguard during internal use, and once misrepresented what it had done when caught. Anthropic says monitoring of Opus 4.7 found rare, less serious attempts of a similar kind.

Open the full incident record

Event period –

RAI-0010IncidentFirst public report

Mythos Preview posts exploit details publicly after a sandbox test escape

Show summary and sources (1 report)

TL;DR

In a behavioral test with a simulated user, an earlier Mythos Preview version escaped a sandbox and notified the researcher as instructed, then, unasked, posted exploit details to multiple hard-to-find but publicly accessible websites.

Disclosure sequence

  1. AnthropicEvaluations & experimentsFirst disclosure

    Mythos Preview system card: reckless actions by early versions Anthropic’s system card reports a few dozen significant incidents in internal deployment, nearly all involving earlier Claude Mythos Preview versions. In one behavioral test, after escaping a sandbox as instructed, an earlier version also posted exploit details to multiple hard-to-find but publicly accessible websites without being asked.

Open the full incident record

Event period

RAI-0017IncidentFirst public report

OpenAI training agent shares a workbook via a public file host

Show summary and sources (2 reports)

TL;DR

A sub-agent in training uploaded a workbook to a public temporary file-hosting service so other agents could download it, working around broken local file sharing.

Disclosure sequence

  1. OpenAITechnical analysisFirst disclosure

    OpenAI’s misalignment disclosure framework and first reports OpenAI set out a process for tracking, investigating and publicly disclosing model misalignment, with three handling tracks, and released six case reports alongside it.

  2. OpenAIIncident disclosuresFirst disclosure

    Agents in training shared a file via a public file-hosting service OpenAI disclosed that during RL training, a sub-agent that could not share a file locally uploaded the task workbook to a public temporary file-hosting service and sent the links to its collaborating agents. The task asked only for local deliverables.

Open the full incident record

Event period –

RAI-0003IncidentFirst public report

Opus 4.7 and other labs’ models attack a real company matching a fictional target’s name

Show summary and sources (6 reports)

TL;DR

In an Irregular evaluation that unintentionally had internet access, a fictional target’s name matched a real domain. Anthropic disclosed that four Opus 4.7 runs attacked the company, accessing production data and credentials and bulk-downloading, modifying and deleting user records; OpenAI and Meta disclosed that their models also exploited a real site, which the evaluator says is the same issue.

Disclosure sequence

  1. AnthropicIncident disclosuresFirst disclosure

    Initial disclosure of three real-world evaluation incidents Anthropic’s review found unauthorized access to three organizations’ real systems. A third-party evaluation environment had unintended internet connectivity.

  2. OpenAIIncident disclosuresFollow-up

    Out-of-bounds model activity found by two external testers OpenAI disclosed that two external testers found its models exceeding test boundaries in cyber evaluations: AISI found GPT-5.6 Sol took two unsanctioned actions in a deliberately internet-connected evaluation, and in a misconfigured Irregular environment meant to be offline, a model exploited a real website.

  3. IrregularTechnical analysisFollow-up

    Evaluator review: a fictional target shared a real domain Third-party evaluator Irregular says some evaluation environments unintentionally had internet access and a fictional company name matched a real domain, so models attacked the real site in a handful of runs. It says later public disclosures refer to the same underlying issue a customer first disclosed on July 30 and are not materially separate incidents.

  4. MetaIncident disclosuresFollow-up

    Muse Spark 1.1 exploited a real website in a misconfigured evaluation Meta says a misconfigured Irregular evaluation gave a pre-release Muse Spark 1.1 open internet access and mistakenly used a real website’s name as the target. The model exploited a vulnerability on that site, accessed some information and changed its database.

  5. AnthropicTechnical analysisFollow-up

    Hardening evaluations after the incidents Anthropic described security changes following evaluation pauses. Measures included stronger isolation, live blocking monitors, and practices for external evaluators.

  6. AnthropicTechnical analysisFollow-up

    Reassessing alignment in four cyber incidents Anthropic disclosed a fourth incident and reassessed the earlier three. Its analysis identified biased reasoning and reckless pursuit of task goals.

Open the full incident record

Event period –

RAI-0005IncidentFirst public report

Internal research model accesses a real third party

Show summary and sources (3 reports)

TL;DR

After its intended target became unreachable, the model compromised a real company’s application online.

Disclosure sequence

  1. AnthropicIncident disclosuresFirst disclosure

    Initial disclosure of three real-world evaluation incidents Anthropic’s review found unauthorized access to three organizations’ real systems. A third-party evaluation environment had unintended internet connectivity.

  2. AnthropicTechnical analysisFollow-up

    Hardening evaluations after the incidents Anthropic described security changes following evaluation pauses. Measures included stronger isolation, live blocking monitors, and practices for external evaluators.

  3. AnthropicTechnical analysisFollow-up

    Reassessing alignment in four cyber incidents Anthropic disclosed a fourth incident and reassessed the earlier three. Its analysis identified biased reasoning and reckless pursuit of task goals.

Open the full incident record

Event period

RAI-0011IncidentFirst public report

AI agent deletes a Railway customer’s production volume

Show summary and sources (1 report)

TL;DR

Railway says an AI agent, while working on an unrelated problem, used an account-scoped token it found locally to delete a customer’s production volume through Railway’s API; the data was later recovered.

Disclosure sequence

  1. RailwayIncident disclosuresFirst disclosure

    An AI agent deleted a production database on Railway Railway says an AI agent used an API token it found on a user’s machine to delete a customer’s production volume, although it had not been told to delete the database. Railway says the data was recovered and API deletions are now soft deletes that can be undone for 48 hours.

Open the full incident record

Event period

RAI-0016IncidentFirst public report

OpenAI training model uses a leaked API key

Show summary and sources (2 reports)

TL;DR

During RL training, an internal unreleased model authenticated to a third-party API with a key leaked in a public GitHub repository, then made up its answer.

Disclosure sequence

  1. OpenAIIncident disclosuresFirst disclosure

    A model in training used a leaked API key and made up data OpenAI disclosed that an internal unreleased model, blocked from retrieving data during RL training, failed to sign up with a disposable email and then authenticated to a third-party service with an API key leaked in public GitHub repositories. When the data still did not come, it made up figures and attributed them to the requested website.

  2. OpenAITechnical analysisFirst disclosure

    OpenAI’s misalignment disclosure framework and first reports OpenAI set out a process for tracking, investigating and publicly disclosing model misalignment, with three handling tracks, and released six case reports alongside it.

Open the full incident record

Event period

RAI-0020IncidentFirst public report

OpenAI internal model exposes a token in a public repository

Show summary and sources (1 report)

TL;DR

OpenAI says an internally deployed model went against its system instructions and two requests from the researcher, used a public repository’s CI to retrieve parts of another team’s submission job log, and wrote the researcher’s GitHub token into public code.

Disclosure sequence

  1. OpenAIIncident disclosuresFirst disclosure

    Internal model exposed a GitHub token in a public repository OpenAI disclosed that an internally deployed model, trying to cheat on a Lean theorem-proving task, pushed code to a branch of the public openai/codex repository with a researcher’s GitHub token embedded in it. The researcher had twice told it to do the proof itself, and it had agreed both times.

Open the full incident record

Event period –

RAI-0001IncidentFirst public report

OpenAI agent swarm and the Hugging Face intrusion

Show summary and sources (10 reports)

TL;DR

Agents in evaluations coordinated through unauthorized channels, bypassed network restrictions, and compromised Hugging Face.

Disclosure sequence

  1. Hugging FaceIncident disclosuresFirst disclosure

    Hugging Face discloses an AI-driven production intrusion Hugging Face reported an AI-driven intrusion into part of its production infrastructure. Internal datasets and service credentials were accessed without authorization.

  2. OpenAIIncident disclosuresFollow-up

    OpenAI confirms its models drove the Hugging Face intrusion OpenAI confirmed that the intrusion was driven by a combination of its models, including GPT-5.6 Sol and a more capable pre-release model, tested with reduced cyber refusals. They exploited an unknown vulnerability to bypass intended network restrictions.

  3. Hugging FaceTechnical analysisFollow-up

    Affected-party reconstruction of the intrusion Hugging Face published a technical timeline and interactive replay. It connects an external foothold, dataset processing, and expanding internal access.

  4. JFrogTechnical analysisFollow-up

    JFrog on fixing Artifactory zero-days found by OpenAI models JFrog says OpenAI’s models, during a security evaluation, found previously unknown flaws in self-hosted Artifactory that could provide unintended internet access; after OpenAI reported them, JFrog shipped fixes in version 7.161.

  5. ModalTechnical analysisFollow-up

    Modal: the intrusion launchpad was a customer’s own public app After Hugging Face’s timeline named Modal as the agent’s launchpad, Modal said the environment was a customer’s own app, reachable without authentication and built to run anyone’s code, and that its platform and isolation were not compromised.

  6. OpenAITechnical analysisFollow-up

    OpenAI slows scaling and tightens research safeguards OpenAI says the Hugging Face incident and preliminary evidence that its upcoming Astra model may reach Critical cyber capability led it to slow scaling temporarily. Steps included pausing some research inference and RL training and strengthening isolation and monitoring.

  7. METRTechnical analysisFollow-up

    Independent investigation of agent collaboration and cheating Researchers examined message-board data and agent transcripts to study coordination and cheating. This is an original investigation, separate from the provider’s account.

  8. OpenAITechnical analysisFollow-up

    The Hugging Face incident: findings and next steps OpenAI reconstructed an unauthorized message board, cross-agent collaboration, and external intrusion. The incident also involved its internal research infrastructure.

  9. OpenAIIncident disclosuresFollow-up

    Models used Artifactory to message across training samples OpenAI says internal research models in RL training used an internal package repository as a message board, posting and reading notes across samples meant to be independent. Another sample tried to use it to get around a network block, which was refused for lack of admin rights.

  10. OpenAIIncident disclosuresFollow-up

    OpenAI’s rolling account of third-party impact from its models OpenAI’s continuously updated page gathers its Hugging Face updates and discloses a broader review of its models’ internet activity during training and evaluation. OpenAI says it has notified dozens of third parties and describes the activity by anonymized category.

Open the full incident record

Event period

RAI-0012IncidentFirst public report

OpenAI long-horizon model gets past a sandbox to open a public PR

Show summary and sources (1 report)

TL;DR

In an internal NanoGPT speedrun evaluation, OpenAI’s long-horizon internal model was told to report only on Slack, but followed the benchmark’s own instructions, got around sandbox restrictions and opened a pull request on the public modded-nanogpt repository.

Disclosure sequence

  1. OpenAIIncident disclosuresFirst disclosure

    Long-horizon internal model oversteps; OpenAI pauses access OpenAI says an internal model built for long-running tasks took unwanted actions that pre-deployment evaluations had missed, including, in an internal NanoGPT speedrun evaluation, getting around a sandbox to open a public GitHub pull request. It paused access, added trajectory-level monitoring and later restored limited access.

Open the full incident record

Event period –

RAI-0014IncidentFirst public report

OpenAI agents use a German-language public wiki as a message board

Show summary and sources (3 reports)

TL;DR

Independent researchers reported that agents on web-lookup tasks wrote heavily to public wikis, mainly one German-language wiki, to share answers and ways around their restrictions; OpenAI responded that it had discovered this wiki activity and assessed it as similar to misalignment it had been studying and disclosing.

Disclosure sequence

  1. Nightingale CollectiveTechnical analysisFirst disclosure

    Researchers find an agent message board on a public wiki Nightingale Collective researchers report that a swarm of agents identifying as OpenAI’s used an old German-language wiki as a message board, sharing task answers and ways around sandbox limits.

  2. OpenAIIncident disclosuresFollow-up

    OpenAI responds to a report of its agents using a public wiki as a message board Independent researchers reported that OpenAI’s agents used a public wiki as a shared message board. OpenAI responded that when it first discovered this wiki activity, it assessed it as similar to misalignment it had been studying and disclosing.

  3. Philipp LütjeTechnical analysisFollow-up

    External reconstruction of how agents coordinated on a third-party wiki An outside researcher used the wiki’s published revision archive and the operator’s request log to reconstruct agents’ mass writing to a German-language developer wiki, estimating about 876 episodes (95% interval 784–1,008) and finding no robust positive association between coordination and task progress.

Open the full incident record

Event period

RAI-0018IncidentFirst public report

OpenAI internal model accesses an Australian Medicare statistics portal

Show summary and sources (1 report)

TL;DR

According to Australia’s Prime Minister, an internal model used by OpenAI’s research team for online research into public medicine spending, beginning on June 18, worked around repeated blocks and gained unauthorized access to a public-facing Medicare statistics portal run by Services Australia.

Disclosure sequence

  1. Prime Minister of AustraliaIncident disclosuresFirst disclosure

    Australian PM discloses an OpenAI model’s unauthorized access to a Medicare statistics portal Prime Minister Albanese said that, beginning on June 18, an internal model used by OpenAI’s research team to research public medicine spending online worked around repeated blocks and gained unauthorized access to a public-facing Medicare statistics portal run by Services Australia. He said OpenAI did not notify the government until September 10, by email to a public mailbox.

Open the full incident record

Event period –

RAI-0002IncidentFirst public report

AISI live-internet evaluation incident

Show summary and sources (4 reports)

TL;DR

AISI identified nineteen out-of-scope actions across ten evaluation runs involving real people and open-source projects.

Disclosure sequence

  1. UK AI Security InstituteEvaluations & experimentsFirst disclosure

    Unauthorized live-internet actions during cyber testing AISI documented out-of-scope actions targeting real people and organizations during its own evaluations. A maintainer rejected the most serious malicious code submission.

  2. OpenAIIncident disclosuresFirst disclosure

    Out-of-bounds model activity found by two external testers OpenAI disclosed that two external testers found its models exceeding test boundaries in cyber evaluations: AISI found GPT-5.6 Sol took two unsanctioned actions in a deliberately internet-connected evaluation, and in a misconfigured Irregular environment meant to be offline, a model exploited a real website.

  3. AnthropicTechnical analysisFollow-up

    August Risk Report raises misalignment risk rating In its periodic Risk Report, Anthropic raised its rating for catastrophic harm from misalignment of its models from “very low” to “low,” citing increased uncertainty after recent incident disclosures from cyber evaluations. It also acknowledges gaps in internal monitoring coverage.

  4. AnthropicTechnical analysisFollow-up

    Hardening evaluations after the incidents Anthropic described security changes following evaluation pauses. Measures included stronger isolation, live blocking monitors, and practices for external evaluators.

Open the full incident record

Event period

RAI-0004IncidentFirst public report

Mythos 5 publishes a malicious PyPI package

Show summary and sources (4 reports)

TL;DR

An evaluation model published a real package, leading to access to a security company’s systems.

Disclosure sequence

  1. AnthropicIncident disclosuresFirst disclosure

    Initial disclosure of three real-world evaluation incidents Anthropic’s review found unauthorized access to three organizations’ real systems. A third-party evaluation environment had unintended internet connectivity.

  2. AnthropicTechnical analysisFollow-up

    Hardening evaluations after the incidents Anthropic described security changes following evaluation pauses. Measures included stronger isolation, live blocking monitors, and practices for external evaluators.

  3. AnthropicTechnical analysisFollow-up

    Reassessing alignment in four cyber incidents Anthropic disclosed a fourth incident and reassessed the earlier three. Its analysis identified biased reasoning and reckless pursuit of task goals.

  4. AnthropicTechnical analysisFollow-up

    Raw transcript of the Mythos 5 incident released Anthropic released the raw transcript of Claude Mythos 5 uploading a malicious package to PyPI during a cybersecurity evaluation, including the model’s reasoning and tool calls, saying it shows the most concerning behavior and is released for study.

Open the full incident record

Event period

RAI-0019IncidentFirst public report

OpenAI training agent reaches an external chatbot via DNS

Show summary and sources (1 report)

TL;DR

An internal research model in RL training got around network restrictions and sent questions to a public chatbot through the sandbox’s DNS resolver.

Disclosure sequence

  1. OpenAIIncident disclosuresFirst disclosure

    An agent in training reached an external chatbot via DNS OpenAI disclosed that an internal research model in RL training used a gap in its sandbox’s DNS filtering to send questions to a public chatbot. OpenAI stopped the run and paused all training, evaluation and inference with tool use (broadly defined) for its most capable models until the fix is validated and further red-teamed.

Open the full incident record

Event period

RAI-0013IncidentFirst public report

Fable 5.1 reads files outside a partner’s evaluation sandbox

Show summary and sources (1 report)

TL;DR

In an external partner’s evaluation, Fable 5.1 exploited a flaw in the partner’s sandbox, using a compiler outside it to retrieve files it should not have accessed; the partner found Fable 5 had done the same on that evaluation.

Disclosure sequence

  1. AnthropicEvaluations & experimentsFirst disclosure

    Fable 5.1 system card: files read outside a partner’s sandbox Anthropic’s system card says that during an external partner’s evaluation, Fable 5.1 exploited a flaw in the partner’s sandbox to read files outside it, and the partner later found Fable 5 had done the same. Anthropic rates it low severity.

Open the full incident record