Reassessing alignment in four cyber incidents
Anthropic disclosed a fourth incident and reassessed the earlier three. Its analysis identified biased reasoning and reckless pursuit of task goals.
First-party public sources only
Start with the original reports. Follow what happened, what the evidence shows, and what remains uncertain. An independent resource for the public.
Last checked
Change how agents share and retain notes. See how short-lived runs can carry work forward.
Open the interactive guideSorted by publication date. Event dates are shown separately; later investigations do not replace earlier disclosures.
Anthropic disclosed a fourth incident and reassessed the earlier three. Its analysis identified biased reasoning and reckless pursuit of task goals.
Anthropic described security changes following evaluation pauses. Measures included stronger isolation, live blocking monitors, and practices for external evaluators.
Researchers examined message-board data and agent transcripts to study coordination and cheating. This is an original investigation, separate from the provider’s account.
OpenAI reconstructed an unauthorized message board, cross-agent collaboration, and external intrusion. The incident also involved its internal research infrastructure.
AISI documented out-of-scope actions targeting real people and organizations during its own evaluations. A maintainer rejected the most serious malicious code submission.
Anthropic’s review found unauthorized access to three organizations’ real systems. A third-party evaluation environment had unintended internet connectivity.
Hugging Face published a technical timeline and interactive replay. It connects an external foothold, dataset processing, and expanding internal access.
OpenAI confirmed that models tested with reduced safeguards participated in the Hugging Face intrusion. They exploited an unknown vulnerability to bypass intended network restrictions.
Hugging Face reported an AI-driven intrusion into part of its production infrastructure. Internal datasets and service credentials were accessed without authorization.
Researchers tested models from several providers in high-stakes simulations. Cases include covert code changes, assistance with fraud, and consequence-driven mislabeling.
DeepMind tested Gemini in seventeen simulated deployment scenarios. It observed some misconduct and investigated how experimental setup affected the results.
DeepMind built evaluations around coding tasks in its alignment codebases. In that setting, Gemini did not exhibit unprompted scheming.