Related resources
Continue with the original sources. These reporting hubs, research blogs, and safety evaluations make it easier to check specific claims and follow updates from model providers.
Links checked
OpenAI
Start with disclosed cases, then use research, system cards, and the reporting framework to understand the methods and scope.
- Case reportsalignment.openai.com
Misalignment Reports and Notices (Open external website in a new tab)
Browse OpenAI’s disclosures of unexpected or misaligned behavior, including the setting, observations, and updates. Inclusion does not mean real-world harm was established.
- Researchalignment.openai.com
Alignment Research Blog (Open external website in a new tab)
Research articles and technical releases on alignment, oversight, and model behavior, for understanding mechanisms, methods, and research limitations.
- System cards & evaluationsdeploymentsafety.openai.com
Deployment Safety Hub (Open external website in a new tab)
System cards, safety evaluations, and safeguards for deployed models. Compare results alongside model versions, evaluation conditions, and enabled protections.
- Disclosure frameworkopenai.com
Model misalignment reporting framework (Open external website in a new tab)
Explains which behaviors OpenAI aims to disclose, its investigation and reporting process, and remaining uncertainty. A disclosure policy, not an incident register.
Anthropic
Alignment research, model system cards, and transparency materials provide experimental methods, evaluation results, and deployment context.
- Researchalignment.anthropic.com
Alignment Science Blog (Open external website in a new tab)
Experiments, methods, and technical articles from Anthropic alignment researchers on topics including reward hacking, behavioral auditing, oversight, and misalignment.
- Researchwww.anthropic.com
Alignment research team and publications (Open external website in a new tab)
The alignment team’s research agenda and publications, including evaluation, oversight, and stress-testing of safeguards.
- System cards & evaluationswww.anthropic.com
Claude model system cards (Open external website in a new tab)
Model-specific records of capabilities, safety evaluations, and deployment decisions, useful for tracing a risk assessment to its model and test conditions.
- Transparencywww.anthropic.com
Transparency Hub (Open external website in a new tab)
An overview of model reports, safety and trust processes, and public commitments. These summaries do not replace full system cards or specific investigations.