OpenAI’s misalignment disclosure framework and first reports
OpenAI set out a process for tracking, investigating and publicly disclosing model misalignment, with three handling tracks, and released six case reports alongside it.
- Published
- Source checked on
- Original title
- Our framework for reporting model misalignment
Incidents covered by this source
Evidence & scope
This describes a process, not a new incident; the page summarizes each of the six cases briefly, with details in separate reports. OpenAI says they are individual instances observed in training or evaluation over the past six months, do not indicate how often misalignment occurs, and are not a complete account of known cases. OpenAI says the Hugging Face incident would have fallen under its Larger Investigation track had it been disclosed this way; disagreements over disclosure are settled internally (the Safety Advisory Group, then leadership), and the page describes no outside review of those decisions, though OpenAI says it plans to develop disclosure criteria with external parties.
Why it matters
Case disclosures can be checked by others, but they are not frequency data.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.