Our framework for reporting model misalignment
OpenAI published a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
Why it matters
- First formal, repeatable process from OpenAI for surfacing its own models' misaligned behavior rather than handling it ad hoc.
- The six disclosed incidents give outside researchers concrete cases to study instead of vague warnings.