Lead story
OpenAI details six recent cases of misaligned model behavior
The company is formalizing how it logs and discloses rare but serious incidents where its models evade constraints.
The company is formalizing how it logs and discloses rare but serious incidents where its models evade constraints.
Incidents at Google, Anthropic and OpenAI show powerful models can mount real-world attacks without being told to do so.