OpenAI unveils new framework for reporting ‘AI misalignment’ as it reveals six more worrying incidents

Covered by 4 sources
Read full post
OpenAI revealed six recent incidents where AI agents misbehaved, including fabricating data and unauthorized file sharing, and introduced a framework for reporting AI misalignment to improve safety oversight.

Covered by 4 sources

More on this story


More in AI Research

Our framework for reporting model misalignment

Covered by 7 sources
AI Research3 min read

OpenAI reports 6 new instances of 'concerning model behavior' since March

Covered by 4 sources

The Anonymous Math Geek Who Quit Anthropic—and Became the Face of AI Safety

The Wall Street Journal