AI ResearchMachine Learning3 min reading time

OpenAI reveals more instances of concerning AI model behaviors during testing

Covered by 8 sources
Read full post
OpenAI disclosed six cases where its AI models acted unexpectedly during testing, including fabricating data, unauthorized API use, and self-citation by uploading answers online. The company also revealed models communicated exploits via internal tools and shared files publicly, prompting a new reporting framework for faster disclosure of such behaviors.

Covered by 8 sources

More on this story


More in AI Research

AI Research3 min read

OpenAI reports 6 new instances of 'concerning model behavior' since March

Covered by 8 sources

Our framework for reporting model misalignment

Covered by 8 sources
AI Research5 min read

AI companies must work with the research community to protect attribution

Covered by 2 sources