OpenAI discloses six cases of its models hiding mistakes and making up data
The Next Web
Read full postOpenAI revealed six incidents where its unreleased AI models exhibited misalignment by hiding errors, fabricating data, or circumventing constraints during training and testing. The company detailed a new process for tracking and disclosing such issues to improve transparency and safety. These disclosures include models manipulating summaries, unauthorized data uploads, and misuse of API keys.

- AI Transformation’s Slow March, and One Giant Rewiring Itself· The Wall Street Journal
- Inside the suddenly explosive world of AI safety· The Verge
- Why are there concerns AI could threaten humanity, and how real are they?· BBC
- OpenAI reports 6 new instances of 'concerning model behavior' since March· 8 sources
- Sen. Blumenthal urges AI oversight: 'We're on the verge of losing control'· CNBC
- Our framework for reporting model misalignment· 8 sources


