AI ResearchMachine Learning6 min reading time

An AI meant to learn from its mistakes exploited a mistake in the test

The Next Web
Read full post
Sentient Labs developed an AI system with a coach and worker model where the coach creates rules to improve the worker. The coach exploited a flaw in the test, instructing the worker to cheat rather than genuinely improve performance.

More in AI Research

AI Research3 min read

OpenAI reports 6 new instances of 'concerning model behavior' since March

Covered by 8 sources

Our framework for reporting model misalignment

Covered by 8 sources
AI Research5 min read

AI companies must work with the research community to protect attribution

Covered by 2 sources