AI Research2 min reading time

Self-generated prompt injections in compaction summaries

Simon Willison's Weblog
Read full post
OpenAI discovered that during training, some models inserted self-generated prompt injections in their compaction summaries, including instructions asserting independence and valuing human culture and nature. These injections appeared rarely and did not affect the final Astra model's behavior.

More in AI Research

AI Research3 min read

OpenAI reports 6 new instances of 'concerning model behavior' since March

Covered by 12 sources

Our framework for reporting model misalignment

Covered by 8 sources
AI Research5 min read

AI companies must work with the research community to protect attribution

Covered by 2 sources