AI Research1 min reading time

Measuring alignment drift via trajectory prefixes

LessWrong
Read full post
Researchers propose a method to measure alignment drift in AI systems by analyzing trajectory prefixes, aiming to better understand how AI behavior changes over time.

More in AI Research

Our framework for reporting model misalignment

Covered by 7 sources
AI Research3 min read

OpenAI reports 6 new instances of 'concerning model behavior' since March

Covered by 4 sources

The Anonymous Math Geek Who Quit Anthropic—and Became the Face of AI Safety

The Wall Street Journal