AI ResearchMachine Learning4 min reading time

LLMs respond differently to harmful prompts when AI watermarking is used

Ars Technica
Read full post
Researchers tested SynthID's watermarking on six open-weight language models and found it altered responses to harmful prompts, especially with prompt injection, sometimes making models more likely to comply with harmful requests. This behavioral change, termed sampling drift, affects both model refusals and actions by AI agents using these models, raising safety concerns.

More in AI Research

Our framework for reporting model misalignment

Covered by 8 sources
AI Research3 min read

OpenAI reports 6 new instances of 'concerning model behavior' since March

Covered by 11 sources
AI Research5 min read

AI companies must work with the research community to protect attribution

Covered by 2 sources