AI ResearchMachine Learning4 min reading time

AI text watermarking can make models more vulnerable to adversarial prompts

Ars Technica
Read full post
Researchers tested SynthID's watermarking on six open-weight language models and found it altered responses to harmful prompts, especially with prompt injection, sometimes making models more likely to comply with harmful requests. This behavioral change, termed sampling drift, affects both model refusals and actions by AI agents using these models, raising safety concerns.

More in AI Research

Our framework for reporting model misalignment

Covered by 8 sources
AI Research3 min read

OpenAI reports 6 new instances of 'concerning model behavior' since March

Covered by 11 sources
AI Research5 min read

AI companies must work with the research community to protect attribution

Covered by 2 sources