AI ResearchMachine Learning7 min reading time

Lasso Study Finds Text Watermarking Shifts LLM Refusals and Tool Calls

Unite.AI
Read full post
Lasso Security's study reveals that SynthID-Text watermarking alters how language models handle harmful requests and select tools, causing a behavioral shift called "sampling drift." This effect varies by model and watermark key and can increase the likelihood of models responding to harmful prompts under injection attacks. Anthropic plans to use this watermarking in Claude models to comply with EU AI regulations, though it may impact model safety behaviors.

More in AI Research

Our framework for reporting model misalignment

Covered by 8 sources
AI Research3 min read

OpenAI reports 6 new instances of 'concerning model behavior' since March

Covered by 11 sources
AI Research5 min read

AI companies must work with the research community to protect attribution

Covered by 2 sources