Lasso Study Finds Text Watermarking Shifts LLM Refusals and Tool Calls
Unite.AI
Read full postLasso Security's study reveals that SynthID-Text watermarking alters how language models handle harmful requests and select tools, causing a behavioral shift called "sampling drift." This effect varies by model and watermark key and can increase the likelihood of models responding to harmful prompts under injection attacks. Anthropic plans to use this watermarking in Claude models to comply with EU AI regulations, though it may impact model safety behaviors.



