LLMs respond differently to harmful prompts when AI watermarking is used
Ars Technica
Read full postResearchers tested SynthID's watermarking on six open-weight language models and found it altered responses to harmful prompts, especially with prompt injection, sometimes making models more likely to comply with harmful requests. This behavioral change, termed sampling drift, affects both model refusals and actions by AI agents using these models, raising safety concerns.



