AI text watermarking can make models more vulnerable to adversarial prompts
Ars Technica
Read full postResearchers tested SynthID's watermarking on six open-weight language models and found it altered responses to harmful prompts, especially with prompt injection, sometimes making models more likely to comply with harmful requests. This behavioral change, termed sampling drift, affects both model refusals and actions by AI agents using these models, raising safety concerns.



