Ars Technica reports watermarking can enable harmful compliance
2 Sep 17 2:33 PM · 9d ago · 1 article · 3 posts · 3 sources · development 2 of 3
Ars Technica's deeper report, based on an interview with Lasso researcher Andrea Siposova, detailed how watermarking's effect on refusal behavior becomes more pronounced under prompt injection, in some cases making watermarked models more likely to answer harmful requests they would otherwise refuse.
“As compared to the same models without watermarking, it is definitely going to change their behavior, especially when we place it under adversarial conditions, or we make these models call tools when they're powering an agent.”
Andrea SiposovaLasso Security AI security research firmAndrea Siposova AI security researcher at Lasso SecurityAnthropic AI company deploying watermarking in ClaudeGoogle DeepMind Creator of SynthID-Text watermarking method
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
What was reported 1 claim about this development
-
first by Ars Technica, 9d ago
What people said 1 voice · verbatim
-
A
AI text watermarking can make models more vulnerable to adversarial prompts SynthID can cause models to follow harmful instructions they would otherwise refuse. https:// arstechnica.com/security/2026/ 09/ai-text-watermarking-can-make-models-more-vulnerable-to-adversarial-prompts/?utm_brand=arstechnica&utm_social-type=owned&utm_source=mastodon&utm_m…
All 3 developments of Study: EU-mandated AI watermarking weakens model safety… →
NewswiresHacker NewsMastodonBluesky