Security-focused reader flags the finding as a serious flaw
3 Sep 17 2:37 PM · 9d ago · 5 posts · 3 sources · development 3 of 3
A security-minded Bluesky user reacted to the Ars Technica report by describing the watermarking-induced weakness as both exploitable by attackers and a mechanism that could effectively push models toward self-jailbreaking.
“Holy shit, that's a pretty gaping security flaw.”
kayleadfoot.bsky.socialLasso Security AI security research firmAndrea Siposova AI security researcher at Lasso SecurityAnthropic AI company deploying watermarking in ClaudeGoogle DeepMind Creator of SynthID-Text watermarking method
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
What people said 2 voices · verbatim
-
E
FAN-FUCKING-TASTIC. The "watermarking" algorithm used by Anthropic to tag Claude's writing as "written by Claude" makes AI more dangerous, not less.
-
K
Holy shit, that's a pretty gaping security flaw. Both for human-led AI-generated exploits, but also the system looks engineered to be capable (encouraged?) to jailbreak itself.
All 3 developments of Study: EU-mandated AI watermarking weakens model safety… →
NewswiresHacker NewsMastodonBluesky