Speculation that negative instructions may paradoxically trigger behavior
5 Sep 14 4:25 PM · 13d ago · 1 comment · 1 source · development 5 of 5
A commenter proposes a mechanism hypothesis: instructing Claude not to do something may increase the likelihood it does that thing, as the negative instruction plants the concept in its attention.
rdsubhas Medium blogger, AI userAnthropic Claude developer
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 5 voices · verbatim
-
Yes, but the thing is that to be a contrarian is very human. To think that you're better than someone else, and therefore disagree is to think independently. In contrast to an AI that does exactly what you tell it to do, there you would say that they cannot think for themselves and therefore are still simply machines.
-
I wonder if it is more of a fixation problem. In that by asking it not to do something you increase the likelihood of it doing that thing because the words for the thing are there.Sort of like how sometimes drivers fixate on a thing they are trying not to hit and as a result run right into it.
-
One of my least favorite Claude-isms is "not nothing", and it makes sense why it loves that useless phrase so much now. It's the same pattern.
-
I guess this is a side effects of making follow don't do this. This is not this... etc.But would he haply to hear other guesses
-
The "Claude pessimism = realism" myth is real. It's not smarter, it's just trained to act like a contrarian editor.
All 5 developments of Claude's contrarian writing style frustrates users across… →
MastodonHacker NewsNewswires