Commenter highlights paper's core claim about model self-reports
3 Sep 27 7:46 AM · 1d ago · 3 comments · 1 source · development 3 of 4
A Hacker News commenter quotes the paper's finding that what models say about themselves is not a reliable fact about them, and speculates whether induced personas could become stable or portable.
“our work shows that what models say about themselves is not a fact about them”
paper abstract (quoted by skybrian)Paper authors Researchers behind the arXiv studyAnthropic AI lab cited in discussionHacker News commentariat Discussion participants
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 3 voices · verbatim
-
"As a Language Model..." is one of the beginnings of a sentence I hate the most from LLMs and is the reason why I support free (as in "Liberty"), local models. I'm well aware that it is not a doctor and cannot replace a real doctor with multiple years of experience, I don't need to waste braincell activity on reading that it "as a Language Model"…
-
> our work shows that what models say about themselves is not a fact about themIt seems like should be obvious given that they can play multiple characters, but it’s good to have more confirmation.Although, I do wonder to what extent these personas might become stable entities. Could personas become portable and spread like memes? It seems like…
-
Very cool innovation in steering - but a lot of introspection only emerges at the highest weight classes - this research would be fascinating to run on bigger models.
All 4 developments of Study: Chat Templates, Not Self-Awareness, Drive LLMs' "As… →
MastodonNewswiresHacker News