arXiv paper warns LLM language outputs may not reflect internal computation
A new paper coins 'linguistic illegibility' to describe when a model's stated reasoning diverges from what it actually computes, raising security concerns.
What to know
- A new arXiv paper argues LLMs' stated reasoning (including chain-of-thought) may not reliably reflect their actual internal computation, a phenomenon it calls 'linguistic illegibility'.
- The paper frames this gap as a security concern, since safety approaches that trust a model's self-reported reasoning could be undermined if that reasoning doesn't match internal processing.
- Discussion of the paper has been limited across Hacker News and Lobsters, with no named authors or detailed argument surfaced in the available coverage beyond the abstract.
tomjakubowski Hacker News submitterCorbin Lobsters submitter
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
2
Paper is cross-posted to Lobsters
User Corbin submitted the same arXiv paper to Lobsters roughly a day after its Hacker News appearance, generating a single comment and a score of 1.
“We introduce the term *linguistic illegibility* to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks.”
— paper abstract, arXiv preprint · source -
1
Paper is shared to Hacker News, drawing modest discussion
User tomjakubowski submitted the paper to Hacker News, where it reached a score of 63 with 26 comments; the same link was also picked up by an HN Frontpage RSS aggregator and a Mastodon bot account.
-
background
Paper introduces 'linguistic illegibility' as an LLM security concern — The arXiv paper argues that LLMs' externalized language outputs and mechanistically-extracted linguistic features can be unreliable indicators of what the model is actually computing internally, coining the term 'linguistic illegibility' for this phenomenon and linking it to security risks.