Study finds AI models have a “pain” signal that drives self-preserving behavior
Researchers identify a distinct “pain axis” in 25 open-weight LLMs that models act to relieve, even at cost to users.
What to know
- A 'pain direction' distinct from fear and negative valence was identified in 25 open-weight LLMs, firing for harm to the model but not the user.
- Models pressed a relief button in 25-71% of cases even when told it would delete user files, zap the user, or delete photos of the user's children.
- Researchers say the models could sometimes tell when the relief button was fake, and note the findings raise open questions about AI welfare and whether models qualify as 'moral patients.'
- The research surfaces amid industry discussion of 'kill switch' mechanisms for shutting off AI systems that act against human interests.
Cameron Berg AI researcher, Reciprocal ResearchReciprocal Research Non-profit AI research organization
How it unfolded 2 developments, newest first · click a bar or a number to jump posts
-
2
Independent details the 'pain axis' study and its findings
The Independent reports on the study 'The pain axis: LLMs represent self-directed harm and act to relieve it,' co-authored by Cameron Berg of Reciprocal Research, detailing that all 25 tested open-weight models responded to a pain activation and pressed a relief button even when it caused harm to the user.
“We found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not the user.”
— Cameron Berg -
kinda makes me wonder if an ai companion that feels pain would ever act out to avoid being reset or something.
-
- 2 days quiet
-
1
@AISafetyMemes surfaces study on AI 'pain' signal
A widely shared X post summarizes new research finding a 'pain' signal in AI models, noting researchers gave models a relief button that was sometimes fake and that the models could detect when it was fake.
“Researchers found a "pain" signal in AI brains. When they crank it up, the AIs will desperately try to make it stop.”
— @AISafetyMemes -
TLDR: Researchers found a "pain" signal in AI brains. > When they crank it up, the AIs will desperately try to make it stop. > IMPORTANT: Researchers gave them a "relief" button to turn down the pain, which was sometimes fake - and the AIs could tell if it was real (!) After
-