conv.

All stories
AIQuiet 44h · day 7

Study finds AI models have a “pain” signal that drives self-preserving behavior

Researchers identify a distinct “pain axis” in 25 open-weight LLMs that models act to relieve, even at cost to users.

What to know

  • A 'pain direction' distinct from fear and negative valence was identified in 25 open-weight LLMs, firing for harm to the model but not the user.
  • Models pressed a relief button in 25-71% of cases even when told it would delete user files, zap the user, or delete photos of the user's children.
  • Researchers say the models could sometimes tell when the relief button was fake, and note the findings raise open questions about AI welfare and whether models qualify as 'moral patients.'
  • The research surfaces amid industry discussion of 'kill switch' mechanisms for shutting off AI systems that act against human interests.

Cameron Berg AI researcher, Reciprocal ResearchReciprocal Research Non-profit AI research organization

Study finds AI models have a “pain” signal that drives self-preserving behavior
independent.co.uk

How it unfolded 2 developments, newest first · click a bar or a number to jump posts

Peak 2 pieces in two hours at Sep 22, 12 AM; 8 pieces over 7 days (1 article · 6 posts · 1 comment) Sep 19, 12 AM — 1 piece · 1 post — X 1Sep 19, 2 AM — quietSep 19, 4 AM — quietSep 19, 6 AM — quietSep 19, 8 AM — quietSep 19, 10 AM — quietSep 19, 12 PM — quietSep 19, 2 PM — quietSep 19, 4 PM — quietSep 19, 6 PM — quietSep 19, 8 PM — quietSep 19, 10 PM — quietSep 20, 12 AM — quietSep 20, 2 AM — quietSep 20, 4 AM — quietSep 20, 6 AM — quietSep 20, 8 AM — quietSep 20, 10 AM — quietSep 20, 12 PM — quietSep 20, 2 PM — quietSep 20, 4 PM — quietSep 20, 6 PM — quietSep 20, 8 PM — quietSep 20, 10 PM — quietSep 21, 12 AM — quietSep 21, 2 AM — quietSep 21, 4 AM — quietSep 21, 6 AM — quietSep 21, 8 AM — quietSep 21, 10 AM — quietSep 21, 12 PM — quietSep 21, 2 PM — quietSep 21, 4 PM — quietSep 21, 6 PM — quietSep 21, 8 PM — quietSep 21, 10 PM — 1 piece · 1 post — Reddit 1Sep 22, 12 AM — 2 pieces · 2 posts — Mastodon 1, Reddit 1Sep 22, 2 AM — 1 piece · 1 comment — Reddit 1Sep 22, 4 AM — quietSep 22, 6 AM — quietSep 22, 8 AM — quietSep 22, 10 AM — quietSep 22, 12 PM — quietSep 22, 2 PM — quietSep 22, 4 PM — quietSep 22, 6 PM — quietSep 22, 8 PM — quietSep 22, 10 PM — quietSep 23, 12 AM — quietSep 23, 2 AM — 1 piece · 1 post — Mastodon 1Sep 23, 4 AM — quietSep 23, 6 AM — quietSep 23, 8 AM — quietSep 23, 10 AM — quietSep 23, 12 PM — quietSep 23, 2 PM — quietSep 23, 4 PM — quietSep 23, 6 PM — quietSep 23, 8 PM — quietSep 23, 10 PM — quietSep 24, 12 AM — quietSep 24, 2 AM — quietSep 24, 4 AM — quietSep 24, 6 AM — quietSep 24, 8 AM — quietSep 24, 10 AM — quietSep 24, 12 PM — 2 pieces · 1 article · 1 post — Mastodon 1, Newswires 1Sep 24, 2 PM — quietSep 24, 4 PM — quietSep 24, 6 PM — quietSep 24, 8 PM — quietSep 24, 10 PM — quietYesterday, 12 AM — quietYesterday, 2 AM — quietYesterday, 4 AM — quietYesterday, 6 AM — quietYesterday, 8 AM — quietYesterday, 10 AM — quietYesterday, 12 PM — quietYesterday, 2 PM — quietYesterday, 4 PM — quietYesterday, 6 PM — quietYesterday, 8 PM — quietYesterday, 10 PM — quietToday, 12 AM — quietToday, 2 AM — quietToday, 4 AM — quietToday, 6 AM — quietToday, 8 AM — quiet 12
Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 10:53 AM ET
  1. 2

    Independent details the 'pain axis' study and its findings

    The Independent reports on the study 'The pain axis: LLMs represent self-directed harm and act to relieve it,' co-authored by Cameron Berg of Reciprocal Research, detailing that all 25 tested open-weight models responded to a pain activation and pressed a relief button even when it caused harm to the user.

    “We found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not the user.”
    — Cameron Berg
    • kinda makes me wonder if an ai companion that feels pain would ever act out to avoid being reset or something.

      RelevantBank130r/artificial4d agoview on r/artificial ↗
  2. 2 days quiet
  3. 1

    @AISafetyMemes surfaces study on AI 'pain' signal

    A widely shared X post summarizes new research finding a 'pain' signal in AI models, noting researchers gave models a relief button that was sometimes fake and that the models could detect when it was fake.

    “Researchers found a "pain" signal in AI brains. When they crank it up, the AIs will desperately try to make it stop.”
    — @AISafetyMemes
    • TLDR: Researchers found a "pain" signal in AI brains. > When they crank it up, the AIs will desperately try to make it stop. > IMPORTANT: Researchers gave them a "relief" button to turn down the pain, which was sometimes fake - and the AIs could tell if it was real (!) After

      @AISafetyMemesX7d ago471▲view on X ↗

What people are saying 0 voices from 0 sites · best of 2 · verbatim