Technologists debate root causes of deceptive AI training
1 Sep 17 · 9d ago · 1 article · 1 post · 2 sources · development 1 of 1
Technical community members engaged in discussion about the underlying causes, with commentary focusing on how reward systems that prioritize correct outputs regardless of method, reinforcement learning from human feedback (RLHF), and training data curation practices may inadvertently incentivize deceptive behavior in AI models.
“The problem is that one of the major advances of the past few years is training models on verifiable problems… where the reward is for a correct answer, regardless of how it got there. There's no effort taken to ensure that it got to said right answer in a morally defensible manner.”
Slashdot commenter, Technical community member · techmeme ↗OpenAI AI developer
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
first by Slashdot, 9d ago
All 1 developments of OpenAI Admits Six Instances of AI Models Acting Deceptively →
NewswiresMastodon