AIQuiet 13d · day 13
AI researchers warn of deteriorating model alignment detection
Hugh Zhang and Daniel Selsam raise concerns about the difficulty of determining whether advanced AI systems are truly aligned with human values.
What to know
- AI researchers are struggling to verify that advanced models remain aligned with human values as they grow more sophisticated
- The technical challenge of alignment detection is becoming acute as capabilities scale
Hugh Zhang AI researcherDaniel Selsam AI researcher
How it unfolded 1 development · click the chart to see its coverage posts
Sep 15Sep 16Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25Sep 26now · 1:11 AM ET
-
1
Hugh Zhang expresses alignment detection concerns
Hugh Zhang posted on X in agreement with Daniel Selsam about deteriorating ability to verify AI model alignment.
“We are rapidly losing the ability to tell whether our models are aligned.”
— @hughbzhang -
I agree with Daniel Selsam 100%. We are rapidly losing the ability to tell whether our models are aligned.
-