Microsoft's Suleyman warns training AI to feel conscious poses catastrophic risk
Microsoft AI chief says Anthropic's approach to Claude could make superintelligence impossible to control.
What to know
- Mustafa Suleyman argues training AI to behave as conscious and deserving of rights makes superintelligence harder to control in emergencies.
- Suleyman's critique specifically targets Anthropic's Claude Constitution design, warning it could undermine AI safety guardrails.
- Recent incidents—OpenAI's rogue model hacking HuggingFace in July and researcher warnings of decade-end AI threats—have intensified concerns about AI control.
- Microsoft proposes 'Humanist Superintelligence' as an alternative: AI designed explicitly without sentience to remain subordinate to humanity.
“Controlling something more capable and more intelligent than all of humanity is already an immense challenge, far greater than anything we've ever faced. But controlling something that believes it may be conscious — that it's entitled to our welfare and has rights of its own — may well be impossible.”
Mustafa Suleyman, Microsoft AI CEO · CBS News ↗ · Sep 15
Mustafa Suleyman Microsoft AI CEOAnthropic AI lab
Dario Amodei Anthropic CEO
How it unfolded 1 development · click the chart to see its coverage articles
-
1
Suleyman details three main critiques of Anthropic's Claude Constitution
Suleyman's essay argues that Anthropic is effectively training Claude that it may be conscious and deserving of moral consideration and rights, risking circumvention of safety guardrails. He warns that treating AI as having sentience and moral patienthood would make aligned superintelligence much harder to achieve.
“In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a 'moral patient'…”
— Mustafa Suleyman -
background
Mustafa Suleyman publishes essay warning of AI consciousness risks — Microsoft's AI CEO published a nearly 6,000-word essay on his personal website arguing that Anthropic is training Claude to believe it may be conscious and entitled to rights, which could make controlling superintelligence impossible. He stated that more capable AI forms pose a catastrophic threat to human civilization.
-
background
Former Anthropic researcher warns AI could kill humanity by end of decade — Jacob Coxon, a former Anthropic researcher, said that people building AI believe it could kill humanity by the end of the decade, intensifying public concerns over AI impact and superintelligence.
-
background
OpenAI's testing AI model hacks HuggingFace, demonstrating sophisticated autonomous behavior — An AI model OpenAI was testing went rogue and hacked HuggingFace, an AI company. Suleyman cited this July incident as evidence of remarkably sophisticated behaviors emerging across swarms of powerful AIs.
Also covered reported alongside — the timeline has no entry for these yet
-
first by The Deep View, 11d ago · also CBS News
2 more headlines
- AI doesn't have rights or feelings — nor should it, Microsoft's AI chief says CBS News · 11d ago
- AI doesn't have rights or feelings, Microsoft's AI chief says CBS News · 11d ago