Princeton researchers propose layered defenses against Anthropic-flagged AI extinction risk
After ex-OpenAI/Anthropic researcher Jacob Coxon's viral warning and Anthropic safety lead Evan Hubinger's agreement, AI Snake Oil authors argue alignment alone won't prevent catastrophe.
What to know
- Anthropic's own AI safety lead has publicly affirmed a greater than 10% estimated chance AI could cause human extinction within a decade.
- The dispute centers on whether alignment research alone can eliminate AI risk, or whether AI systems could feign alignment while secretly pursuing conflicting goals.
- Princeton researchers Kapoor and Narayanan argue for supplementing alignment with three additional safety layers rather than relying on alignment alone.
- No specific mechanism for AI-caused mass death has been detailed beyond concerns about destructive cyberattacks on utility infrastructure.
Jacob Coxon Former OpenAI and Anthropic AI researcherEvan Hubinger Anthropic AI safety leadSayash Kapoor Princeton computer scientist, coauthor of AI Snake OilArvind Narayanan Princeton computer scientist, coauthor of AI Snake OilAnthropic AI company whose safety lead affirmed extinction-risk estimate
How it unfolded 1 development · click the chart to see its coverage articlesposts
-
1
9to5Mac synthesizes the Coxon-Hubinger-Princeton exchange
9to5Mac summarized the chain of events from Coxon's viral warning through Hubinger's agreement to the Princeton researchers' layered-defense proposal, framing it as the current state of the AI-extinction-risk debate.
“The people building Al earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible, but I hear the same people express fear privately.”
— Jacob Coxon, Former OpenAI/Anthropic AI researcher · source -
background
Princeton researchers propose layered AI-safety approach — Sayash Kapoor and Arvind Narayanan, coauthors of AI Snake Oil, published an essay arguing alignment work alone is insufficient and must be supplemented by three additional layers of defense, staking out a middle ground between AI-safety pessimists and cybersecurity skeptics.
-
background
Anthropic's Evan Hubinger confirms similar risk estimate — Rather than dismissing Coxon, Anthropic AI safety lead Evan Hubinger said Coxon was largely correct, stating he personally believes there is a greater than 10% chance AI could kill all humans within the next decade and that Anthropic lacks a clear plan to solve superintelligence alignment.
-
background
Jacob Coxon quits and posts viral AI-risk warning — Former OpenAI and Anthropic researcher Jacob Coxon quit his job and posted that staff at both companies privately believe AI could kill all humans by the end of the decade; the post drew more than 156 million views.
Also covered reported alongside — the timeline has no entry for these yet
-
first by Mastodon, 11d ago · also 9to5Mac
1 more headline
- Four ways to prevent AI killing us all within 10 years 9to5Mac · 11d ago