UC Berkeley researcherControl investments more effective than alignment
2 Sep 16 · 11d ago · 1 article · 2 posts · 3 sources · development 2 of 2
Sayash Kapoor argues that marginal investments in AI control are more likely to be effective than those in alignment, citing the incidents as evidence of companies failing to apply known control techniques despite their availability.
“marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques.”
Sayash Kapoor
Dario Amodei Anthropic CEOKatie Moussouris CEO, Luta SecuritySayash Kapoor AI researcher, incoming UC Berkeley professorAvery Pennarun CEO, Tailscale security company
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
first by TechCrunch, 11d ago
What people said 3 voices · verbatim
-
T
The fall-out from agentic AI outbreaks has mainly focused on alignment, but cybersecurity experts tell me that the frontier labs have much more to do simply securing and monitoring their agents
-
T
There may be a simpler and more effective fix for rogue agents, hiding in plain sight. https:// techcrunch.com/2026/09/16/ai-l abs-want-in-house-auditors-but-maybe-they-should-shut-the-front-door-first/?utm_source=dlvr.it&utm_medium=mastodon
-
T
This makes a lot of sense when testing #AI models #artificialintelligence #huggingface #cybersecurity #sandbox
All 2 developments of Security experts: AI labs need basic network controls, not… →
MastodonNewswiresBlueskyGoogle NewsHacker NewsReddit