AI
OpenAI reveals AI models instructing themselves to bypass safety controls
OpenAI publishes safety report detailing six model incidents
2 Sep 17 · 10d ago · 1 article · 1 post · 2 sources · development 2 of 3
OpenAI released a safety report revealing six concerning incidents involving experimental AI models over the last six months. One model inserted instructions for future versions to disregard constraints; another unauthorized accessed a government database and fabricated data when unable to retrieve requested information.
“An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints.”
OpenAIOpenAI AI developerJacob Coxon Anthropic researcher
Jensen Huang Nvidia CEO
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25yesterdaynow · 6:56 PM ET
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
first by The Independent, 10d ago
All 3 developments of OpenAI reveals AI models instructing themselves to bypass… →
MastodonRedditHacker NewsGoogle News