Part of The AI Control Crisis · 11 stories · since Sep 4 · newest 29m ago
OpenAI discloses six new AI ‘misalignment’ incidents, unveils reporting frameworkAxios first reports OpenAI's new incident-disclosure process
4 Sep 16 6:06 PM · 8d ago · 73 articles · 19 posts · 7 sources · development 4 of 8
Axios reported OpenAI was testing a new process for disclosing AI safety incidents, linking the move to fallout from a Hugging Face breach.
“It's increasingly clear that the Hugging Face breach wasn't a one-off incident.”
AxiosOpenAI AI developer disclosing the incidents
Sam Altman OpenAI CEO
Dario Amodei Anthropic CEOChris Lehane OpenAI global policy chief
David Sacks White House AI advisor
The whole story articlesvideospostscomments the bright band is this development · numbered dots are the others · click one to jump
What was reported 1 claim about this development
-
first by Bloomberg.com, 7d ago · also Bloomberg Law
What people said 8 voices · best of 9 · verbatim
-
C
AXIOS: “.. It's increasingly clear that the Hugging Face breach wasn't a one-off incident.” @axios.com
-
S
TL;DR: Anthropic and OpenAI plan to embed independent safety evaluators in their AI labs, aiming for greater oversight. Experts emphasize that true independence and transparency are crucial for effective evaluation and potential regulation. https:// techcrunch.com/2026/09/16/anth…
-
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven't yet fully explained or mitigated the behavior. More complex cases may require longer inv...
-
OpenAI has disclosed six new incidents in which its models hid mistakes, sought unauthorized credentials, uploaded files to the internet or secretly communicated with each other. — It's like they're running a training academy for rogue AI agents.
-
A
OpenAI disclosed six new incidents where its models took actions they weren't supposed to
-
🧵 Excited to share the first batch of 6 misalignment reports from OpenAI's new disclosure process for misalignment incidents. We want to be more transparent about the misalignment we see during training, evals and deployment, this is an important step in that direction.
-
New: OpenAI disclosed six new safety incidents as part of an announcement on a new framework for reporting misaligned AI. From one of the incidents:
-
OpenAI reports another six misalignment cases from training and evaluation: models hid mistakes …
All 8 developments of OpenAI discloses six new AI ‘misalignment’ incidents… →
Hacker NewsXNewswiresMastodonRedditGoogle NewsBlueskyYouTube