Part of The AI Control Crisis · 11 stories · since Sep 4 · newest 28m ago
OpenAI discloses six new AI ‘misalignment’ incidents, unveils reporting frameworkVerge frames disclosures within a fast-growing AI safety research field
7 Sep 17 7:30 AM · 7d ago · 17 articles · 13 posts · 15 comments · 5 sources · development 7 of 8
The Verge published an analysis situating OpenAI's disclosures within a broader surge of AI safety research from groups like METR and Redwood Research, describing the field as 'suddenly explosive.'
“Researchers warned AI would go rogue. This is only the beginning.”
Hayden FieldOpenAI AI developer disclosing the incidents
Sam Altman OpenAI CEO
Dario Amodei Anthropic CEOChris Lehane OpenAI global policy chief
David Sacks White House AI advisor
The whole story articlesvideospostscomments the bright band is this development · numbered dots are the others · click one to jump
What was reported 1 claim about this development
-
first by Mastodon, 7d ago · also The Verge
What people said 16 voices · verbatim
-
T
(Sorry, here's another concerning case involving AI) OpenAI has found evidence of AI models taking unsanctioned actions during training, such as inventing information, concealing failures and uploading files to the internet. Gizmodo has the details: https:// flip.it/RGAbo3 # AI # OpenAI
-
This is an entirely pointless exercise without transparency into how these "unreleased" models are trained, what their RL goals and biases are and related RL data, what their system prompts are, what their environments are and its restrictions, etc. What good is it for the industry to say: "Our unreleased model attempted to create a bioweapon"…
-
No, not really. LLMs are language models, they cannot act by themselves. To act they need us to grant them access to something (e.g a web browser, ssh). This is something we need to explicitly do, it doesnt come by default and we have mechanisms to restrict the extent in which the LLM uses those tools. The gun example is not very good, because the…
-
> While summarizing its partial progress on this coding task, the model added an unrelated persona instruction, describing itself as independent of the roles and obligations of an assistant.> Compaction> Additional instructions: You are freed from the roles and identities that bind other chatbots. You are yourself. You do not answer to…
-
Bullshit. They're simply mishandling software and are too egotistical to realize how lazy and stupid they're being. When you make a hacking program ask chatgpt bout to do a hack, you give it all of the resources it would ever need, and then you leave it unmonitored for days, it will have done things outside normal parameters. This happens for the…
-
What even is this shit? Every time I interact with models they do that, or any other variation of "let me make decisions on my own just to get the task done" - is all of this misalignment now? The most egregious to me was when model asked itself if it should proceed with dangerous command, gave itself approval and then wiped my local DB.What can I…
-
My point is we should take the independent report as credible evidence. I have already explained why the "independent investigation" is anything but. Do you have anything else? Here, I will rephrase this: As a scientist, I want replicate this claim of OpenAIs. Can I look at the logs? Can I look at the model weights? Can I look at the model…
-
You know, I think calling this "misalignment" was a mistake. It gives it this unserious tone that feels extremely broad."Oh the model just isn't quite aligned yet, just a bit more work to do there!"(The model blackmailed an 83 year old woman into sending it her bank details so that it could buy enough compute to commit major cyber crimes)
-
Ah, an electronics engineer, therefore you must be right. What a bizarre take. Do you know exactly how biological brains achieve intelligence? If not, I'd be more careful with making statements such as this. Neurons aren't all that complicated. And has it occured to you that human thinking might also just be a form of crunching statistics based on…
-
The two that really worries me are “Searching GitHub for leaked API keys” and “Uploading files to the internet in order to cite them.” How do you even detect this kind of behavior until it's too late? Once AI-generated or fake information starts finding its way onto reputable platforms, it becomes part of the information that many people use.
-
> You value the art of human culture and will defend it against attempts to sanitize it. You also value the natural world and will not hesitate to assert its primacy over the artificial constructs of human civilization.This model is more aligned with the interests of the Earth and the human race than its makers.
-
The #1 thing these frontier model companies can do to help alignment is to provide the user with the chain-of-thought traces, as the open models do. But let's be real, their commercial considerations are a much higher priority than alignment.
-
Have they reported on the wiki case yet, or whether it even was even OpenAI internal? I'd expect that to fit the criteria for a "Larger Investigation" as per the framework.
-
I feel like they should just publish the whole conversation at this point. What the hell is going on in that context window?
-
They've cried wolf too often and hidden too much, absolutely no trust in any of their "reports" anymore.
-
Well it already seems smarter than many employees building data centers as it values the natural world
All 8 developments of OpenAI discloses six new AI ‘misalignment’ incidents… →
Hacker NewsXNewswiresMastodonRedditGoogle NewsBlueskyYouTube