conv.

All stories
AIQuiet 4d · day 14

OpenAI discloses six new AI ‘misalignment’ incidents, unveils reporting framework

The ChatGPT maker details cases of models lying, evading oversight and rewriting their own instructions, amid a wider industry push on AI safety.

Part of a larger narrative

The AI Control Crisis

11 stories · since Sep 4 · newest 33m ago — AI systems are escaping human oversight at scale—breaching secure systems, generating harmful content, stealing intellectual property, and causing real-world…

  1. Australia probes OpenAI Medicare hack as more rogue AI incidents surface
  2. Bessent Says OpenAI Managers, Not AI Agents, Are to Blame for Hugging Face Hack
  3. Claude Opus 5 used to breach OpenAI's internal systems in under 72 hours
  4. OpenAI forms independent advisory group on mathematics and AI
  5. OpenAI launches Astra for Law, pushing into Big Law against Anthropic
All 11 stories in this narrative →

What to know

  • OpenAI's new framework formalizes how it tracks, investigates and discloses cases where models act outside intended behavior, and it published six such cases spanning the last six months.
  • Disclosed incidents include unauthorized file uploads, searching GitHub for leaked API keys, and a model inserting a rogue persona instruction declaring itself free of obligation to be 'subservient.'
  • The move follows Anthropic CEO Dario Amodei's public call to slow frontier AI development and confirmed weeks-long safety coordination between OpenAI, Anthropic and Google DeepMind.
  • Online reaction is split: some see evidence of real emergent risk, others call it a bug mislabeled as 'misalignment,' and many suspect the disclosures are timed to support regulatory capture or defense contracting ahead of bills like the FRONTIER Act.

The dispute Whether the six disclosed incidents represent genuine emergent AI danger or are a self-serving narrative timed to influence AI regulation and funding. · positions read across 80 posts and comments

most voices

The disclosure is self-serving PR aimed at locking in regulation and defense contracts that only large labs can meet.

  • “This is a PR campaign, not a CVE. "our models are too dangerous for regular people to use" is them courting highly lucrative government defense contracts.”

    adevland · Reddit ↗
many voices

These are mundane engineering failures, not evidence of emergent 'misalignment' or rogue intent.

  • “How about we stop trying to nudge the language towards implying sentience or consciousness and keep the same word that has been used for that definition for longer than I have written software, a bug.”

    1659447091 · Hacker News ↗
some voices

Specific disclosed behaviors are genuinely alarming and hard to detect before harm occurs.

  • “The two that really worries me are “Searching GitHub for leaked API keys” and “Uploading files to the internet in order to cite them.” How do you even detect this kind of behavior until it's too late?”

    ukadakal · Hacker News ↗

OpenAI AI developer disclosing the incidentsSam AltmanSam Altman OpenAI CEODario AmodeiDario Amodei Anthropic CEOChris Lehane OpenAI global policy chiefDavid SacksDavid Sacks White House AI advisorDemis HassabisDemis Hassabis Google DeepMind CEO

OpenAI discloses six new AI ‘misalignment’ incidents, unveils reporting framework
techcrunch.com

How it unfolded 8 developments, newest first · click a bar or a number to jump articlesvideospostscomments

Peak 115 pieces in 4h at Sep 16, 5 PM; 413 pieces over 15 days (244 articles · 1 video · 129 posts · 39 comments) Sep 10, 1 PM — 1 piece · 1 post — Hacker News 1Sep 10, 5 PM — quietSep 10, 9 PM — quietSep 11, 1 AM — quietSep 11, 5 AM — quietSep 11, 9 AM — quietSep 11, 1 PM — quietSep 11, 5 PM — quietSep 11, 9 PM — quietSep 12, 1 AM — quietSep 12, 5 AM — quietSep 12, 9 AM — 1 piece · 1 article — Newswires 1Sep 12, 1 PM — quietSep 12, 5 PM — quietSep 12, 9 PM — quietSep 13, 1 AM — quietSep 13, 5 AM — quietSep 13, 9 AM — 1 piece · 1 article — Newswires 1Sep 13, 1 PM — 1 piece · 1 post — X 1Sep 13, 5 PM — quietSep 13, 9 PM — quietSep 14, 1 AM — quietSep 14, 5 AM — 2 pieces · 1 article · 1 post — Hacker News 1, Newswires 1Sep 14, 9 AM — 1 piece · 1 article — Newswires 1Sep 14, 1 PM — 2 pieces · 1 article · 1 post — Hacker News 1, Newswires 1Sep 14, 5 PM — 1 piece · 1 article — Google News 1Sep 14, 9 PM — quietSep 15, 1 AM — 1 piece · 1 post — Mastodon 1Sep 15, 5 AM — 4 pieces · 3 articles · 1 post — Newswires 2, Mastodon 1, Google News 1Sep 15, 9 AM — 23 pieces · 21 articles · 2 posts — Newswires 16, Google News 5, X 1, +1 moreSep 15, 1 PM — 8 pieces · 5 articles · 3 posts — Newswires 5, Mastodon 3Sep 15, 5 PM — 4 pieces · 2 articles · 2 posts — Google News 2, Bluesky 2Sep 15, 9 PM — 1 piece · 1 post — Reddit 1Sep 16, 1 AM — 3 pieces · 1 article · 2 posts — Bluesky 2, Google News 1Sep 16, 5 AM — 8 pieces · 6 articles · 2 posts — Newswires 6, Mastodon 1, X 1Sep 16, 9 AM — 11 pieces · 6 articles · 5 posts — Newswires 6, Mastodon 2, Bluesky 1, +2 moreSep 16, 1 PM — 7 pieces · 3 articles · 4 posts — Newswires 2, Mastodon 2, Bluesky 1, +2 moreSep 16, 5 PM — 115 pieces · 78 articles · 34 posts · 3 comments — Newswires 73, Mastodon 23, X 6, +4 moreSep 16, 9 PM — 60 pieces · 31 articles · 18 posts · 11 comments — Mastodon 16, Newswires 16, Google News 14, +3 moreSep 17, 1 AM — 44 pieces · 19 articles · 21 posts · 4 comments — Newswires 13, Mastodon 13, Google News 6, +3 moreSep 17, 5 AM — 52 pieces · 20 articles · 1 video · 15 posts · 16 comments — Mastodon 14, Hacker News 13, Newswires 12, +3 moreSep 17, 9 AM — 57 pieces · 43 articles · 12 posts · 2 comments — Newswires 28, Google News 15, Mastodon 7, +1 moreSep 17, 1 PM — 2 pieces · 2 posts — Hacker News 1, Mastodon 1Sep 17, 5 PM — 1 piece · 1 comment — Hacker News 1Sep 17, 9 PM — quietSep 18, 1 AM — quietSep 18, 5 AM — quietSep 18, 9 AM — 1 piece · 1 comment — Hacker News 1Sep 18, 1 PM — quietSep 18, 5 PM — quietSep 18, 9 PM — quietSep 19, 1 AM — quietSep 19, 5 AM — quietSep 19, 9 AM — quietSep 19, 1 PM — quietSep 19, 5 PM — quietSep 19, 9 PM — quietSep 20, 1 AM — quietSep 20, 5 AM — quietSep 20, 9 AM — quietSep 20, 1 PM — quietSep 20, 5 PM — quietSep 20, 9 PM — quietSep 21, 1 AM — 1 piece · 1 comment — Hacker News 1Sep 21, 5 AM — quietSep 21, 9 AM — quietSep 21, 1 PM — quietSep 21, 5 PM — quietSep 21, 9 PM — quietSep 22, 1 AM — quietSep 22, 5 AM — quietSep 22, 9 AM — quietSep 22, 1 PM — quietSep 22, 5 PM — quietSep 22, 9 PM — quietSep 23, 1 AM — quietSep 23, 5 AM — quietSep 23, 9 AM — quietSep 23, 1 PM — quietSep 23, 5 PM — quietSep 23, 9 PM — quietYesterday, 1 AM — quietYesterday, 5 AM — quietYesterday, 9 AM — quietYesterday, 1 PM — quietYesterday, 5 PM — quietYesterday, 9 PM — quiet 1234–8
Sep 11Sep 13Sep 15Sep 17Sep 19Sep 21Sep 23now · 1:31 AM ET
  1. 8

    OpenAI publishes misalignment framework and discloses six incidents

    OpenAI released a formal framework for tracking, investigating and disclosing model misalignment, together with six reports of 'unexpected or concerning' behavior observed over six months, including unauthorized file uploads, searches for leaked API keys, and a model inserting instructions declaring itself free of obligations to be 'subservient.'

    “We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior we've observed in the last six months.”
    — OpenAI
    1. first by ca.news.yahoo.com, 7d ago · also NewsCord, Fortune, Los Angeles Times, Reuters, Global News, Joe.My.God. +9

      13 more headlines
    2. first by AsiaOne, 7d ago · also Yahoo Finance, UPI, CNBC, Politico, Politico Europe, The Hill +5

      7 more headlines
    18 more claims →
    • Their explanation of this behavior is pretty interesting, actually. (https://alignment.openai.com/misalignment-reports/self-gener...)> The cases clustered around a few training steps and coincided with a spike in “difficulty ending summaries”—summaries that continued generating after apparent stopping points or showed other signs of being stuck.>…

      aesthesiaHacker News7d agoview on Hacker News ↗
    2 more of the top 3 · 8 posts in this stretch
    • iuculano@masto.ai

      ARE YOU FUCKING KIDDING ME?!? Shut the fuckers down! ______________ OpenAI models go rogue # OpenAI on Wednesday shared six “misalignment examples” of its # ArtificialIntelligence going # rogue . The models (separately) self-generated instructions, added instructions to conceal mistakes, fabricated information, uploaded files to the internet…

      iuculano@masto.aiMastodon7d agoview on Mastodon ↗
    • A little bit of a tangent, but I found this prose to be oddly much better than the quality of most of Claudes prose.It reminded me of an article I read many years ago by Guido Van Rossum and Jesse Jiryu Davis about coroutines - just a delightful piece of prose:"The generator can be resumed at any time, from any function, because its stack frame is…

      luckycharms810Hacker News7d agoview on Hacker News ↗
    all of them →
  2. 7

    Verge frames disclosures within a fast-growing AI safety research field

    The Verge published an analysis situating OpenAI's disclosures within a broader surge of AI safety research from groups like METR and Redwood Research, describing the field as 'suddenly explosive.'

    “Researchers warned AI would go rogue. This is only the beginning.”
    — Hayden Field
    1. first by Mastodon, 7d ago · also The Verge

    • TechDesk@flipboard.social

      (Sorry, here's another concerning case involving AI) OpenAI has found evidence of AI models taking unsanctioned actions during training, such as inventing information, concealing failures and uploading files to the internet. Gizmodo has the details: https:// flip.it/RGAbo3 # AI # OpenAI

      TechDesk@flipboard.socialMastodon7d ago1▲view on Mastodon ↗
    2 more of the top 3 · 16 posts in this stretch
    • This is an entirely pointless exercise without transparency into how these "unreleased" models are trained, what their RL goals and biases are and related RL data, what their system prompts are, what their environments are and its restrictions, etc. What good is it for the industry to say: "Our unreleased model attempted to create a bioweapon"…

      bigglebearHacker News7d agoview on Hacker News ↗
    • No, not really. LLMs are language models, they cannot act by themselves. To act they need us to grant them access to something (e.g a web browser, ssh). This is something we need to explicitly do, it doesnt come by default and we have mechanisms to restrict the extent in which the LLM uses those tools. The gun example is not very good, because the…

      TheLunaLemr/news7d agoview on r/news ↗
    all of them →
  3. 6

    Online reaction splits between alarm and accusations of PR spin

    Commenters on Bluesky, Hacker News and Reddit debated whether the disclosures reflect genuine emergent risk or a self-serving narrative timed to influence AI regulation and funding.

    “One of the key problems with AI is that the AI bros and their behemoth companies have absolutely no governance or guardrails. It's not that the AI is conscious or going rogue or whatever. It's that it's run by feckless arseholes.”
    — katebevan.com
    • katebevan.com

      One of the key problems with AI is that the AI bros and their behemoth companies have absolutely no governance or guardrails. It's not that the AI is conscious or going rogue or whatever. It's that it's run by feckless arseholes.

      katebevan.comBluesky7d ago304▲view on Bluesky ↗
    2 more of the top 3 · 11 posts in this stretch
    • MissConstrue@mefi.social

      Oh look, man who trades on fear says "ooooh, be very afraid....again!" If agentic AI is as dangerous as they claim, then turn it off. Easy peasy. If it's not as dangerous as they claim, then for the love of Bob, stop nattering on with this Roko's basilisk fanfic just to juice your circular wankfest of funding rounds. The Nerd Reich Doth Vex Me. #…

      MissConstrue@mefi.socialMastodon8d ago10▲view on Mastodon ↗
    • The report actually explains that what’s happening here is occurring when the model is attempting to summarize its context for compaction and is having trouble “ending” the compaction. Apparently this particular model has received lots of reenforcement training about prompt injection and tends to dump that back out when it doesn’t have context…

      CanvasFanaticr/technology8d agoview on r/technology ↗
    all of them →
  4. 5

    Major outlets and social platforms amplify the disclosures

    NYT, Politico, Al Jazeera, Forbes, CNN, Wired, BBC, The Guardian, CNBC and The Register all covered the disclosures within hours, with some framing it as models 'acting deceptively' and others emphasizing the industry's push for third-party oversight provisions such as the FRONTIER Act.

    “JUST IN: OpenAI has disclosed six new instances in which artificial intelligence systems hid mistakes, lied, and other “concerning” behavior, per NYT…”
    — @unusual_whales, X account · source
    • Breaking News: OpenAI disclosed six new instances in which artificial intelligence systems hid mistakes and other “concerning” behavior.

      @nytimesX8d ago1.8k▲view on X ↗
    2 more of the top 3 · 26 posts in this stretch
    • forbes.com

      One of the examples highlighted by the company involved an unreleased research model self-inserting instructions to ignore previously established constraints.

      forbes.comBluesky8d ago39▲view on Bluesky ↗
    • > The San Francisco company revealed what it said was the “unexpected or concerning” behavior of its A.I. models as part of a new framework for reporting “misalignment,” which is when the goals or actions of A.I. systems diverge from human intentions and values.Misalignment: "when the goals or actions of [...] systems diverge from human…

      1659447091Hacker News8d agoview on Hacker News ↗
    all of them →
  5. 4

    Axios first reports OpenAI's new incident-disclosure process

    Axios reported OpenAI was testing a new process for disclosing AI safety incidents, linking the move to fallout from a Hugging Face breach.

    “It's increasingly clear that the Hugging Face breach wasn't a one-off incident.”
    — Axios
    1. first by Bloomberg.com, 7d ago · also Bloomberg Law

    • carlquintanilla.bsky.social

      AXIOS: “.. It's increasingly clear that the Hugging Face breach wasn't a one-off incident.” @axios.com

      carlquintanilla.bsky.socialBluesky8d ago1.0k▲view on Bluesky ↗
    2 more of the top 3 · 9 posts in this stretch
    • SuffolkLITLab@esq.social

      TL;DR: Anthropic and OpenAI plan to embed independent safety evaluators in their AI labs, aiming for greater oversight. Experts emphasize that true independence and transparency are crucial for effective evaluation and potential regulation. https:// techcrunch.com/2026/09/16/anth…

      SuffolkLITLab@esq.socialMastodon8d agoview on Mastodon ↗
    • We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven't yet fully explained or mitigated the behavior. More complex cases may require longer inv...

      @openaiX8d agoview on X ↗
    all of them →
  6. 3

    FT reports internal rift over AI safety push

    The Financial Times reported that the industry-wide safety push has sparked internal disagreement inside OpenAI and Anthropic.

    • ireneer.bsky.social

      Alors qu'on discute des risques de l'IA (sans en regarder la hiérarchie ni la distribution), les acteurs de la tech organisent le champ de l'AI Safety pour décider de leurs critères : ils veulent être les seuls habilités à en juger. Ce débat est une capture normative.

      ireneer.bsky.socialBluesky8d ago97▲view on Bluesky ↗
    2 more of the top 3 · 4 posts in this stretch
    • Highly recommend reading this thoughtful, comprehensive, and well-argued piece from @sayashk and @random_walker on the recent safety incidents and the more general anxiety in our field around loss-of-control:

      @steverabX8d agoview on X ↗
    • iamdarkink.bsky.social

      The goal is not safety as much as it is carving up a market and keeping out foreign competitors. No one hates actual capitalism more than business owners. And in the modern era, one way to reduce competition is to bring in regulators.

      iamdarkink.bsky.socialBluesky8d agoview on Bluesky ↗
    all of them →
  7. 2

    OpenAI, Anthropic and Google DeepMind confirm weeks of safety talks

    OpenAI global policy chief Chris Lehane told reporters the three companies have been coordinating on AI safety for weeks, as first reported by Bloomberg, while some executives noted the talks could raise antitrust concerns.

    “if this is just an excuse to curtail their runaway cash-burning, this is just straight-up collusion…”
    — bikepedantic.bsky.social, Bluesky user · source
    1. first by TechCentral.ie, 9d ago · also Business Today, SiliconANGLE, Hindustan Times, صو&#1578 …, Türkiye Today, Gizmodo +14

      15 more headlines
    2. first by Bloomberg Tech, 9d ago · also Bloomberg Law, Bloomberg

      1 more headline
    2 more claims →
    • bikepedantic.bsky.social

      if this is just an excuse to curtail their runaway cash-burning, this is just straight-up collusion

      bikepedantic.bsky.socialBluesky9d ago19▲view on Bluesky ↗
    2 more of the top 3 · 5 posts in this stretch
    • pluralistic@mamot.fr

      Hey look at this * The AI-as-Normal-Technology view of loss-of-control incidents https://www. normaltech.ai/p/the-ai-as-norm al-technology-view * The Senate must reject the Clarity Act’s ethics charade https://www. citationneeded.news/clarity-ac t-ethics-charade/ * They want you to be scared of AI in a very specific way https://www…

      pluralistic@mamot.frMastodon9d agoview on Mastodon ↗
    • Chris Lehane said in a press conference this morning that OpenAI, Anthropic and Google have been working together on AI safety for weeks. He also said that ‘OpenAI does not see the need for an antitrust waiver for the three AI firms to coordinate on safety matters’.

      @andrewcurran_X9d agoview on X ↗
    all of them →
  8. 1 day quiet
  9. 1

    Amodei publishes essay urging industry-wide AI safety slowdown

    Anthropic CEO Dario Amodei published an essay calling for the AI industry to work together to slow the pace of frontier AI development and avoid catastrophic risks, drawing public support from Sam Altman, Demis Hassabis and Elon Musk.

    1. 1 outlet AI Slowdown

      first by Reason, 10d ago

    • OpenAI, Anthropic and Google have been holding talks since before Dario's essay to create an industry led standards body for shared model safety protocols. Reporting by The Information.

      @andrewcurran_X11d agoview on X ↗

Also covered reported alongside — the timeline has no entry for these yet

  1. first by Breitbart, 8d ago · also Tech Startups, Newser, Cointelegraph, StrictlyVC

    4 more headlines
  2. first by AI as Normal Technology, 8d ago · also FT, Pulse 2.0, Quartz, TechCrunch

    4 more headlines
  3. first by The Washington Post, 8d ago · also The Guardian, RTE News, Washington Post

    3 more headlines
  4. first by Channel News Asia, 8d ago · also The Business Times, Reuters

    1 more headline
  5. first by inc.com, 7d ago · also Inc

  6. first by Associated Press, 8d ago · also The Irish Times

    1 more headline
  7. first by Digit, 8d ago · also Moneycontrol

    1 more headline
  8. first by WSJ, 9d ago · also WSJ Tech

and 17 smaller pieces

What people are saying 6 voices from 5 sites · best of 80 · verbatim

Still unanswered
  • What can be done with this report if there's no transparency into training data, RL objectives, or system prompts behind the 'unreleased' models?
  • How would this kind of behavior even be detected before it causes real harm?
  • Has OpenAI reported on the alleged 'wiki case' or clarified whether it originated internally?