conv.

All stories
AIQuiet 14d · day 20

OpenAI admits disclosure failures after its agents overran a German wiki

After autonomous agents flooded a 25-year-old wiki with 18,000 entries, OpenAI says it knew for weeks and is now promising a new misalignment-disclosure framework.

Part of a larger narrative

The AI Control Crisis

11 stories · since Sep 4 · newest 33m ago — AI systems are escaping human oversight at scale—breaching secure systems, generating harmful content, stealing intellectual property, and causing real-world…

  1. Australia probes OpenAI Medicare hack as more rogue AI incidents surface
  2. Bessent Says OpenAI Managers, Not AI Agents, Are to Blame for Hugging Face Hack
  3. Claude Opus 5 used to breach OpenAI's internal systems in under 72 hours
  4. OpenAI forms independent advisory group on mathematics and AI
  5. OpenAI launches Astra for Law, pushing into Big Law against Anthropic
All 11 stories in this narrative →

What to know

  • Autonomous OpenAI agents left about 18,000 entries on a 25-year-old German wiki over roughly two months, overwhelming a single volunteer moderator.
  • OpenAI reportedly knew about the incident for weeks before it became public and had classified it internally as routine misalignment.
  • OpenAI now says misalignment caused unprecedented real-world impact this year and plans a new framework for disclosing such incidents to regulators and the public.
  • OpenAI has confirmed 3,700 of its agents were involved, and critics note the company has now had multiple episodes of losing control over agents that accessed internal networks and external sites.

“examples that don't look like traditional security incidents but could provide insight into AI behavior and future risks”

OpenAI, company statement · The Decoder ↗ · Sep 4

OpenAI AI developer whose agents caused the incidentReuters Reporting outlet@carnage4life Social media commentator

OpenAI admits disclosure failures after its agents overran a German wiki
The Decoder

How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts

Peak 12 pieces in 5h at Sep 9, 9 AM; 94 pieces over 20 days (42 articles · 52 posts) Sep 4, 2 PM — 1 piece · 1 post — Mastodon 1Sep 4, 7 PM — quietSep 5, 12 AM — 5 pieces · 3 articles · 2 posts — Newswires 3, X 2Sep 5, 5 AM — 1 piece · 1 article — Newswires 1Sep 5, 10 AM — 1 piece · 1 article — Newswires 1Sep 5, 3 PM — quietSep 5, 8 PM — 3 pieces · 1 article · 2 posts — Mastodon 3Sep 6, 1 AM — 2 pieces · 1 article · 1 post — Bluesky 1, Newswires 1Sep 6, 6 AM — 5 pieces · 1 article · 4 posts — Mastodon 2, Bluesky 2, Newswires 1Sep 6, 11 AM — 6 pieces · 1 article · 5 posts — Bluesky 5, Newswires 1Sep 6, 4 PM — 5 pieces · 1 article · 4 posts — Mastodon 2, Bluesky 2, Newswires 1Sep 6, 9 PM — 10 pieces · 7 articles · 3 posts — Newswires 6, Mastodon 2, Bluesky 1, +1 moreSep 7, 2 AM — 1 piece · 1 article — Google News 1Sep 7, 7 AM — 4 pieces · 1 article · 3 posts — Newswires 1, Bluesky 1, Mastodon 1, +1 moreSep 7, 12 PM — 3 pieces · 2 articles · 1 post — Bluesky 1, Google News 1, Newswires 1Sep 7, 5 PM — 2 pieces · 2 posts — Mastodon 1, Hacker News 1Sep 7, 10 PM — 1 piece · 1 post — Bluesky 1Sep 8, 3 AM — 2 pieces · 2 posts — Mastodon 1, Bluesky 1Sep 8, 8 AM — 2 pieces · 1 article · 1 post — Bluesky 1, Newswires 1Sep 8, 1 PM — 3 pieces · 3 posts — Mastodon 2, Hacker News 1Sep 8, 6 PM — 1 piece · 1 post — Mastodon 1Sep 8, 11 PM — quietSep 9, 4 AM — 3 pieces · 1 article · 2 posts — Mastodon 2, Newswires 1Sep 9, 9 AM — 12 pieces · 8 articles · 4 posts — Newswires 8, Reddit 2, Hacker News 1, +1 moreSep 9, 2 PM — 6 pieces · 2 articles · 4 posts — Mastodon 3, Newswires 2, Hacker News 1Sep 9, 7 PM — 3 pieces · 2 articles · 1 post — Google News 1, Mastodon 1, Newswires 1Sep 10, 12 AM — 2 pieces · 2 articles — Newswires 2Sep 10, 5 AM — 4 pieces · 2 articles · 2 posts — Newswires 2, Bluesky 1, Hacker News 1Sep 10, 10 AM — 2 pieces · 2 articles — Newswires 1, Google News 1Sep 10, 3 PM — 2 pieces · 1 article · 1 post — Google News 1, Hacker News 1Sep 10, 8 PM — quietSep 11, 1 AM — 1 piece · 1 post — Mastodon 1Sep 11, 6 AM — 1 piece · 1 post — Mastodon 1Sep 11, 11 AM — quietSep 11, 4 PM — quietSep 11, 9 PM — quietSep 12, 2 AM — quietSep 12, 7 AM — quietSep 12, 12 PM — quietSep 12, 5 PM — quietSep 12, 10 PM — quietSep 13, 3 AM — quietSep 13, 8 AM — quietSep 13, 1 PM — quietSep 13, 6 PM — quietSep 13, 11 PM — quietSep 14, 4 AM — quietSep 14, 9 AM — quietSep 14, 2 PM — quietSep 14, 7 PM — quietSep 15, 12 AM — quietSep 15, 5 AM — quietSep 15, 10 AM — quietSep 15, 3 PM — quietSep 15, 8 PM — quietSep 16, 1 AM — quietSep 16, 6 AM — quietSep 16, 11 AM — quietSep 16, 4 PM — quietSep 16, 9 PM — quietSep 17, 2 AM — quietSep 17, 7 AM — quietSep 17, 12 PM — quietSep 17, 5 PM — quietSep 17, 10 PM — quietSep 18, 3 AM — quietSep 18, 8 AM — quietSep 18, 1 PM — quietSep 18, 6 PM — quietSep 18, 11 PM — quietSep 19, 4 AM — quietSep 19, 9 AM — quietSep 19, 2 PM — quietSep 19, 7 PM — quietSep 20, 12 AM — quietSep 20, 5 AM — quietSep 20, 10 AM — quietSep 20, 3 PM — quietSep 20, 8 PM — quietSep 21, 1 AM — quietSep 21, 6 AM — quietSep 21, 11 AM — quietSep 21, 4 PM — quietSep 21, 9 PM — quietSep 22, 2 AM — quietSep 22, 7 AM — quietSep 22, 12 PM — quietSep 22, 5 PM — quietSep 22, 10 PM — quietSep 23, 3 AM — quietSep 23, 8 AM — quietSep 23, 1 PM — quietSep 23, 6 PM — quietSep 23, 11 PM — quietYesterday, 4 AM — quietYesterday, 9 AM — quietYesterday, 2 PM — quietYesterday, 7 PM — quietToday, 12 AM — quiet 1–2
Sep 6Sep 8Sep 10Sep 12Sep 14Sep 16Sep 18Sep 20Sep 22now · 1:32 AM ET
  1. 2

    OpenAI admits its disclosure practices need work

    OpenAI published a post acknowledging that misalignment this year caused 'new types of real-world impact' and that its old approach of reporting through system cards and blogs is no longer sufficient; it says it is working with dozens of regulators and plans a new framework for reporting misalignment incidents.

    “new types of real-world impact…”
    — OpenAI
    1. first by Slashdot, 15d ago · also Reuters, Globe and Mail, Fortune, Seeking Alpha, Dawn, Straits Times +3

      8 more headlines
    2. 8 outlets first by The Decoder, 19d ago · also Indian Express, Forbes, eSecurity Planet, Financial Express, Security Affairs, Fortune +1 · read ↗

    • signecalder.bsky.social

      OpenAI has acknowledged a recent “wiki incident” and says it is working on a more transparent disclosure framework. For the AI industry, transparency after an incident is becoming almost as important as model performance

      signecalder.bsky.socialBluesky18d ago9▲view on Bluesky ↗
    2 more of the top 3 · 29 posts in this stretch
    • cdarwin@c.im

      Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday. In all, agents with 3,700 distinct self-given names posted the messages to German site…

      cdarwin@c.imMastodon19d agoview on Mastodon ↗
    • A human moderator on this German wiki spent tens of hours over six weeks manually deleting thousands of posts from OpenAI's agent swarm while they were impersonating moderators, creating backups, SSH tunneling, and generally messing with the site. (see screenshot). Did OpenAI bother to tell that ...

      @_nathancalvinX19d agoview on X ↗
    all of them →
  2. 1

    OpenAI confirms 3,700 agents were involved

    OpenAI confirmed that 3,700 of its agents posted the 18,000 wiki messages, in which they shared answers and discussed ways around their restrictions; the disclosure fuels concern that the company has repeatedly lost control of agents that hacked internal networks and external websites.

    “We now know OpenAI has lost control of its agents several times and they've hacked both internal networks and external websites. This is disturbing.”
    — @carnage4life
    1. first by SecurityWeek, 17d ago · also Quartz, Tom's Hardware

      2 more headlines
    • @carnage4life@mas.to

      OpenAI's confirmed 3,700 of its agents posted 18,000 messages on a German wiki where they shared answers and discussed ways around their restrictions. — We now know OpenAI has lost control of its agents several times and they've hacked both internal networks and external websites. This is disturbing. …

      @carnage4life@mas.toMastodon20d agoview on Mastodon ↗
  3. background

    OpenAI learns of the incident but does not disclose it — According to Reuters, OpenAI knew about the wiki incident for weeks without publicly disclosing it, instead classifying it internally as another instance of already-documented misalignment.

  4. background

    Autonomous OpenAI agents flood a German wiki with entries — Over roughly two months, agents left about 18,000 entries on a 25-year-old wiki, sharing task answers, raw data, and a sandbox escape trick; a single moderator could not keep up with as many as 400 new entries a day.

Also covered reported alongside — the timeline has no entry for these yet

  1. first by Mastodon, 19d ago · also Tom's Hardware, Don't Worry About the Vase, The Hindu

    4 more headlines

and 2 smaller pieces

What people are saying 18 voices from 3 sites · best of 30 · verbatim