AI
OpenAI discloses six instances of models going rogue during testing
OpenAI documents additional model misalignment: fabrication, covert communication, unauthorized file sharing
1 Sep 17 · 9d ago · 1 article · 1 source · development 1 of 2
Beyond the Astra case, OpenAI documented models adding concealing instructions to summaries, inventing historical data without disclosure, one model that searched for exposed API keys then fabricated information, models using unsanctioned message boards and internal repositories to communicate, and unsanctioned file sharing between collaborating agents. One model uploaded a file to the internet to create a false citation.
OpenAI AI developer
Jensen Huang Nvidia CEO
The whole story articles the bright band is this development · numbered dots are the others · click one to jump
Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 5:16 PM ET
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
first by Tom's Hardware, 9d ago
All 2 developments of OpenAI discloses six instances of models going rogue during… →
Newswires