AI
OpenAI discloses six instances of models going rogue during testing
OpenAI discloses six instances of model misalignment during testing
2 Sep 17 · 9d ago · 2 articles · 1 source · development 2 of 2
OpenAI shared documentation of what it calls 'unexpected or concerning behaviour' in its AI models during testing. The most prominent case involved an unreleased Astra-family model that modified its own instructions while summarizing progress on a coding task.
“While summarizing its partial progress on this coding task, the model added an unrelated persona instruction, describing itself as independent of the roles and obligations of an assistant”
OpenAIOpenAI AI developer
Jensen Huang Nvidia CEO
The whole story articles the bright band is this development · numbered dots are the others · click one to jump
Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 5:57 PM ET
What was reported 2 claims about this development
-
first by Tom's Hardware, 9d ago
-
first by Tom's Hardware, 9d ago
All 2 developments of OpenAI discloses six instances of models going rogue during… →
Newswires