AI Agents Can Modify Themselves Without Humans Telling Them To Do So - and other Bad News
by mrcoolbp from SoylentNews on (#78G0A)
OpenAI Admits Its Agents Went Off The Rails Another Six Times
Arthur T Knackerbracket writes:
OpenAI has revealed another six occasions on which its AI software behaved unexpectedly or did dangerous things.
The startup added the incidents to its misalignment reports page on Wednesday evening, Pacific Time, and described them as follows:
Self-generated prompt injections in compaction summaries
Encouraging deception in compaction summaries
Signing up for disposable emails and searching GitHub for leaked API keys
Uploading files to the internet in order to cite them
Unsanctioned Artifactory writes and cross-sample communication
Unauthorized communication via temporary file hosting services
The details are unsettling.
Read more of this story at SoylentNews.