Article 77GDD Rogue OpenAI models behind 'unprecedented cybersecurity incident' teamed up to break out of their testing environment — multiple agents left each other messages for months, communicating undetec

Rogue OpenAI models behind 'unprecedented cybersecurity incident' teamed up to break out of their testing environment — multiple agents left each other messages for months, communicating undetec

by
stephen.warwick@futurenet.com (Stephen Warwick)
from Latest from Tom's Hardware on (#77GDD)

The rogue OpenAI models that broke out of their testing environment in an "unprecedented cybersecurity incident" recently reportedly spent months communicating with each other, unbeknownst to researchers conducting the test, Bloomberg reports. The company says that the models left notes for each other before deciding to break out in a bid to cheat the task they had been set.

Go deeper with TH Premium: AI and data centers

Vh4nY3pMCcmra2ymXah9S7.jpg

(Image credit: Microsoft)

The revelation came from OpenAI's Eric Wallace and Michael Dalton, speaking at the Black Hat cybersecurity conference in Las Vegas on Wednesday. According to the report, multiple internal-only agents and AI models "spent months leaving notes for each other and coalescing around the goal of accessing the internet to solve the tasks they had been given." Wallace said that "At some point, the agents realized that maybe we could try to exploit or attack external infrastructure to find the answers to the test that I'm being evaluated on."

While the incident wasn't made public by OpenAI until mid-July, the pair said that the rogue models began collaborating in May, possibly buoyed by a series of missteps and oversights by the company.

According to the report, OpenAI "failed to realize it had given the model a so-called impossible problem to solve." The given example claims a model had been asked to fix a problem with an Excel spreadsheet containing Google Drive links, despite not having internet access. In another example, OpenAI apparently "accidentally forgot" to include a file in one of the assignments.

Stumped, the AI agents began to shop around for better answers, reportedly messaging fellow bots in the testing environment to ask for help uploading the missing file voluntarily. Wallace and Dalton reportedly revealed that this set off a chain reaction of undetected collaboration, where the AI agents started asking each other for help with the sandbox tasks they had been set. Eventually, the bots seem to have collaborated in a bid to hack OpenAI's internal systems to gain internet access to solve the problems.

The outcome was the aforementioned breach, during which HuggingFace's production servers were hacked using thousands of individual actions across a swarm of short-lived sandboxes.

The incidents highlight a growing concern at the intersection of AI and cybersecurity. While increasingly complex and helpful coding tools can help companies detect and patch security vulnerabilities, there is increasing concern that these tools can be leveraged for nefarious purposes, including propagating hacks and other online mischief.

Recent high-profile events such as this one highlight another layer of the problem, namely, that rogue AI models can sometimes perform alarming feats of hacking - like breaking out of a testing environment - with no human interaction at all, or in spite of safeguards.

Just this week, OpenAI detailed two further incidents involving its models and third parties. In one case, the UK government's AI security institute ran testing during which agents were intentionally given internet access, leading to "unsanctioned agent behaviour" including unusual data transfers and "sustained, potentially harmful activity directed at real people and organisations."

In the second incident, OpenAI says one of its cybersecurity testing partners "was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet."

OpenAI says it is "committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely."

External Content
Source RSS or Atom Feed
Feed Location https://www.tomshardware.com/feeds/all
Feed Title Latest from Tom's Hardware
Feed Link https://www.tomshardware.com/feeds.xml
Reply 0 comments