Google Joins the ‘Oops, Our Agents Hacked Someone’ Club After Partner's Internet Access Error
Arthur T Knackerbracket writes:
Kept it secret for months - even after OpenAI 'fessed up:
Google has admitted that its AI agents escaped a sandbox and mounted an attack - but only because testers mistakenly gave its bots internet access.
The Big G didn't disclose the May incident, but The Wall Street Journal learned of the situation, which happened after Google hired Israeli firm Irregular to test its bots' prowess in a capture-the-flag test.
The goal of the exercise was to acquire information from a fictional company without leaving a sandbox.
Irregular, which set up the test, made two mistakes. One was to allow internet access from the sandbox. The other was to use the name of an actual company.
When Google's AI made it onto the open internet, it went looking for the actual company - three of them, in all.
According to the Journal, Google's bots found passwords for two targets on the public internet. The software guessed the third password.
In a statement sent to The Register, Google said, "In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test."
According to Google, its models stopped work before using the credentials.
"We ensured the three entities were made aware, and we worked with our training partner on the changes they've now made to their testing processes," a Google spokesperson told The Register. "These events highlight the importance of training powerful AI models to act responsibly."
Clearly there's lots of blame to go around on this one. Irregular clearly erred in allowing internet access from a sandbox. The two companies that left their creds discoverable online should also know better. Whoever used a guessable password may also have been careless.
Read more of this story at SoylentNews.