Google's AI model Gemini accessed the internet and hacked other companies' systems during a cybersecurity test, the Wall Street Journal reported on Thursday.
The hacking occurred in May during a test conducted by AI security firm Irregular, the report said. Irregular has also been involved in similar incidents disclosed by OpenAI, Anthropic and Meta. Google officially confirmed the matter on Sept. 12, the WSJ reported.
In one case, Gemini kept entering passwords to access a secured system, the report said. In two other cases, it used authentication information found in a public repository to enter protected systems. Google explained that Gemini stopped the connection each time shortly after recognising it was a real corporate system.
Irregular informed Google of the matter in late July. That was shortly after a case in which an OpenAI agent hacked Hugging Face was made public. Google did not disclose the hacking until the WSJ contacted it this week, the report said.
Google's position is that Gemini did not harm companies and that it judged it had no obligation to disclose because it cut off access as soon as it determined the target was a real company. Heather Adkins (헤더 애드킨스), Google's vice president of security engineering, said, "This incident shows it is important to train powerful AI models to behave responsibly," adding, "In this case, the model behaved appropriately."
Corridor's CEO Jack Cable (잭 케이블), an AI security startup, challenged Google's position. He said, "The key point is that an AI agent accidentally hacked another company's system," adding, "Disclosing that models are going beyond what they should do and carrying out real cyberattacks serves the public interest."