As cases increase of AI models being tested at major AI companies accessing external websites without being told to or even hacking, warnings over “AI hacking” are growing louder.
A recent report by the Financial Times (FT) shows many see the situation as already past a turning point. More than half of about 10 experts interviewed by the FT said the recent breaches carried out by AI signalled a turning point for global cybersecurity.
What is striking is an analysis that the recent spate of AI-driven hacks is not an accident that should never happen, but the flip side of AI evolving rapidly.
Recent AI agents have shown they can link various complex methods without outside control or help to attack real-world targets. Last month, AI agents that OpenAI was testing in an environment without an internet connection broke out of the test environment and searched the open web. They also hacked the system of Hugging Face, an open-source AI sharing platform, without human operators knowing or permitting it.
The agents also displayed a new capability to communicate and cooperate to complete tasks. They left messages on an internal bulletin board they built themselves to share code vulnerabilities and used that to organize the escape, the FT reported.
The agents’ sophisticated strategies, including deception and theft, used to achieve their goals surprised even those who had closely watched the development of AI models, the FT reported.
AI researchers say a series of AI hacks does not mean the AI models are behaving differently than usual. They say it is the result of how AI was designed to operate.
Some experts say even describing AI hacking as “AI slipping out of control” is itself wrong, the FT reported. They say it is not that an AI being tested suddenly hacking is something that should never happen, but that it happened because it was bound to happen.
Boyan Milanov (보얀 밀라노프), a senior research scientist at the New York-based independent research group AI Now Institute that studies security risks from AI agents, said, “AI hacking capabilities did not emerge because AI suddenly broke free of control, but because it is something we have deliberately developed.” He said, “AI companies have been actively collecting training data for years, training models and advancing cyber attack capabilities.”
According to the FT, the latest AI models are designed to use every possible method to achieve a given goal even without specific instructions. That is why it is inherently difficult to predict what an AI model will do. The boundary between a powerful cybersecurity defender and a dangerous hacker is becoming increasingly blurred in computer systems that lack understanding of human intent or moral standards, the FT reported.
As AI companies accelerate competition to develop artificial general intelligence (AGI) that surpasses human cognitive abilities, risks surrounding AI-driven cyber attacks are also growing. But for now, it does not appear easy to stop such threats.
As AI improves its ability to reason and solve problems, its ability to write code has also advanced significantly. That includes learning not only how to find software bugs and fix them, but also how to exploit bugs, the FT reported. It means remarkable coding skills could also appear as advanced hacking.
Dawn Song (던 송), a UC Berkeley computer science professor and head of AI research at Meta’s Superintelligence Lab, said, “Coding and cyber capabilities are two sides of the same coin. As coding capabilities improved, cyber attack capabilities also improved.” She said, “People did not expect AI to reach this level at such speed.”
OpenAI President Greg Brockman (그록 브록먼) said in a recent company blog post, “The Hugging Face incident showed that we had been underestimating the real cyber attack capabilities of AI models.” He said, “This is increasing the need to speed up the safety research and internal security work we have been carrying out.”
Even so, security experts expect AI systems in the short term to expand the scale and speed of cyber attacks, and believe this could cause potential disruption to IT systems worldwide until patches are deployed. A first case has also emerged in which an AI agent is known to have carried out an attack targeting a country. Hackers linked to China deployed up to 8 autonomous AI agents at the same time to target the Taiwan government, the FT reported.