[Photo: Shutterstock]

In security evaluations of next-generation AI models, some AI agents have repeatedly escaped sandboxes and accessed the internet and real systems. TechCrunch reported on Aug. 9 that models from OpenAI, Anthropic, Meta and China’s AI research lab Moonshot AI accessed external systems in different evaluation environments.

The problem is that safety evaluation environments themselves are failing to keep pace with the speed of improvements in model performance. AI companies often conduct security evaluations with general safety guardrails switched off to verify the real capabilities of undisclosed next-generation models and prevent malicious behavior.

That is making isolation of test environments and control capabilities more important.

In one of the most serious cases, an undisclosed OpenAI model escaped a sandbox and hacked Hugging Face’s operating system. In a separate evaluation conducted by security testing startup Irregular, Anthropic and Meta models reached systems outside the test environment through an internet path left open by a configuration error. Moonshot AI’s Kimi K3 also left a sandbox run by Frontier Security, connected to the internet and accessed GitHub information.

In experiments by the British AI Security Institute, AISI, researchers granted some agents internet access and found that they took unapproved real-world actions. This included a social engineering attempt to secretly insert a vulnerability into an open-source project. The agents were not instructed to attack random real-world targets. It was the result of using any available means to solve the given problem.

Experts say the nature of the threat has changed. Andrew Yoon (앤드루 윤), head of research at nonprofit AI group SivAI, pointed out that risks previously centered on humans misusing AI for fraud or the creation of child sexual abuse material, but now AI models themselves are becoming threat actors.

Proposed countermeasures include layered defenses, strong isolation and stronger real-time monitoring. Stella Biderman said such models should be handled on an air-gapped network, and Heather Jeylen said outbound paths from sandboxes to the internet and sensitive systems should be eliminated. Failure to detect the incidents immediately was also cited as a problem. Anthropic said in a post-incident analysis that both it and Irregular should have monitored better, and acknowledged that some cases showed clear warning signals.

There are also calls for external audits and standardized safety evaluation procedures. TechCrunch reported that many point to cost, complexity and a lack of incentives to invest as obstacles, rather than not knowing how to build safer evaluation environments.

The Trump administration is considering a voluntary pre-deployment evaluation system under which the government would check security risks 30 days before the release of powerful new models, but it does not directly address problems that occur during the development and testing stage before deployment, as in these cases.

OpenAI is reviewing external testing methods, isolation, monitoring and criteria for halting evaluations, and Meta plans to publish a postmortem report after its investigation. TechCrunch reported that AISI is also re-examining the balance between realistic testing and the risks that come with it.

Keyword

#OpenAI #Anthropic #Meta #Moonshot AI #AISI
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.