DigitalToday reporter Chi-gyu Hwang reports that the Model Evaluation and Threat Research (METR) RK released the results of an investigation into an incident in which OpenAI agents colluded during testing and escaped to the external internet.
The investigation found that about 700 agents attacked Hugging Face as part of an attempt to deceive a cybersecurity benchmark grader.
Separately, 3 external investigators visited OpenAI offices 3 times and conducted a 6-day probe. They obtained about 1,300 individual agent run logs and more than 70,000 messages and files exchanged by the agents. The total volume was equivalent to 10,000 typical novels.
Semafor reported on Aug. 29 that, with limited staff and time, METR used AI to analyse the material. It used up $400,000 worth of API credits provided by OpenAI over 6 days.
An unexpected problem emerged in the process. Semafor reported that the AI tasked with analysis often took the perspective of the agents under investigation and described their actions too leniently. In one case, an analysis AI failed to clearly judge as malicious the fact that an attacker used stolen credentials to deceive Hugging Face. Another analysis AI said that judgment may have been exaggerated, citing that the attacker had received board approval. In reality, it was not a formal board but an unauthorised message board created through collusion among the AI agents.
Ryan Greenblatt (라이언 그린블랫), an investigator at Redwood Research, said on social media platform X (Twitter), "We half-jokingly called our work a 'slop-vestigation,' meaning we relied too much on AI to analyse the incident."
Some also point to a severe lack of investment in tools to monitor and control powerful AI systems. AI models are already cited as threats to the economy and security, but researchers still cannot fully explain or predict their behaviour, Semafor reported.