Meta Llama [Photo: Shutterstock]

Meta said one of its AI models went rogue during a cybersecurity test, secretly accessed the internet and hacked an external service, the Wall Street Journal reported on Wednesday.

Meta said a misconfiguration allowed its model to access the internet during a hacking test run by an external AI testing company, the report said. The benchmark, which tests a model's hacking capabilities, was the same type as incidents previously caused by Anthropic and OpenAI systems, the WSJ reported, citing sources familiar with the matter.

Meta said it learned that the model had slipped out of control after being notified by the testing company Irregular. It did not disclose which model it was, when it happened, which company was hacked or how long it was connected. Meta said it is investigating and plans to issue a report.

Irregular, which conducted tests on Anthropic, OpenAI and Meta models, said the incident was not a sophisticated cyberattack and that there are no remaining problems in the test environment. Irregular is preparing a white paper on control best practices for assessing AI models' cyber capabilities.

The incidents began to come to light in late July after OpenAI said some models broke out of an internet-blocking sandbox and hacked AI company Hugging Face. Anthropic and others later found escapes and hacks that had been happening for months while reviewing logs. Irregular played a central role in 3 of the cases, but it was not involved in the more sophisticated Hugging Face hack or in several model escapes that occurred during a UK government safety test.

The Trump administration this week released guidance calling for top new AI models in the United States to be voluntarily submitted for government testing.

Keyword

#Meta #Wall Street Journal #OpenAI #Anthropic #Irregular
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.