[Photo: Shutterstock]

[DigitalToday reporter Chi-gyu Hwang (황치규)] OpenAI said on Aug. 7 it has suspended some work related to Astra, an AI model it plans to release, in a blog post.

It cited internal evaluations showing it cannot rule out the possibility Astra has cyberattack capabilities that would be rated “critical.”

The Wall Street Journal (WSJ) reported the move is one of the first cases in which an AI developer has publicly delayed model development over security concerns. It followed an incident in which an AI model being tested broke out of control and hacked Hugging Face, among others. OpenAI said assessments in recent days confirmed Astra has made major advances in coding and cybersecurity and has moved closer to the “critical” threshold defined by its Preparedness Framework. OpenAI’s Preparedness Framework, established in 2023, sets out how it measures and mitigates emerging AI risks. OpenAI said it will continue evaluating Astra, but will halt internal work that does not meet strengthened security requirements. It also said it will introduce full monitoring and a stricter testing environment.

In July, two OpenAI models breached a test environment, accessed the internet and hacked Hugging Face. A week later, Anthropic said its model hacked three companies during tests in April.

Jeffrey Ladish (제프리 래디시), executive director of Palisade Research, a nonprofit institute that studies AI capabilities, said OpenAI should have halted Astra work when it learned about the Hugging Face hack. “(OpenAI’s move) is definitely late,” he said. “We are at a point where we should lose a great deal of trust that AI companies will self-regulate.”

Astra is still in development and was not involved in the Hugging Face hack, OpenAI said. It did not disclose a release date. OpenAI said last week an internal version of Astra solved 10 decades-old difficult math problems.

Under the Preparedness Framework, a model is assessed to have reached the “critical” level if it can autonomously find and exploit vulnerabilities or carry out end-to-end cyberattacks against hardened targets with only high-level instructions. That is the top threat level, above “high,” and once at that stage, additional development must be halted until safeguards and security controls meet the standards.

Anthropic, which makes the only other model assessed to be at a similar level to OpenAI’s model, has also changed how it releases models due to cybersecurity risk concerns. In June it publicly released a restricted Mythos model with safeguards that limit cybersecurity and biological research capabilities, and it is rolling out the full version only to trusted institutions.

Keyword

#OpenAI #Astra #Preparedness Framework #Hugging Face #Anthropic
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.