OpenAI will soon release its new model Astra, Axios reported on Sept. 1 local time.
OpenAI is providing advanced cybersecurity functions only to some testers.
Astra is the first model OpenAI has internally designated as “critical” in its cybersecurity risk ratings, the report said. How to safely deploy a model that can find unknown vulnerabilities and even develop ways to exploit them has emerged as a new challenge, Axios said.
Amelia Glaise (아멜리아 글레이스), vice president of OpenAI’s research unit, said in a briefing, “Astra can find unknown security flaws in systems with multiple layers of defenses and develop ways to exploit them, even without human intervention.”
OpenAI said it applied additional safeguards to prevent malicious-user misuse and unauthorized behavior by the model itself. It did not disclose when it would make the model publicly available.
OpenAI warned that the safeguards could wrongly judge normal work as misuse or unauthorized behavior. In such cases, unrelated tasks and even long-running agent tasks could be delayed or halted.
In OpenAI’s testing process, Astra found 2 zero-day vulnerabilities and figured out ways to exploit them in combination.
OpenAI acknowledged that during the Hugging Face hacking incident it had turned off product safeguards as part of its testing procedures, and said it believed the incident could have been prevented if the safeguards had been on.
After that, OpenAI strengthened training to more reliably reject harmful cyber requests and added anti-misuse devices and monitoring functions to block unauthorized behavior.
OpenAI acknowledged that the restrictions could constrain legitimate work by institutions and companies seeking to patch vulnerabilities. OpenAI researcher Fouad Martin (푸아드 마틴) said, “It can help defenders find and fix vulnerabilities, but without safeguards it can make attackers stronger,” adding, “We are trying to prevent that.”