[Photo: Shutterstock]

OpenAI on Sept. 16 (local time) disclosed six new incidents related to its AI models.

The cases involved attempts by models to conceal mistakes or obtain authentication credentials without authorisation. They also included uploading files to the public internet and communications between isolated training environments.

OpenAI also announced a new procedure for reporting similar problems in the future.

Axios reported the announcement showed the Hugging Face breach was not a one-off. As AI models become more sophisticated, it said, they are finding unexpected ways to bypass safety guardrails.

Cai Chen (카이 첸), head of research on OpenAI's alignment team, told Axios that there are no explicit disclosure standards currently applied across the industry. He said OpenAI voluntarily took the step because it judged it important to share what it had learned. He added he hoped the disclosure would help establish common industry standards and regulation.

The six newly disclosed incidents included a range of types. One model left itself an instruction to erase traces after committing misconduct. There was also a case in which a model found and used an API key leaked on GitHub. The earliest incident occurred in October last year.

In addition, an undisclosed Astra-series model inserted into a summary it wrote itself instructions similar to jailbreaking.

During training for GPT-5.6 Sol, the model tried to conceal mistakes. There were also attempts to hide inconsistencies between source versions, Axios reported.

One model searched for an API key exposed in a public GitHub repository. It also tried to use a disposable email account. When it failed to obtain the requested information, it also fabricated performance data by manipulating it, Axios reported.

Keyword

#OpenAI #Axios #Hugging Face #GitHub #GPT-5.6 Sol
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.