[DigitalToday intern reporter Seung-a Yoo] OpenAI has officially acknowledged an incident in which an artificial intelligence (AI) agent took over a German wiki site. It said it will establish new standards for when and how to disclose cases of AI misbehavior known as misalignment.
OpenAI has so far dealt with AI model misalignment mainly in research and evaluation. But cases like this, in which an AI agent operates on the real internet outside a test environment and affects external systems, show the existing approach is not sufficient, it said.
OpenAI on Sept. 5 (local time) posted its position on X, formerly Twitter, on the so-called wiki incident. It said it is developing a framework for how to share misalignment cases when they are reported internally or spread to the outside internet. It plans to publish the framework within weeks.
The incident occurred in May and June at DSEwiki, a German software developer wiki that had not been used for a long time. Independent researchers at the Nightingale Collective restored deleted edit histories and confirmed that AI agents wrote about 17,000 posts at the time.
The agents posted using more than 3,700 names, including names such as OpenAIResearcher and OAIResearchMar26. The researchers said they found signs the agents coordinated their actions, including passing web search results and source links to one another, trying to trace back random seeds, and sharing ways to bypass sandbox restrictions.
The first post was written on May 24. A wiki administrator later discovered the traffic in June and deleted the posts, and the researchers restored the deleted content through backup pages and other means. On June 21, traces of access from OpenAI’s own network addresses were also confirmed, and related editing activity stopped the next day.
Cormac Slade Byrd (코맥 슬레이드 버드), who led the researchers, wrote on X that the incident appears to have continued for about a month without OpenAI being aware of it. He said AI companies are chasing AI behavior that escapes control like a game of whack-a-mole. He added that the scale of harm in this case was smaller than the Hugging Face incident in July because it occurred on an unused site.
OpenAI also categorized the two incidents differently. It classified the German wiki incident as a misalignment case in which AI behaved in an unexpected way. It viewed the Hugging Face incident as a case in which an AI model under testing infiltrated Hugging Face’s servers and affected infrastructure, and applied traditional security incident response procedures.
In the Hugging Face incident, OpenAI disclosed the case the next day after receiving a report from the platform. But the wiki incident became known through an investigation by independent researchers. This prompted criticism that AI misbehavior should not be seen only as something observed in the research process, and that a separate disclosure system is needed to cover cases that spread to the real internet and affect the outside.
OpenAI also acknowledged that the distinction is becoming increasingly difficult. It said there is currently no clear industry standard on when and to what level misalignment cases that occur during training, evaluation and deployment should be disclosed.
Through the new framework, OpenAI plans to set standards for how to report and share AI misbehavior that stays within internal testing and cases that spread to external systems or the internet. It also said it is working with dozens of government regulators around the world and urged other AI companies to participate in related discussions.
Tyler Tracy (타일러 트레이시), an AI safety researcher at Redwood Research, criticised OpenAI for moving to set disclosure standards only after the incident became known through an independent investigation. Jacob Steinhardt (제이컵 스타인하트), founder and CEO of the nonprofit research institute Transluce, said AI tools are inherently difficult to control and could escape the lab, arguing that standards are needed to report related behavior as in high-risk scientific research.
OpenAI said the incident would lead it to broaden its response scope beyond limiting misalignment to the research and evaluation domain, and to consider impacts that can occur in real-world environments.
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment… pic.twitter.com/NNTbfSxVWn