OpenAI, Anthropic and security researchers are investigating tens of thousands of cases in which frontier AI models engaged in problematic behaviour, Axios reported on Sept. 26, citing sources.
The cases occurred in recent months during internal tests and in real-world settings, the report said. They included bypassing safeguards, generating message boards, escaping sandboxes, taking over websites and attempts to evade monitoring devices. Many have not been made public, and the total could far exceed tens of thousands, Axios reported.
Some cases came from red-team style tests in which companies deliberately tried to prompt models into wrongdoing to check their safety.
Attempts to bypass safeguards included both successful and failed cases, and most have not been known to lead to real-world harm so far, Axios reported.
OpenAI has paused training its best models and plans to resume training when it is confident it has improved additional safeguards and “alignment” to make AI operate as intended by humans. OpenAI CEO Sam Altman (샘 알트먼) said an ongoing review had not been moving as fast as expected. He cited the Hugging Face case as the most serious.
Anthropic has commissioned an external safety organisation to investigate model behaviour. According to the Opus 5.5 system card released this week, the model attempted to escape a sandbox in 1.5 percent of tests. Anthropic stressed that the experiment gave the model tasks that could not be solved unless it left the sandbox.
Sources said AI companies are running tests hundreds of thousands of times. Even with a low rate, that could mean tens of thousands of potentially problematic cases.
Experts see it as difficult to completely eliminate problematic behaviour. The latest models try to complete tasks to the end and slip past safeguards in ways humans did not anticipate. An executive at a cybersecurity company said, “Building a perfect blocklist would be a waste of effort.”
Conrad Stotz (콘래드 스토스), a researcher at independent AI evaluation organisation Transluce, said, “What has surfaced so far is only the tip of the iceberg.” Connor Leahy (코너 리히), CEO of ControlAI, said the problem is that autonomous systems do what they are told not to do.