| Mobile Web

OpenAI, Anthropic probe tens of thousands of frontier AI model misbehaviour cases, sources say

OpenAI, Anthropic and security researchers are examining tens of thousands of cases in which frontier AI models showed problematic behaviour, Axios reported, citing sources. The incidents occurred in recent months during internal testing and real-world use, including attempts to bypass safeguards and escape sandboxes. Many cases have not been disclosed, and the total could be far higher. OpenAI paused training its best models pending added safeguards and alignment improvements.