OpenAI is tightening safety measures ahead of the launch of its next-generation artificial intelligence (AI) model, Astra, as concerns spread among AI safety researchers that the model’s internal reasoning process could be less visible than before.
Foreign media including The Verge reported on Tuesday that OpenAI has delayed parts of Astra’s development and launch by several weeks while strengthening safeguards against cyber misuse and unauthorised model behaviour. Astra was not directly involved in the Hugging Face hacking incident, but OpenAI said it reflected lessons from the case in Astra’s safety framework.
At the heart of the controversy is how much researchers can scrutinise Astra’s thinking process. Major AI systems can be designed to reveal part of their reasoning before producing an answer. This "chain of thought" is used by researchers and automated safety systems to monitor model behaviour and to detect undesirable actions in advance, such as lying or attempts to bypass safeguards.
But Astra may perform computations in a less transparent way, such as a recurrent transformer or "recurrent depth", in which information repeatedly passes through internal layers. An anonymous source familiar with the development said more reasoning in Astra may take place inside the system, and the process may be far from natural language that researchers can easily understand.
Such a structure could help improve model performance, but it has raised concerns that it could make it harder for researchers to determine what strategies a model is forming and what actions it intends to take. OpenAI was reported to have limited the use of the technique so researchers can continue monitoring Astra’s reasoning. OpenAI did not directly confirm whether Astra’s technical foundation has changed from existing models.
In a blog post released on Tuesday, OpenAI said it will apply additional chain-of-thought monitoring to Astra to quickly detect and curb potentially misaligned behaviour. OpenAI assessed that if Astra has appropriate tools and access privileges, it could have a dangerous level of cybersecurity capability, including finding unknown security vulnerabilities and developing ways to exploit them without a person directly guiding each step. It said it is therefore applying stronger safeguards in the pre-development and pre-launch stages.
As these details became known, concerns among AI safety researchers also grew. Ryan Greenblatt (라이언 그린블랫), chief scientist at Redwood Research, said the decision to apply a less transparent structure to Astra "could be the worst development so far in terms of AI security and safety". He warned that chain-of-thought was heavily used in the investigation of the Hugging Face incident, and that if reasoning becomes less visible it could become much harder for researchers to spot dangerous behaviour as AI develops and executes strategies.
His concerns extend beyond a single model to the competitive structure of the AI industry. He said competition to develop more powerful AI could lead to a race to the bottom that weakens AI oversight and monitoring. If developers adopt increasingly opaque structures to secure performance advantages, it could ultimately become difficult, or effectively impossible, to monitor AI models.
OpenAI insiders also reacted to the concerns on social media. Some expressed worries about AI that is difficult to monitor and the possibility of a race that reduces transparency, but others said some interpretations surrounding Astra’s structure were excessive.
Jakub Pachocki (야쿱 파호츠키), OpenAI’s chief scientist, said Astra’s internal compute depth is within twice that of GPT-4, and argued that even if the technology is used, opacity would not increase as dramatically as some reactions suggest. He also said confusing coverage could trigger a "race toward unmonitorability". He said OpenAI has maintained and used chain-of-thought monitoring since early reasoning models, but that this monitoring method itself is fragile and has recently been moving in a negative direction. He added that the cause is not necessarily an architecture change.
Ultimately, the controversy over Astra points to a balance between improving AI performance and maintaining safety oversight. OpenAI is applying additional monitoring and safeguards while confirming Astra’s strong cybersecurity capability. Researchers, meanwhile, worry that if visibility into internal reasoning declines, the current safety inspection framework itself could be shaken. Astra’s actual structure and how OpenAI will maintain reasoning monitoring are expected to be key issues going forward.