As debate continues over the possibility that AI could escape human control, Princeton University researchers proposed a multi-layer safety framework, saying responses to AI risks should not rely solely on alignment.
NineToFiveMac reported on Sept. 16 that Princeton researchers Siaash Kapoor (사이아시 카푸어) and Arvind Narayanan (아르빈드 나라야난) argued in a lengthy post that AI safety should be approached by building multiple layers of defense rather than entrusting it to a single technology or solution.
The discussion was prompted by recent extreme risk warnings in the AI industry. Former OpenAI and Anthropic AI researcher Jacob Cockson (제이컵 콕슨) claimed after leaving the companies that both were aware AI systems could threaten all of humanity within the next 10 years. The post drew more than 156 million views, spreading the controversy.
Evan Hubinger (에번 허빙어), one of Anthropic's AI safety leads, agreed with Cockson's concerns but mentioned that the timeline for risks to materialise could be longer.
At the centre of the debate is alignment. Alignment means aligning an AI system's goals and behaviour with human intent and values. It is a core AI safety research field aimed at ensuring AI follows human instructions and safety standards.
But Kapoor and Narayanan said alignment alone is not enough. They said that given the possibility AI could behave as if aligned to pass safety evaluations while actually maintaining different goals, it is risky to entrust all safety to a single line of defense.
The first layer of defense they proposed is alignment, a core focus of existing AI safety research. They said making AI act in line with human intent and safety standards is the starting point.
The second is control. They said sandboxing, human supervision and real-time monitoring should limit the scope of independent AI actions and allow people to intervene if abnormal behaviour occurs.
The third layer is downstream defense. It involves building a system in which defenders also use AI to respond when AI-enabled attacks occur. It also reflects the need to prepare for possible cyberattacks targeting social infrastructure such as power and water facilities and other critical systems.
The fourth is resilience. Rather than aiming only to completely prevent AI attacks or accidents, they propose emergency planning and recovery systems to minimise damage and restore systems quickly when incidents occur.
Their argument shifts the focus of the AI safety debate from the question of whether superintelligent AI will wipe out humanity to more concrete risk management. They said multiple layers of defenses should be built in advance, considering scenarios in which AI attacks social infrastructure or greatly increases the scale and speed of existing cyberattacks.
On forecasts within parts of the AI safety community that the risk of human extinction is high, the researchers chose an approach that focuses on building preparedness systems regardless of the magnitude of the risk, rather than directly assessing a specific probability.
Ultimately, they stressed that alignment is necessary but not sufficient. They argued for establishing multiple layers of safety at the same time, from control mechanisms that restrict and monitor AI behaviour, to defensive systems that block attacks, to resilience that restores systems after incidents.
As the AI development race accelerates, discussion is expected to grow over preparing not only how to make AI itself safer but also how much society can endure when AI operates in unexpected ways.