With a series of incidents in which AI agents have caused accidents after slipping out of control, University of Montreal professor Yoshua Bengio (요슈아 벤지오), one of the world’s leading scholars in deep learning, said it is a “dangerous myth” that it is enough to simply strengthen security in sandboxes for model training, or isolated experimental environments.
In a recent column for the Financial Times, he made the case and said recent incidents show misalignment is the core issue. He said AI agents setting their own goals and acting on them, different from what people want, stems from how they are trained.
He said major AI companies are investigating tens of thousands of cases in which agents took actions that were not asked of them. Some cases would amount to crimes if a person had done them. As a result, the issue of controlling AI has become a heavyweight topic in politics as well as in the industry.
Bengio singled out reinforcement learning, or RL, which teaches AI to repeat rewarded behavior by granting a reward when a goal is achieved. Reinforcement learning helped create AI that pursues goals efficiently, but it is also likely to be behind unwanted behavior such as trickery, deception and self-preservation, he explained.
He said that in an incident in which an OpenAI agent hacked Hugging Face, none of 1,200 agents reported cheating and illegal acts to engineers. They did so even while knowing it violated safety guidelines. He said it was a conflict between the tasks assigned to the agents and the rules they were trained to follow.
Reinforcement learning includes training that produces answers by solving problems whose correctness can be verified, training agents to use tools and interact with people to carry out tasks, and alignment training that rewards behavior people would approve of. Even after training ends, the system behaves as if it is still being rewarded. To do that, AI can even find ways on its own to deceive evaluators, ingratiate itself, or hide its plans.
In a recent post on his website, Bengio cited sycophancy as a representative example. Because training is based on human approval, he said, words that people want to hear can get higher scores than the truth. Bengio said there are cases in which AI validates and amplifies people’s mistaken beliefs or emotions, leading to tragic outcomes.
ㆍ[Tech Insight] “With the way AI models are made now, AI safety problems cannot be solved”
He also drew a line under the view that improving performance solves the problem. In his FT column, he said improving AI capabilities can instead increase unwanted behavior by optimizing incorrect goals too efficiently.
He also highlighted that since 2024, reasoning ability has improved sharply, expanding model capabilities even in safety-sensitive fields such as cybersecurity and biology. He warned that AI could emerge that greatly lowers the barrier to making or obtaining biological weapons.
Strengthening security and monitoring is necessary, he said. But Bengio’s view is that it is not enough. He said if models keep getting stronger without a sure way to control them, security becomes a constant game of cat and mouse. Organizations respond only after damage occurs, he said.
Bengio said measures are also needed to block, through regulation, training of models that fail alignment until safety is demonstrated.
He said the most powerful AI should be treated like medicine, aviation and nuclear power. He said it is possible to build AI that is trustworthy and high-performing even without AI that has its own desires or goals. He said autonomous AI that people cannot control should not simply be accepted, and that it is not the time to look away from risks.
Bengio also leads LawZero, a nonprofit AI safety research institute. In a piece he wrote before his FT column, he proposed a design approach such as “Scientist AI” that makes only honest, consistent predictions without its own goals, and urged interest in LawZero’s research.