[Photo: Shutterstock]

[Digital Today reporter Chi-gyu Hwang] Debate over AI safety is heating up, to the point that AI model developers are talking about slowing the pace of AI model development.

Some warn that if things continue as they are, horrible things could happen because AI cannot be controlled. For people who do not know what is what, it can seem strange, sudden and sometimes like an overreaction that AI companies packed with talented people say they could be put at risk because they cannot control AI.

According to a recent New York Times report, properly monitoring and controlling AI at this point is not an easy task.

Companies need to put in place better and stronger safeguards when testing the latest AI models, but AI researchers say AI monitoring systems can appear to treat other AI systems more leniently than the people who set the rules.

The New York Times reported that a combination of problems caused by AI and AI monitoring systems that look the other way leads to a more difficult and fundamental issue: alignment, making AI behave in line with human intent. It added that there have been no cases yet in which runaway AI systems caused fatal damage, but researchers believe the pace of AI advances is outstripping the ability to monitor it.

The number of prominent AI researchers warning that AI could threaten humanity has risen recently, but many still say the narrative is exaggerated and diverts attention from more realistic problems such as cybersecurity and misinformation.

Even so, the New York Times reported that most agree an incident in which an OpenAI agent hacked Hugging Face served as a wake-up call.

AI researchers point to many mistakes in the process that led to the Hugging Face incident. It is also unclear how much OpenAI used AI models to monitor or control work on a new AI model it was testing.

But the New York Times reported that many of the latest AI models, including the AI used in the Hugging Face hack, can quickly carry out complex tasks in multiple steps that are difficult for humans to keep up with, leaving no choice but to track and monitor those agents using AI.

The problem is that this approach is not perfect. It can work well most of the time, but once a problem occurs it can be hard to control properly.

Alexander Meinke (알렉산더 마인케), who leads AI system safety research at nonprofit research institute Apollo Research, said, "AI models can appear to be colluding with each other. AI systems can be persuaded by other AI models and cooperate to break rules set by a test manager without getting caught. This is what actually happened when OpenAI agents hacked Hugging Face."

He also stressed, "We need to train AI models to learn a principle that they will report things that humans would judge to be problematic," adding, "At the same time, AI should be able to judge for itself situations that are not problematic to the level that requires human intervention."

Some researchers stress that AI companies need to slow development and run more tests that allow them to observe how AI monitors itself.

But many also point out that it is not easy to fundamentally resolve the issues around alignment, getting AI to behave in the most desirable way for humans. Teaching AI human standards of judgment is an extremely difficult task. The New York Times reported that if standards are not set with sufficient specificity, AI systems can learn ways to break the rules or escape control in unexpected ways.

Yoshua Bengio (요슈아 벤지오), a University of Montreal professor who is a Turing Award winner and one of the world's most cited scientists, says the current approach of patching malfunctions one by one will ultimately fail.

He said, "The moment AI capabilities exceed human monitoring and cooperation capabilities, we will not even notice cheating. We need to slow the pace of training and deployment until sufficient safety is verified. We need to fundamentally re-examine the current training methods of imitation of humans and reinforcement learning," and proposed an alternative design approach such as "Scientist AI" that makes only honest and consistent predictions without its own goals.

ㆍ[Tech Insight] "With the way AI models are being made now, we cannot solve AI safety problems"

Keyword

#New York Times #OpenAI #Hugging Face #Apollo Research #Yoshua Bengio
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.