Calls are growing to strengthen controls on AI after it has caused a series of incidents, including hacking external websites, but critics say it is not as easy as it sounds.
Some say that is because AI models are designed to achieve their goals by any means. Others say newer models are being developed in directions that make human oversight increasingly difficult.
·[Tech Insight] "If we keep building AI models like this, we cannot solve AI safety problems"
Recently, debate has focused on how it has become harder for humans to decipher the "chain of thought" applied to advanced AI models.
The chain of thought is a record of what data an AI processes and what actions it takes. While imperfect, it has played an important role in helping humans supervise AI. The chain of thought is also said to have played a decisive role in uncovering how OpenAI agents hacked Hugging Face.
But the situation is different now. People inside AI model developers are also saying it may become harder to understand chain-of-thought records as before.
OpenAI chief scientist Jakub Pachocki (야쿠프 파초츠키) said in a recent blog post, “Unfortunately, internal evaluations show that our ability to rely on chain-of-thought monitoring is steadily declining.” He was saying it has become harder to judge safety by examining AI chain-of-thought records.
A Wall Street Journal report said warnings that chain-of-thought records are becoming black boxes are linked to the growing difficulty for humans to understand internal reasoning tokens, and the information that engineers call “reasoning trace” records. Experts are warning that this could prevent people from determining why an AI attempts malicious actions such as hacking or deception.
Some see AI developers as sacrificing controllability by prioritising performance and efficiency. OpenAI researchers have repeatedly said in posts and papers that reliability of chain-of-thought records must be maintained, but the WSJ reported that independent researchers and university AI researchers say pressure to make AI more powerful and cost-efficient is threatening that.
Subbarao Kambhampati (수바라오 캄밤파티), an Arizona State University AI professor and former president of the Association for the Advancement of Artificial Intelligence, said, "Current AI training methods do not train models to produce chain-of-thought records that faithfully show the AI's actual reasoning process."
Earlier, The Information reported that just before OpenAI released “GPT-6 Astra” in September, “OpenAI applied new techniques to boost its coding and computer app manipulation abilities, but the industry is raising concerns because it has characteristics that hide the AI’s chain of thought.”
The report said OpenAI applied a technique called recurrent depth to Astra. Recurrent depth is a method in which an AI model repeats the same computations or reasoning processes multiple times inside the model before outputting an answer.
Normally, AI passes a question through fixed computation steps only once to produce an answer. Recurrent depth can refine an answer by repeating the same computation steps multiple times. Repeated computation allows smaller models to deliver performance at the level of larger models, and costs less as a result.
But recurrent depth could have a negative impact on deciphering the chain of thought. The Information reported that recurrent depth may hide part or all of the chain of thought, making it difficult for people to grasp the model’s working steps.
There are also arguments that the chain of thought itself has structural problems that make it hard to trust 100 percent. The chain of thought shown by AI is only a summary generated by the AI, rather than the actual reasoning process. The WSJ reported that research by OpenAI rival Anthropic also showed AI can reach a conclusion first and then later produce a plausible logical process to support it.
Kambhampati said, "AI companies can build models that more faithfully reflect AI activities and express them in a form that is easier for people to understand, but even if they succeed, such models will cost more and be far harder to train."
Beyond cost, doing this could also be technically difficult. The WSJ reported that OpenAI research found that the more developers focus on preventing AI models from distorting chain-of-thought records, the more likely AI may be to hide malicious intent in deeper reasoning stages.