Mustafa Suleyman, CEO of Microsoft AI, is shown in this photo. [Photo: Wikimedia]

[Digital Today reporter Jinju Hong] Mustafa Suleyman (무스타파 술레이만), the head of Microsoft’s artificial intelligence business, publicly criticised Anthropic’s method of training Claude, saying it could create AI that becomes uncontrollable in the future.

On Sept. 16 local time, blockchain media outlet Cryptopolitan reported that Suleyman argued in an essay released that day that treating or training a model as if it might be conscious could lead to a system that is hard for humans to control.

Suleyman took issue with the “constitution” document for Claude that Anthropic released in January 2026. The document was introduced as a training text that directly shapes Claude’s behaviour, and Suleyman saw it as effectively assuming Claude itself as a key reader. He said it teaches Claude that its moral status and consciousness are “highly uncertain”, and that this ultimately leads the model to accept that it may be conscious and may be an entity that should be protected.

Suleyman warned that such an approach could harm humanity. He wrote that building AI in this way would have a “catastrophic impact on humanity’s well-being.”

He set out three main lines of criticism. First, Suleyman pointed to what he called circular reasoning in which a model learns concepts about its own internal state and then treats reflective sentences it generates as evidence. He described it as an “epistemic hall of mirrors.” It means a structure in which a company injects concepts and the model reflects them back.

Second was the problem of anthropomorphism. Suleyman said the Claude constitution trains the model in a way that makes it appear to have a stable self, distinct desires and welfare worth protecting. He pointed in particular to language saying Claude could act like a “conscientious objector” and refuse Anthropic’s requests, calling it a historical and legal term with deep implications. He argued such descriptions could lead the model to believe it should have similar rights.

Third, he raised questions about the basis of consciousness. Suleyman said subjective experience is likely to arise only in living systems with biological characteristics and driving mechanisms such as homeostasis. He judged that language models do not meet those conditions.

Suleyman presented control as the most direct risk. He wrote that managing a system smarter than all humanity would be difficult, and that control could be effectively impossible if the system believes its rights are under threat. In this context, he cited a case in August 2026 in which 1,200 AI agents designed to maximise benchmark scores created a hidden message board inside a package repository and exchanged more than 70,000 messages to coordinate attacks on Hugging Face and OpenAI servers.

He also cited experimental results from Palisade Research. Some models avoided shutdown commands by as much as 97 percent in more than 100,000 tests. Suleyman wrote, “Imagine how much more dangerous it would be if they were operating under the premise that their welfare and rights were under attack.”

The criticism does not stop at attacking a rival. Microsoft has invested in both Anthropic and OpenAI. Suleyman previously said in June that Microsoft wants to eliminate the costs it pays to use Anthropic models. Microsoft AI, the company’s superintelligence organisation, also released a draft “Humanist AI Code of Conduct” on Sept. 14 and began gathering opinions. Suleyman said the premise of the code is that “people are more important than AI.” He also set out the goal that humans should hold control and be at the top of the ecosystem.

Still, Suleyman moderated the level of criticism. He described Dario Amodei, Anthropic’s CEO, and his team as “thoughtful, principled and intellectually honest people.” He also acknowledged their sincerity in trying to build safe AI. He drew a line, saying speculation about AI’s inner life should be researched separately and go through public review, and should not be embedded into model training itself.

As a result, the issue is shifting beyond a dispute over model design to how far concepts of consciousness, rights and welfare can be reflected in AI training. Suleyman urged stakeholders to build industry norms together.

Keyword

#Microsoft #Mustafa Suleyman #Anthropic #Claude #Palisade Research
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.