Many of China’s most advanced artificial intelligence (AI) models are still being trained on Nvidia chips, a report said.
On Aug. 10, the Hong Kong-based South China Morning Post (SCMP) reported that while the Chinese government is pushing semiconductor self-reliance, local AI developers cannot easily change their training infrastructure because of the cost and technical burden of switching to domestic chips.
The key obstacle was cited as the software ecosystem. Nvidia’s CUDA has long been established as a standard for AI development, but moving to Huawei’s alternative platform, CANN, requires large-scale code rewrites and optimisation, it said. A source familiar with industry conditions said, "Among Chinese AI developers, it is still common to train large language models (LLMs) on Nvidia chips."
James Wang (제임스 왕), who develops AI models at a research institute affiliated with a university in Shanghai, said existing training pipelines depend on CUDA. "CUDA code cannot run directly on Ascend and requires extensive rewriting," he said. "Moving existing workflows to Huawei Ascend chips could take at least 50 percent more time and cost."
The burden of switching grows depending on model architecture and how much is made public. An engineer in Beijing who took part in moving LLM work to Ascend said open-source models such as DeepSeek and their distilled models can use support from the existing ecosystem, making training possible if 2 to 3 additional engineers are assigned for about a month compared with Nvidia-based systems. Models such as Moonshot AI’s Kimi K3, which do not 공개 source code and only release weights, may require about 10 engineers to do more than 6 months of additional work, the engineer said.
The trend shows that domestication in China’s AI industry is advancing first in inference rather than training. The media outlet noted that AI development is divided into training and inference, with training complex and resource-intensive while inference is relatively easier to migrate. DeepSeek-V4 and Moonshot’s Kimi K3 have been successfully adapted for Huawei and Alibaba Group Holding platforms, the report said. Alibaba Group Holding owns SCMP.
Which chips were used in the training stage remains sensitive. DeepSeek and Moonshot did not disclose the AI chips used to train their models. Kimi K3 was reported in July by U.S. outlet The Information to have been trained on chips including Nvidia’s Blackwell processors. The report also raised that even as Chinese AI companies try to expand domestic alternatives, they still rely on Nvidia hardware for training front-line models.
Attempts to train on domestic chips are not absent in China. Meituan said in June, when it unveiled Longcat-2.0 with 1.6 trillion parameters, that it completed both training and running the model on a domestic computing cluster of 50,000 units, without naming the supplier.
This has again highlighted that the key to China’s AI semiconductor self-reliance lies less in chip performance itself than in the cost of switching development ecosystems. Domestic chips are broadening their use in inference, but how quickly a CUDA-centric development environment can be replaced for training cutting-edge models is expected to be a future dividing line in competitiveness.