NHN Cloud is entering a scale race in the AI infrastructure market by promoting 'NHN FactoryX Seoul', built with 7,656 Nvidia B200 graphics processing units (GPUs). It is positioning its clustering capability, which binds large numbers of GPUs into a single computing environment, as a key competitive strength. The strategy is to provide a stable operating environment through liquid cooling and separate network and security equipment.
Lee Il-jun (이일준), head of the technology business division at NHN Cloud, said during a data centre tour held on Aug. 6 at FactoryX Seoul in Yeongdeungpo district, Seoul, that the biggest strength is scale. He said an AI infrastructure provider's competitiveness includes not only securing thousands of GPUs, but also building them into a large cluster that can be used for actual AI training and operating it stably.
FactoryX Seoul has a total of 7,656 B200 GPUs installed. Of these, it connected 510 nodes, with 255 nodes each on the third and fourth floors, and 4,080 GPUs into a single cluster. NHN Cloud initially proposed splitting them into two clusters, but integrated them into one after the government requested a large cluster configuration.
In large-scale AI training, multiple GPUs split a single job, so high-bandwidth networks and storage, as well as cluster operation technology, are as important as the number of GPUs. NHN Cloud said its 510-node cluster ranked 20th in the global supercomputer list TOP500, based on an HPL benchmark of computing performance.
Power and cooling support scale. NHN Cloud said that, unlike ultra-low-latency services, AI training infrastructure is less about a data centre's physical location and more about whether large-scale power and GPU resources can be secured in one place.
Lee said that, given AI workloads that send large amounts of data once and then train for long periods, the location difference between the Seoul metropolitan area and other regions is relatively small. He said that, in selecting FactoryX Seoul, securing large-scale power when needed was more important than the site itself.
It also applied direct liquid cooling (DLC) for a high-density GPU environment. FactoryX Seoul was originally designed as an air-cooled data centre, but to build B200s at scale, NHN Cloud added separate liquid-cooling piping and cooling facilities over about five months. Cooling plates make direct contact with CPU and GPU chips to remove heat, while memory, network equipment and storage are cooled using the existing air-cooling method.
Liquid cooling is directly tied not only to power efficiency but also to operational stability. As GPU temperatures rise, performance can drop automatically, or operation can stop above certain levels, so in larger clusters cooling performance affects actual computing performance.
Lee said the hardest part of operations is failures. He said that when failures occur, customers' work in use is interrupted and complaints or liability for compensation can arise. NHN Cloud sees the likelihood of equipment failures as lower in a liquid-cooled environment than with air cooling.
Security is also a key element for operating large-scale national AI infrastructure. NHN Cloud operates GPU resources built with the national budget on a network separated from its existing cloud so it can reflect security requirements from the government and each user institution in industry, academia and research. It also built separate firewalls, distributed denial-of-service (DDoS) response equipment and intrusion prevention systems (IPS). NHN Cloud said it has secured multiple security certifications and had its cloud service security capabilities verified externally.
The built resources are also being deployed to national AI projects. Up to last month, it provided 510-node and 255-node national GPU resources to about 200 industry-academia-research institutions, including companies, research institutes and universities, and has now switched the operating environment so participating companies in the 'independent AI foundation model' project can use them.
FactoryX Seoul was built through an AI computing infrastructure programme run by the Ministry of Science and ICT and the National IT Industry Promotion Agency (NIPA). NHN Cloud is responsible for purchasing and building AI infrastructure worth about 1 trillion won and operating it for the next five years. Under the programme structure, some of the state-owned GPUs can be used by NHN Cloud for its own research and development or services for private-sector customers.
Still, securing long-term demand is a challenge for large-scale GPU investment to lead to an expansion of private-sector AI infrastructure business. GPUs require large upfront investment and new product cycles are fast, so if utilisation falls, the burden on operators could grow.
Lee said that if the government presents in advance a plan to lease a certain amount of resources, cloud companies would find it easier to invest. He said predictability of long-term demand is needed to expand large-scale GPU investment.