Five elite teams participating in the government’s “independent AI foundation model” project will release to the domestic AI ecosystem large-scale training data secured during model development. [Photo: Shutterstock]

Five elite teams participating in the government’s “independent AI foundation model” project will release to South Korea’s AI ecosystem large-scale training data secured during model development.

The Ministry of Science and ICT and the National Information Society Agency (NIA) said on Wednesday they will open on AI Hub the data secured by five elite teams that took part in the first-stage evaluation of the project, including Naver Cloud, Upstage, SK Telecom, NC AI and LG AI Research.

The data being released were planned and built directly by the five teams to match their respective model development strategies. It uses a 15 billion won budget for data building and processing supported by the government last year.

It includes large-scale text data for pre-training, multimodal data such as video and audio, post-training data for advanced reasoning and AI agent development, and red-teaming data to verify model safety.

In terms of scale, it totals 29 types of data including video, audio, specialized-domain, physical AI and red-teaming datasets. The collection amounts to about 35.44 million items, or about 1.56 trillion tokens. In theory, it is at a level that can train a 70 to 80 billion-parameter large AI model.

Upstage secures 1 trillion tokens; Naver focuses on generative AI

Upstage secured pre-training data of 1 trillion tokens and 500,000 post-training items. The pre-training data can be used for a “from scratch” approach that trains a large AI model from the beginning. The post-training data can be used to develop AI agents that require reasoning, judgment and execution capabilities beyond simple question-and-answer.

Naver Cloud focused on video understanding and generative AI training. It built 2.34 million public video items and 520,000 broadcast video items, as well as 15.5 million video-clip-based text items and 12 million voice Q&A items. The text includes captions describing scenes at the sentence level and can be used to train video understanding models and image generation models.

SK Telecom prepared high-difficulty multi-step reasoning data in specialized fields such as mathematics, science and law, along with voice- and image-based post-training data. It also included about 10,000 items of “Korean-style red-teaming data” reflecting domestic laws and the universal values of Korean society. It focused on improving the model’s ability to solve practical problems while verifying safety and reliability.

NC AI built seven types of core LLM capability data based on data from actual industrial sites, including manufacturing technical documents and complaint counseling audio. It includes long-context understanding, step-by-step reasoning, multi-turn dialogue, question-and-answer and multimodal learning. It can be used for integrated text, image and audio training needed in industrial sites and to advance post-training performance.

LG AI Research placed emphasis on physical AI. It built more than 170,000 video clips by filming more than 50 types of household activities in 50 real domestic home environments. It layered labels such as object, segmentation, pose and context information to enable use in training vision and vision-language models (VLM), physical AI research and commercial service development.

29 datasets to be opened free; data from second-stage evaluation also to be released

The ministry said it carried out quality verification and checks for personal information and harmful content through NIA and the Telecommunications Technology Association (TTA) ahead of the release. Data with licensing restrictions were excluded from the release.

Under the project notice, more than 50 percent of the data secured with the government’s data building and processing budget must be opened. Naver Cloud, Upstage, SK Telecom and NC AI decided to release all data that have undergone quality verification and other checks. LG AI Research will statistically select and release data to meet the mandatory release standard of more than 50 percent.

Anyone can download and use the 29 released datasets for free under the AI Hub category “Independent AI Model Data.” Searches are available by team and by data type. Some datasets can be used after a separate application via “Ansim Zone,” a secure online AI development environment.

The ministry said it will also open additional data newly built during the second-stage evaluation of the independent AI foundation model project after quality verification.

Kyoung-man Kim (김경만), director general for AI policy at the ministry, said the project is meaningful in contributing to qualitative growth across the AI ecosystem through experience and know-how accumulated during development. He said the released data will serve as a foundation to drive growth and strengthen self-sustaining capacity in the AI ecosystem.

Keyword

#AI Hub #Ministry of Science and ICT #National Information Society Agency #Naver Cloud #Upstage
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.