Kakao CI [Photo: Kakao]

Kakao unveiled efficiency technologies at an international conference that reduce training costs for large AI models and cut model size.

Kakao said on Wednesday it presented training optimisation techniques for large mixture-of-experts (MoE) models and results from its model lightweighting research at the international language modelling conference COLM 2026.

COLM is an international conference on large language models (LLMs) and language modelling that was launched in 2024. Kakao introduced its research results at the main conference and at a tokenisation workshop, Tokshop.

At the main conference, it presented a methodology to predict the optimal learning rate for large MoE models at about 1 percent of total training cost. The learning rate determines how much parameters are adjusted during model training.

Kakao applied a "mu-parameterisation" technique, which scales training settings found in small models to large models, to suit the MoE architecture. It increased model size and the number of training tokens in stages to search for the optimal learning rate, and verified whether settings secured over a short training segment are maintained in large-scale training.

The method was also applied to Kakao's pretraining of "Kanana-2.6-155b-a17b". The company said it enabled stable training up to 1 trillion tokens.

At Tokshop, it unveiled a "BBT" structure that boosts performance while reducing model size. While existing byte pair encoding (BPE)-based models store word or word-piece information, BBT was designed to use bytes, a smaller unit, while retaining the advantages of processing multiple characters as a bundle.

In comparative tests, the number of parameters fell 20.2 to 40.4 percent and test performance improved 2.7 to 6.6 percent. Kakao said robustness to typos and character perturbations also improved, along with transfer performance to languages not seen in training.

Kakao sees BBT as potentially useful for reducing memory burdens in on-device environments and delivering stable performance in multilingual inputs or conversation settings where typos are frequent.

It plans to expand the scope of application to various model architectures as well as multimodal and multilingual environments, and further reduce training costs for large LLMs.

A Kakao official said, "This research suggests a way to reduce training and operating costs while maintaining or improving AI model performance," and added, "We will expand research to various model architectures and multimodal and multilingual environments to raise the possibility of using it in real services."

Keyword

#Kakao #COLM 2026 #MoE #BBT #Kanana-2.6-155b-a17b
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.