AI Automated Translation.

Font Size

Share

Kakao, AI Model 'Uses Less, Learns Better'… Unveils Efficiency Technology at COLM 2026

Kakao, AI Model 'Uses Less, Learns Better'… Unveils Efficiency Technology at COLM 2026

Predicts Optimal Learning Rate for MoE to Cut Costs… Reduces Model Size by Up to 40.4% While Improving Performance

Kakao has unveiled research results that reduce the training costs of large-scale AI models and lower model size while enhancing performance.

On the 8th, Kakao announced that it presented technologies for improving the training efficiency of Mixture-of-Experts (MoE) models and model lightweighting at the international conference on language modeling, 'COLM 2026.'

COLM is an international conference specializing in large-scale language models and language modeling research, established in 2024. Kakao introduced its research results at the conference's main conference and the tokenization workshop 'TokShop,' respectively.

At the main conference, it presented a methodology for predicting the optimal learning rate required for pre-training large-scale MoE models at a level of approximately 1% of the total training cost.

The learning rate is a value that determines how much an AI model adjusts its parameters during the training process. If the value is too high, training becomes unstable; if it is too low, the training speed may slow down, making it a crucial setting in large-scale model training.

Kakao applied the 'muP (mu-parameterization)' technique, which extends optimal learning settings found in small-scale models to large-scale models, to fit the MoE structure. It explored the optimal learning rate by gradually increasing model size and training token volume, and verified whether settings found in short training intervals remained valid for large-scale training.

This methodology was actually applied to the pre-training of Kakao's 'Kanana-2.6-155b-a17b' model. Kakao explained that through this, it achieved stable training up to 10 trillion tokens.

It also unveiled lightweighting research aimed at reducing model size. At TokShop, Kakao presented 'BBT: BPE-Guided Byte Transformer.' While the existing BPE method stores word or subword information to process sentences, BBT is designed to utilize smaller units—bytes—while maintaining the advantages of existing compression methods.

In comparative experiments, it reduced the number of parameters by 20.2% to 40.4% while improving test performance by 2.7% to 6.6%. Robustness against typos and character perturbations, as well as transfer performance to untrained languages, were also improved.

Kakao believes that BBT can be utilized to reduce the memory burden of AI models in on-device environments and to achieve stable performance even in conversational environments with frequent multilingual input or typos.

Going forward, Kakao plans to expand the scope of research application to various model structures and further reduce the training costs of large-scale LLMs.

A Kakao official stated, "The significance of these studies lies in presenting practical solutions that can lower training and operational costs while maintaining or improving AI model performance," adding, "We will expand our research to various model structures and multimodal and multilingual environments to increase the potential for actual service application and strengthen global competitiveness."

"This article was translated using AI and may differ slightly from the original."