Technology

Kakao Finds Optimal Value at 1% AI Training Cost… Applies Training to 10 Trillion Kanana Tokens

Major Breakthroughs in MoE Learning Optimization Unveiled at COLM 2026 Predicting the Optimal Learning Rate at About 1% of Total Training Costs Applied to Pre-training for ‘Kanana 2.6’… Stable Learning with 10 Trillion Tokens BBT Reduces Model Size by 40.4% and Improves Performance by Up to 6.6%

Lee So-Hyun
2026-10-08 09:37:41
[Edaily Reporter Lee So-Hyun ] Kakao(035720)has unveiled an optimization technology that reduces the training costs of large-scale artificial intelligence (AI) models while enhancing their stability. The company predicted the optimal training rate using only about 1% of the total training cost and applied it to the pre-training of its “Kanana” model. In a separate study, it reduced the model size by up to 40.4% while improving test performance by up to 6.6%.

Kakao announced on the 8th that it presented research findings on technology to improve the training efficiency of large-scale Mixture-of-Experts (MoE) models and on model lightweighting at the international language modeling conference “COLM (Conference on Language Modeling) 2026.”

COLM is an international conference specializing in large language models (LLMs) and language modeling research, established in 2024. Kakao presented its research findings at the conference’s main session and at the tokenization workshop “TokShop.”



Predicting the Optimal Learning Rate at Approximately 1% of the Training Cost

At the main conference, Kakao presented a methodology for predicting the optimal learning rate required for pre-training large-scale MoE models at a cost of approximately 1% of the total training cost.

The learning rate is a value that determines how much an AI model adjusts its internal parameters during the training process. If it is too high, training can become unstable; if it is too low, training speed may slow down, making the process of finding the appropriate value crucial.

MoE models, which select only the necessary parts from multiple expert modules to perform computations, are increasingly being used in high-performance LLMs. However, since there is relatively little published research on training settings compared to existing dense models, the large-scale training process can involve significant trial and error and cost burdens.

Kakao applied the “μ-parameterization” technique—which scales optimal training settings identified in small-scale models to large-scale models—to the MoE architecture.

By gradually increasing the model size and the number of training tokens, the team explored the optimal learning rate and verified whether the settings identified during short training intervals remained valid during large-scale training.

This method was actually applied to the pre-training of Kakao’s “Kanana-2.6-155b-a17b” model and was used to achieve stable training up to 10 trillion (10T) tokens.

Model Size Reduced by 40.4%, Performance Up by Up to 6.6%

At the tokenization workshop “TokShop,” researchers presented their “BBT (BPE-Guided Byte Transformer)” study, which improves performance while reducing model size.

Existing BPE (Byte Pair Encoding)-based models process text by dividing it into tokens in the form of words or word fragments.

BBT is designed to convert this into a structure based on “bytes”—an even smaller basic unit—while retaining the advantages of the existing BPE method, such as efficiently grouping and processing multiple characters.

In comparative experiments, it reduced the number of parameters by 20.2–40.4% while improving test performance by 2.7–6.6%.

Kakao explained that the model demonstrated more stable performance even in the presence of typos or character corruption, and that its transfer performance to untrained languages also improved.

Kakao believes that BBT will reduce the memory burden of models in on-device environments and increase their usability even in conversational settings where multilingual input or typos are common.

Kakao plans to expand the scope of this research to various model architectures and further refine the technology to enhance efficiency in multimodal and multilingual environments.

A Kakao spokesperson stated, “This research is significant in that it presents a practical method for reducing training and operational costs while maintaining or improving the performance of AI models,” adding, “We will expand our research to include various model architectures and multimodal and multilingual environments to increase the potential for real-world service applications.”

Economy

Corporation

IT·Science

Economy

LG Energy Solution, Third-Quarter Revenue of 9.6 Trillion Won Sets Record… Operating Profit Exceeds Forecasts by More Than Double (Comprehensive)

LG Energy Solution surpassed 9 trillion won in revenue in the third quarter of this year, achieving its highest quarterly revenue to date. Operating profit also reached 756 billion won, marking two co…
2026-10-08 09:45:33

Corporation

Kim Seong-yeon, Largest Shareholder of OSCOTEC Inc., Seeks Board Seat: “No Will to Change the Board; I Will Take Action Myself” [only-EDAILY]

Kim Seong-yeon, a director at Genosco and the only son of former OSCOTEC Inc. founder Kim Jeong-geun as well as the largest shareholder of OSCOTEC Inc.(039200), is moving to restructure the board of d…
2026-10-08 08:31:01

IT·Science

NEWEN AI Joins KOLON CORPORATION’s AI Ecosystem… Pushing Ahead with On-Site Commercialization

#NEWEN AI, a provider of industry-specific artificial intelligence (AI) analytics, is joining KOLON CORPORATION’s AI collaboration ecosystem. The two companies will combine their technologies and cust…
2026-10-08 10:12:17