AI

Kakao Unveils MoE Training and Model Lightweighting Research at COLM 2026

Kakao presented a method to predict optimal learning rates for MoE pretraining at 1% of training cost, plus a byte-level tokenizer called BBT.

Kakao announced that it presented research results on training optimization for Mixture-of-Experts (MoE) models and on model lightweighting at COLM 2026, the international language modeling conference, and at its tokenization workshop TokShop, according to aitimes.com.

At the main conference, Kakao presented a methodology that predicts the optimal learning rate for pretraining large-scale MoE models at roughly 1% of the total training cost. The work applies mu-parameterization, a technique that transfers optimal training settings from small models to large models, to MoE architectures.

Kakao said the methodology was used in the pretraining of the Kanaana-2.6-155b-a17b model, which trained stably at a scale of 10 trillion tokens.

At TokShop, the company introduced BBT (BPE Guided Byte Transformer), a lightweighting technology that uses bytes as smaller units instead of the existing BPE method. Comparative experiments showed parameter counts reduced by 20.2-40.4% and performance improved by 2.7-6.6%.

BBT also showed improved robustness against typos and character perturbations, along with better transfer performance to unseen languages. Kakao expects the approach to reduce memory burden in on-device environments and to ensure stable performance in multilingual and typo-prone input environments.

Quick answers

What did Kakao present at COLM 2026?

Kakao presented a methodology that predicts the optimal learning rate for pretraining large-scale Mixture-of-Experts models at roughly 1% of the total training cost.

What is BBT?

BBT (BPE Guided Byte Transformer) is a lightweighting technology that uses bytes as smaller units instead of the existing BPE method. Experiments showed parameter counts reduced by 20.2-40.4% and performance improved by 2.7-6.6%.

Which model used Kakao's MoE training methodology?

The methodology was used in the pretraining of the Kanaana-2.6-155b-a17b model, which trained stably at a scale of 10 trillion tokens.

Source