Evolution and Adaptation of Large Language Models for Bahasa Indonesia

Authors

  • Allan Desi Alexander Universitas Bhayangkara Jakarta Raya, Jakarta, Indonesia
  • Siti Setiawati Universitas Bhayangkara Jakarta Raya, Jakarta, Indonesia

DOI:

https://doi.org/10.38035/dit.v4i1.3623

Keywords:

Language Adaptation, Bahasa Indonesia, Large Language Models, Vocabulary Expansion, Continual Pre-Training, Cultural Alignment

Abstract

This study evaluates the systematic evolution and computational adaptation of pre-trained language models and Large Language Models (LLMs) for Bahasa Indonesia and its low-resource regional dialects. Initially centered on bidirectional encoder-based representations like IndoBERT, the regional natural language processing (NLP) field has transitioned toward generative sequence-to-sequence structures and massive decoder-only architectures. This paper investigates the engineering methodologies of cross-lingual vocabulary adaptation, parameter initialization heuristics, and language-adaptive pre-training strategies designed to address text overfragmentation, representational misalignment, and tokenization cost inefficiencies. Through extensive structural benchmarks, this analysis compares discriminative and generative performances across tasks including sentiment classification, extractive question answering, text style normalization, domain-specific retrieval-augmented pipelines, and entity linking. While localized generative models such as Komodo, Sailor, and the SEA-LION suite improve contextual reasoning, colloquial style transfers, and regional dialect preservation, they remain susceptible to architectural anomalies like template leakage and entity hallucination. This study provides foundational benchmarks and methodological frameworks for adapting massive language models to morphologically rich, culturally diverse, and low-resource linguistic environments.

References

Cahyawijaya, S., Winata, G. I., Wilie, B., Vincentio, K., Li, X., Kuncoro, A., Ruder, S., Lim, Z. Y., Bahar, S., Khodra, M. L., Purwarianti, A., & Fung, P. (2021). IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation. EMNLP 2021 - 2021 Conference on Empirical Methods in Natural Language Processing, Proceedings, 8875–8898. https://doi.org/10.18653/v1/2021.emnlp-main.699

Chakrabarty, A., Chaturvedi, A., & Garain, U. (2020). NeuMorph. ACM Transactions on Asian and Low-Resource Language Information Processing, 19(1), 1–19. https://doi.org/10.1145/3342354

Chizhov, P., Arnett, C., Korotkova, E., & Yamshchikov, I. P. (2024). BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training. EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, 16587–16604. https://doi.org/10.18653/v1/2024.emnlp-main.925

Di Marco, M., & Fraser, A. (2024). Subword Segmentation in LLMs: Looking at Inflection and Consistency. EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, 12050–12060. https://doi.org/10.18653/v1/2024.emnlp-main.672

Ding, W., Wang, W., Kwok, S. H. D., Liu, M., Fang, T., Bai, J., Liu, X., Yu, C., Li, Z., Luo, C., Yin, Q., Yin, B., He, J., & Song, Y. (2024). IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Language Models in E-commerce. Findings of the Association for Computational Linguistics: EMNLP 2024, 2247–2266. https://doi.org/10.18653/v1/2024.findings-emnlp.123

Fridkin, S., & Bendersky, M. (2026). Interpretable Machine Learning: A Comprehensive Review of Foundations, Methods, and the Path Forward. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 16(1). https://doi.org/10.1002/widm.70075

Fujii, K., Nakamura, T., Loem, M., Iida, H., Ohi, M., Hattori, K., Shota, H., Mizuki, S., Yokota, R., & Okazaki, N. (2024). Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities. http://arxiv.org/abs/2404.17790

Hasmawati, & Ade Romadhony. (2023). Similar Questions Identification on Indonesian Language Subject Using Machine Learning. Jurnal Nasional Pendidikan Teknik Informatika (JANAPATI), 12(2), 196–202. https://doi.org/10.23887/janapati.v12i2.62582

Jap, B. A. J., & Arumsari, C. (2017). Adaptation of the Token Test in Standard Indonesian. Makara Human Behavior Studies in Asia, 21(1), 44. https://doi.org/10.7454/mssh.v21i1.3499

Jin, B. (2024). Multilingual Neural Machine Translation with Integrated Language Adapters. Proceedings of the 2024 International Conference on Generative Artificial Intelligence and Information Security, 29–35. https://doi.org/10.1145/3665348.3665355

Jo, E., Cho, E., Lee, Y., Song, S., & Joo, H. J. (2025). Domain and Language adaptive pre-training of BERT models for Korean-English bilingual clinical text analysis. BMC Medical Informatics and Decision Making, 25(1). https://doi.org/10.1186/s12911-025-03262-7

Kartika, B. V., Alfredo, M. J., & Kusuma, G. P. (2023). Fine-Tuned IndoBERT based model and data augmentation for indonesian language paraphrase identification. Revue d’Intelligence Artificielle, 37(3), 733–743. https://doi.org/10.18280/ria.370322

Lee, K. S. J. W., Qorib, M. R., Soegeng, A. I., & Ng, H. T. (2026). Equity with Efficiency: An Empirical Study of Tokenizers for Multilingual Large Language Models. http://arxiv.org/abs/2606.15044

Linder, M. R. (2026). Vocabulary Expansion of Large Language Models via Kullback-Leibler-Based Self-Distillation. http://arxiv.org/abs/2508.15807

McGiff, J., & Nikolov, N. S. (2025). Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review. http://arxiv.org/abs/2505.04531

Mu, L., Wang, X., Ni, L., Li, Y., Wu, Z., Jin, P., & Zhang, Y. (2025). DenseLoRA: Dense Low-Rank Adaptation of Large Language Models. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 1, 10198–10211. https://doi.org/10.18653/v1/2025.acl-long.503

Owen, L., Tripathi, V., Kumar, A., & Ahmed, B. (2024). Komodo: A Linguistic Expedition into Indonesia’s Regional Languages. http://arxiv.org/abs/2403.09362

Petrov, A., La Malfa, E., Torr, P. H. S., & Bibi, A. (2023). Language Model Tokenizers Introduce Unfairness Between Languages. http://arxiv.org/abs/2305.15425

Provilkov, I., Emelianenko, D., & Voita, E. (2020). BPE-dropout: Simple and effective subword regularization. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 1, 1882–1892. https://doi.org/10.18653/v1/2020.acl-main.170

Ridwan, M. (2018). National and Official Language: The Long Journey of Indonesian Language. Budapest International Research and Critics Institute (BIRCI-Journal) : Humanities and Social Sciences, 1(2), 72–78. https://doi.org/10.33258/birci.v1i2.14

Subhan Mahendrasyah, M., & Hariguna, T. (2024). Analisis Sentimen Pengguna Aplikasi Bukalapak di Platform Playstore Menggunakan Metode Naïve Bayes. Building of Informatics, Technology and Science (BITS), 6(2), 733–745. https://doi.org/10.47065/bits.v6i2.5528

Xu, Y., & Che, W. (2026). Evolving LLMs from Next-Token Prediction to Multi-Token Prediction via Self-Distillation. Electronics (Switzerland), 15(7). https://doi.org/10.3390/electronics15071533

Yamaguchi, A., Villavicencio, A., & Aletras, N. (2026). How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text? Computational Linguistics, 52(1), 295–330. https://doi.org/10.1162/COLI.a.581

Zhang, L., Lou, Z., Ying, Y., Yang, C., & Zhou, H. (2025). Efficient Fine-Tuning of Large Language Models via a Low-Rank Gradient Estimator. Applied Sciences (Switzerland), 15(1), 1–16. https://doi.org/10.3390/app15010082

Downloads

Published

2026-07-20

How to Cite

Alexander, A. D., & Setiawati, S. (2026). Evolution and Adaptation of Large Language Models for Bahasa Indonesia . Dinasti Information and Technology, 4(1), 85–93. https://doi.org/10.38035/dit.v4i1.3623